SaaS Product (B2C)

Live Project & Prototype

Mumble is a voice-first workspace for people who work by talking. You hit record, talk, and it transcribes in real time, pulls out the action items, applies tags, and writes a summary. Then it reads the whole thing back to you, if listening is easier than reading. It handles solo capture and live meetings, and it separates speakers as they talk. The design leads with clarity. Anything the machine generated stays visually distinct from your own words, and the app never presents a guess as a certainty.

Mumble is a voice-first workspace for people who work by talking. You hit record, talk, and it transcribes in real time, pulls out the action items, applies tags, and writes a summary. Then it reads the whole thing back to you, if listening is easier than reading. It handles solo capture and live meetings, and it separates speakers as they talk. The design leads with clarity. Anything the machine generated stays visually distinct from your own words, and the app never presents a guess as a certainty.

Mumble is a voice-first workspace for people who work by talking. You hit record, talk, and it transcribes in real time, pulls out the action items, applies tags, and writes a summary. Then it reads the whole thing back to you, if listening is easier than reading. It handles solo capture and live meetings, and it separates speakers as they talk. The design leads with clarity. Anything the machine generated stays visually distinct from your own words, and the app never presents a guess as a certainty.

Role:

Product Designer

Product Designer

Year:

2026

2026

  • Explore the full story –

Challenge

Speaking is faster than typing. It works while you are driving, walking between meetings, or thinking out loud, and it does not ask you to slow your thoughts down to the speed of your hands. Dictation solved that part. It did not solve what comes after. Every tool hands back a transcript, which is just another wall of text you now have to read, sort, and retype somewhere useful. That cost lands on everyone who captures by voice, and it lands hardest on people who read slowly or have dyslexia. So people end up running four apps to do one job: one to capture voice, one for notes, one for tasks, one for meetings. Nothing joins them up.

Objective

One capture, four outputs, and a transcript you can listen to instead of read. Read aloud sits above the transcript rather than buried in a menu, and it reports position as "Line 4 of 18" instead of a percentage, because a reader following along needs to know which line. An extracted task is marked with a filled pill. The line currently being read gets a left edge bar. Two different mechanisms, so a line can be both and stay legible. Anything the AI wrote sits on a tinted panel with an accent edge, so it is always separable from your own words.

Results

Sixteen screens across desktop and mobile, built on twenty components and a three-tier token system. Four apps collapse into one capture. A transcript can be listened to rather than read. Tasks leave the conversation without being retyped. Speaker detection is least accurate in the first seconds, before the model has a voiceprint to compare against, so that turn is flagged as low confidence and can be corrected. The correction then applies to every turn with the same voice. Designing what happens when the model is wrong, rather than only the happy path, is the decision I would defend first.