SaaS Product (B2C)
Role:
Year:
Explore the full story –
Challenge
Speaking is faster than typing. It works while you are driving, walking between meetings, or thinking out loud, and it does not ask you to slow your thoughts down to the speed of your hands. Dictation solved that part. It did not solve what comes after. Every tool hands back a transcript, which is just another wall of text you now have to read, sort, and retype somewhere useful. That cost lands on everyone who captures by voice, and it lands hardest on people who read slowly or have dyslexia. So people end up running four apps to do one job: one to capture voice, one for notes, one for tasks, one for meetings. Nothing joins them up.
Objective
One capture, four outputs, and a transcript you can listen to instead of read. Read aloud sits above the transcript rather than buried in a menu, and it reports position as "Line 4 of 18" instead of a percentage, because a reader following along needs to know which line. An extracted task is marked with a filled pill. The line currently being read gets a left edge bar. Two different mechanisms, so a line can be both and stay legible. Anything the AI wrote sits on a tinted panel with an accent edge, so it is always separable from your own words.
Results
Sixteen screens across desktop and mobile, built on twenty components and a three-tier token system. Four apps collapse into one capture. A transcript can be listened to rather than read. Tasks leave the conversation without being retyped. Speaker detection is least accurate in the first seconds, before the model has a voiceprint to compare against, so that turn is flagged as low confidence and can be corrected. The correction then applies to every turn with the same voice. Designing what happens when the model is wrong, rather than only the happy path, is the decision I would defend first.






