Make room for the unfinished thought

A pause is not always permission to act.

Google speech · Design account, with new interaction studies

People pause to find a word, change their minds, or add the part that matters. An assistant that treats every pause as an endpoint makes the person work around its timing.

I was the lead designer for a nascent Speech team at Google. My work covered interruption, partial fulfillment, and helping people speak at their own pace. Two decisions capture the problem: how to show an unfinished request, and how long to leave room before acting.

Read the original Google speech account

Two decisions from the original work

These decisions are documented in the original account; not every exploration is presented as shipped.

Try the timing decision

“Add dried mango … with no sugar added.” Act at the pause and the person has to repair the request. Wait without feedback and they may wonder whether they were heard. My current recommendation is to hold the change, make listening visible, and ask a specific question if the pause continues.

Reconstructed from earlier design material. These studies illustrate the tradeoff; the recommendation still needs evaluation in a live system.

A separate question: how should it sound?

These two stock voices read the same script in my current VoiceEnroll prototype. Listen for the difference in delivery. This is an illustrative voice audition, separate from the historical work and the timing studies above.

Candidate A

Azure Speech · en-US-AndrewNeural

Candidate B

Azure Speech · en-US-AvaNeural

Transcript for both clips: “Hello. This is a sample of how I sound. You can pick the voice that fits you best, and change it any time.”

A comparison instrument, not a verdict

I built a voice-comparison tool that lays candidates side by side with the same script, pitch contours, waveforms, spectrograms, and synthesis cost. That structure makes a difference easier to locate and discuss without pretending the chart decides which voice is right.

Voice-comparison tool with three voice candidates, pitch tracks, waveforms, spectrograms, playback controls, and notes.
The comparison tool keeps candidates at the same scale and exposes synthesis cost. Word timing is marked unavailable rather than estimated. This capture is static.

Pitch tracks can point to a pause or contour worth hearing again. They cannot establish warmth, trust, or fit for a task. That judgment needs listeners in the relevant context, with interruption and recovery included in the test.

Copilot Vision

I helped ship vision for Microsoft Copilot. Over the past year, my role has been Voice AI Designer.

Case account in development. Specific design decisions, collaborators, and outcomes will be added when they can be stated with supporting evidence.

Sources, conditions & what remains to test

Conditions: the same English script; default voice delivery with no requested style, rate, pitch, or emphasis changes; mono MP3 at 24 kHz and 96 kbps. The repository generator identifies Azure Speech as the source; the generation date was not preserved. This page does not record a preference or collect research data.

Evidence still needed: listener observations from a matched, task-relevant evaluation; the selection criteria and final product decision; and documented outcomes. The clips above demonstrate a comparison method, not a winning voice. The earlier studies do not establish how these two synthetic voices perform in use.

Separate contemporary reference: Siri recording study.

Related: Learning without leaving the work