Transcription on its own is a commodity. The value is in what happens next — turning a rambling two-minute voice note into a structured set of tasks, or a searchable log.

1. Record in the browser

The media recording API captures audio without any plugin. The pieces that catch people: asking for microphone permission properly, handling refusal gracefully, and picking a format the transcription service accepts.

voice-capture.txt
Build browser audio capture for transcription.

- Request microphone permission at the moment of use, not on page
  load, and handle refusal with a clear explanation.
- Record in a format my transcription API accepts. Say which and why.
- Visible recording indicator and a level meter, so people know it's
  working.
- Cap the length, and warn before hitting it.
- Handle: no microphone, permission denied, the tab losing focus
  mid-recording, and the file being too large to upload.

Show the user their audio before it's sent anywhere.

2. Transcribe server-side

The API key stays on your server, as always. The browser uploads the audio to your endpoint; your endpoint calls the transcription service.

For anything over a few seconds, do this in the background and notify when done — transcription takes real time and a hanging request is a bad experience.

3. Then do the useful part

Raw transcription of speech is messy: filler words, false starts, no punctuation in the right places. On its own it's often harder to read than listening again.

The value is the step after:

voice-to-actions.txt
Here's a transcript of a voice note: <transcript>

Extract:
- Any tasks or commitments, with who and when if stated.
- Any decisions made.
- Anything explicitly flagged as important.
- A one-line summary of what this note was about.

Rules: only what's actually in the transcript. Speech is messy —
ignore false starts and filler. If something is ambiguous, keep the
ambiguity rather than resolving it. Return JSON.

4. Keep the audio

Always. Transcription gets things wrong, especially names and technical terms, and the ability to listen back is what makes the feature trustworthy.

Test with a recording of yourself talking normally, not carefully. Real voice notes have pauses, corrections and background noise, and that's the actual input.