Voice Notes
Speaking is faster than typing, works while walking or driving, and requires no interface. Transcription is now good enough that the recording arrives as searchable text without effort.
That is the lowest-friction capture ever available, and low friction is not straightforwardly good. The friction was doing filtering work, and removing it removes the filter.
A related workplace concept is online timesheets, which offers a useful contrast with the personal knowledge-work problem discussed here.
What voice notes are genuinely good at
Catching things that would otherwise be lost. Walking, driving, in the middle of something. The alternative is not a typed note; it is nothing.
Thinking out loud. Some ideas only form in speech, and the rambling is the process. This is a real cognitive difference and not a lesser form of writing.
Emotional and situational material. Tone carries information that typed notes lose, and for anything journal-shaped that can matter.
For an independent external reference related to this topic, see Otter.
And volume of raw material when you are working something out — twenty minutes of talking produces a lot to sift.
What goes wrong
Speech is four to six times longer than the equivalent note. A two-minute recording transcribes to several hundred words of circling, restating and false starts. Multiply by a year of capture and you have an archive of transcripts nobody will read, including you.
No selection happens. Writing forces condensing, and condensing is where the established benefit lives. Speaking captures the whole stream, including the parts you would have dropped.
Titles do not happen either. Voice notes arrive as timestamps or as first-line fragments, which means nothing in the list gives you anything to attempt recall against and the collection is unscannable.
And transcription errors are invisible. A misheard name or number sits in the text looking exactly like everything else, and you will not catch it later because you will not remember what you said.
The one habit that fixes it
Listen back once, and write three lines.
Not a transcript, not a cleanup — three lines: the point, why it mattered, what to do about it. Ninety seconds after a two-minute recording.
This does the whole job. It performs the condensing that speaking skipped, it produces a scannable title, and it decides whether the recording was worth anything — which is often no, and knowing that is worth ninety seconds.
Then delete the audio, or keep it attached to the three lines rather than instead of them.
When to do it
Same day. The condensing requires remembering the context, and the context is the part that decays fastest.
In a batch. Five recordings in ten minutes is easier than five separate ninety-second sessions, and batching is how it actually gets done.
And if you will not do it, record less. A pile of unprocessed transcripts is not an archive with a backlog; it is material that never became notes. Being honest about this is better than accumulating for a processing session that never comes.
The transcription question
Automatic transcription is now good and it is not a substitute for the three lines. A perfect transcript of rambling is a perfect record of rambling.
Where you transcribe matters. Voice notes are frequently the most personal material anyone captures — thinking out loud, said alone. Sending them to a hosted service is a decision worth making deliberately, and it is the same question that applies to any tool holding your notes.
And do not trust names, numbers or quotations from a transcript without checking. Those are precisely what transcription gets wrong and precisely what you will later cite.
The short version
- Voice capture is the lowest-friction there is, and the friction was doing useful filtering
- It is genuinely good for material that would otherwise be lost, for thinking out loud, and for tone
- Speech runs four to six times longer than the equivalent note, and no selection happens
- Voice notes arrive without titles, so the collection is unscannable and offers nothing to recall against
- The fix is ninety seconds: listen back once, write three lines — the point, why it mattered, what to do
- If you will not do that, record less; unprocessed transcripts never became notes