ON YOUR DEVICE · FREE

Live Transcription

It listens, it types, and no sound leaves the room

How this compares with Otter.ai ProHow this compares with Descript

  1. 1

    Choose what it listens to, and in what language

    The microphone, or another tab. A microphone hears the room, which with headphones on is your half of a conversation; a tab is how you catch everyone else, and the browser asks which one. Then say which language will be spoken, because it does not guess.

  2. 2

    Talk

    Words appear as you speak and the last line keeps rewriting itself until the sentence is done, which is what live captioning looks like. Once a stretch of audio is finished with, its text stops moving and joins the transcript above.

  3. 3

    Take the transcript

    Copy it or download it as plain text. Nothing was recorded and nothing was stored, so closing this page is the end of it.

FAQ

Is my voice uploaded?
No, and this is the only transcription on this site that can say so. The other two send the sound to a server and say that plainly. This one downloads a speech model the first time you press start, about 196 MB, and then everything happens on your own machine. No request carries audio anywhere, which you can check by opening the network panel and watching while you talk.
Is it recording me?
No. What exists is the audio since the last few seconds, held in memory so the model has something to read, and the text on the page. Neither is written to disk and there is nothing to delete afterwards.
Can it hear the other people in a meeting?
Only if their voices reach your microphone or the meeting is open in a browser tab. Choose the tab option and the browser asks which one to share, with its sound. It cannot reach an application outside the browser, so a call in the Zoom or Teams desktop app is out of reach. Open the same meeting in a tab and it works.
How fast is it, really?
Faster than talking, which is the only speed that matters here. Measured on this build: eleven seconds of speech comes back in under a second on a machine with a graphics card, and in five seconds without one. Both are quicker than the speech itself, so it does not fall behind.
Why does the last line keep changing?
Because it is a guess at an unfinished sentence. The model reads a stretch of audio rather than one word at a time, so as more of the stretch arrives it revises what it thought. Everything above the last line is fixed: that audio has been used and set aside.
Why do I have to pick the language?
Because it does not guess, and the way it fails is not the way you would expect. Given no language it assumes English, and English against Chinese speech does not come back as nonsense, it comes back as a translation: ask it in Chinese and read English. That is a different tool from the one you pressed start on, so it asks instead. The list opens on the language this page is written in.
Which languages can it do?
The model knows around ninety nine and the list offers the sixteen most asked for. If yours is missing it is a line of code rather than a limitation, so say so.

More tools like this one