Let AI do the listening — accurate, punctuated transcripts from any recording.
Powered by OpenAI Whisper, 99 languages, timestamps. Free to start.
The link must be publicly accessible.
Modern speech AI turns recordings into clean text far faster and cheaper than a human typist — here's how.
Drop in any recording or paste a link from YouTube, TikTok, Instagram, Vimeo and others. There is no model to configure and no language to declare — both are decided from the audio itself.
The recording is decoded, split if it is long, and sent to OpenAI Whisper, which returns punctuated, capitalised text broken into phrase-level segments with a start and end time on each one.
The transcript sits beside the player, so any segment can be replayed and corrected. Export to TXT, Markdown, DOCX, PDF, SRT or VTT — the same writer used by the standalone subtitle tools.
Speech AI handles volume, speed, and languages that manual transcription can't keep up with.
A long recording is cut into pieces and transcribed several pieces at a time, so an hour of audio does not take an hour of waiting.
One model covers almost any language, and the language is read from the audio rather than picked from a menu you can get wrong.
Sentences, commas and capitals come back with the text, which is the difference between a transcript you can read and a wall of lowercase words.
Each phrase carries the time it was spoken, which is what makes click-to-play, subtitle export and search inside the recording possible.
No tiring, no drift — file fifty comes out to the same standard as file one, at the same cost per minute.
Interviews, calls, lectures, voice memos, field recordings — if there is speech in it, the model writes it down.
Manual transcription is slow and expensive — roughly four hours of work per hour of audio. AI transcription with OpenAI Whisper does the same job in minutes, adds punctuation and timestamps, and handles 99 languages without breaking stride. You stay in control: the transcript sits next to the audio so you can verify and fix anything the model missed.
The engine, the choices behind it, and the places it still gets things wrong.
OpenAI has faster and more accurate speech models — gpt-4o-transcribe and its mini variant — and we do not use them, on purpose. They return plain text and nothing else. whisper-1 is the only one that returns per-segment timestamps, and those timestamps are what the player, the SRT and VTT export and the translation feature are all built on. Swapping in the more accurate model would raise the word accuracy and silently break three features, so the trade goes the other way.
Audio past a certain length is not sent in one piece. The file is scanned for silences and cut at the last quiet point before each boundary, so the split lands between sentences instead of through a word, and the pieces are transcribed several at a time rather than one after another. If a job is interrupted, the cut points are already saved, so it resumes from the chunk it stopped on instead of starting the whole file again.
The language is detected on the first chunk and then reused for the rest of the file. That is deliberate: left to detect itself on every chunk, the model can flip language mid-recording on a quiet or accented passage and hand back a transcript that changes language halfway. The cost of the rule is that a genuinely bilingual recording is transcribed as its opening language.
Accuracy is a property of your audio, not of the marketing page, so we do not publish a percentage. What moves it: overlapping speakers, background noise and music, phone-quality compression, heavy accents, and proper nouns the model has no way to know — names, drug names, internal product names. The practical answer is the editor: the text stays beside the audio, every segment is clickable, and fixing a handful of names in a transcript takes a fraction of the time typing it would.
A file longer than the free display limit is still transcribed all the way through — nothing is truncated in the model call — but only the first 30 minutes are shown and exported until the account is on Pro. Your allowance is charged for what you were shown, capped at 30 minutes, not for the full length of the file. The limit is stated before the upload starts rather than discovered afterwards, and Pro removes it entirely.
Transcription runs on OpenAI whisper-1, picked over the newer and more accurate gpt-4o-transcribe models because it is the only one that returns per-segment timestamps, which the player, the subtitle export and translation all depend on.
Accuracy depends on the recording rather than on the plan: clean speech comes back close to verbatim, while overlapping voices, background music, phone-quality audio and unfamiliar proper nouns are where errors appear. Every segment is clickable next to the audio, so corrections are a matter of minutes rather than retyping.
Most files finish in minutes rather than in the length of the recording, because a 60-minute file is cut into chunks and several of them are transcribed at the same time instead of end to end. Long uploads keep running in the background even if you close the tab.
Speech in 99 languages, recognised automatically from the audio with no setting to pick. The language is detected on the first chunk and then held for the rest of the file, so a recording that switches language midway is transcribed as the language it opened in.
The transcript exports to TXT, Markdown, DOCX, PDF, SRT and VTT, all generated from the same timed segments, so the subtitle files line up with the audio to the millisecond.
Speaker labels are not part of the output today — the transcript is what was said, in order, with times. For an interview, the usual workaround is that turn-taking is visible in the segment breaks.
Transcription writes down what was said and nothing more; an AI summary is a separate action you trigger yourself, so the transcript is never quietly rewritten or shortened on your behalf.
The whole file is transcribed, but a non-Pro account sees and exports the first 30 minutes, and is charged at most 30 minutes of its allowance for it. The cap is shown before the upload, not after the wait.
A single file can run to 60 minutes on Free and 10 hours on Pro, within 5 GB of upload either way.
AI transcription is free within a monthly allowance of 60 minutes with no card required; Pro raises it to 30 hours a month, lifts the display cap and removes the daily upload limit.
Go Pro to transcribe hours of audio and video, with more monthly minutes and unlimited uploads — accurate text in 99 languages.
See Pro plans