Turn any recording into an editable Word document — transcribed for you.
99 languages, timestamps, one-click DOCX export. Free to start.
The link must be publicly accessible.
Get a recording into Word without retyping it — transcribe, then export a ready-to-edit DOCX.
Drop in an MP3, WAV, M4A or other recording, or paste a link to one.
The speech becomes punctuated, timestamped text with the language detected automatically — fix any stray word next to the player, and turn timestamps off if the document shouldn't carry them.
Download a .docx with one paragraph per spoken segment — open and edit it in Microsoft Word, Google Docs, or Pages.
Plenty of work has to land in a Word document — start from the recording and skip the typing.
Turn a recorded interview into a Word doc ready for editing and quoting.
Convert a recorded meeting into a DOCX of minutes you can format and share.
Dictate a draft, then export a Word document to polish.
Get field recordings into an editable doc for analysis.
Skip transcribing by hand — the Word doc is generated for you.
Transcribe in almost any language and export to Word.
If your deliverable is a document, plain text isn't the finish line — you need it in Word. SlayScribe transcribes your audio and exports a clean, editable DOCX, so an interview, meeting, or dictated draft lands in Microsoft Word ready to format and share. No copy-paste, no reformatting, no manual typing.
The document is assembled from the transcript in a fixed shape. Here is that shape, and where it stops.
Nothing is rendered on our side. When you press Word, the document writer is fetched on demand — it stays out of the main bundle because most visits never export anything — and the .docx is assembled locally out of the transcript already on your screen, then handed straight to your download folder. The file takes the name of the recording with its extension swapped, so interview-03.m4a comes back as interview-03.docx. Two things follow from that: the export works on a machine with no Office installed, and what you see on screen is exactly what lands in the file, because the file is made out of it.
The transcript is not poured in as one block. Speech comes back split into phrase-level segments, and each segment becomes its own Word paragraph with a 6 pt gap after it — that is why the document reads as speech rather than as a wall of text. With timestamps on, a paragraph opens with a bold [mm:ss] run in front of the words: a bold run, not a field, a table cell, or a frame, so removing the marks is ordinary editing. The minutes run straight through instead of wrapping at sixty, so minute 80 of a recording reads [80:14] and not [1:20:14] — the mark is a position in the audio, not a clock time.
When a summary exists and is not hidden, the file opens with an “AI Summary” heading, then the summary, then a “Transcript” heading and the segments. Both are real Word Heading 1 styles, so the document has a navigation pane and a table of contents can be built from it. The summary body is flattened on the way in, though, because it is written as Markdown and Word has no Markdown: bold marks are dropped to plain text, a heading line inside the summary becomes a bold paragraph rather than a heading style, and a bullet becomes a literal “•” character starting the line — text that looks like a list without being a Word list, so list indent and renumbering have nothing to act on.
No font, margin, or page style is set anywhere in the export. The document arrives in Word's own default body style and gives way to your template the moment you apply one — which is the reason not to style it in the first place. Also absent by design: speaker labels (the transcript is what was said, in order, not who said it), page numbers, headers and footers, a link back to the audio, and any table layout. And if a transcript has no segments at all, the whole text lands as a single paragraph — paragraphs come from the segment structure, not from the length of the text.
Export while the translated view is open and the translated text is what goes into the .docx; the original stays in the app rather than being overwritten, so both versions remain one click apart. Hide the summary before exporting and it is left out of the document entirely. On a free account a recording longer than the display limit is still transcribed all the way through, but the document carries the first 30 minutes, and your allowance is charged for what you were shown — never more than 30 minutes. That limit is stated before the upload starts, not discovered at the export button.
Yes. Transcribe the audio, then download a .docx that opens in Microsoft Word, Google Docs, or Pages.
The Word file holds one paragraph per spoken segment, in order, with a 6 pt gap after each and a bold timestamp in front when timestamps are on. If a summary was generated, an “AI Summary” section sits above a “Transcript” heading; if not, the document starts at the first spoken line.
The Word file keeps them: by default each paragraph opens with a bold [mm:ss] mark. Turn timestamps off before exporting and the DOCX comes out as plain prose.
Because minutes are counted straight through the recording rather than wrapped into hours, so minute 80 stays 80:14. The mark answers “how far into the audio is this line”, which is what you need when checking a quote against the recording.
The two section titles do — “AI Summary” and “Transcript” are Heading 1, so Word's navigation pane lists them and a table of contents can be generated. The transcript paragraphs themselves carry no style, which is what lets a corporate template take them over untouched.
The summary's bullet points reach the document as characters, not as a Word list. The summary is written as Markdown and flattened on its way into the document, so each bullet arrives as a “•” starting the line. Word's list tools have nothing to renumber or re-indent there, though turning the block into a real list is one click in Word.
Speaker labels are not part of the export today; the document is what was said, in order, with times. In an interview the turn-taking is usually visible from the segment breaks, and a name can be typed in once in Word.
The export sets no font, no margins, and no page styles, so the paragraphs adopt whatever style you apply. Paste them into your .dotx template, or apply the template to the downloaded file, and the transcript takes its formatting from your side rather than ours.
Yes — export while the translated view is open and the translation is what goes into the .docx. The original transcript is not overwritten, so you can switch back and export the source language as a second document.
Yes — the DOCX download is included in the free monthly allowance of 60 minutes, with no card required. On a long recording a free account gets the first 30 minutes in the document and is charged at most 30 minutes for it; Pro lifts that and raises the allowance to 30 hours a month.
Go Pro to transcribe hours of audio and video, with more monthly minutes and unlimited uploads — accurate text in 99 languages.
See Pro plans