Convert Audio to Word

Turn any recording into an editable Word document — transcribed for you.
99 languages, timestamps, one-click DOCX export. Free to start.

Drop Your Videoor Audio
Drop your file here to upload
First 30 minutes free — unlock the full recording with Pro →
or

The link must be publicly accessible.

Audio to a Word document in three steps

Get a recording into Word without retyping it — transcribe, then export a ready-to-edit DOCX.

  1. Step 1

    Upload your audio

    Drop in an MP3, WAV, M4A or other recording, or paste a link to one.

  2. Step 2

    Transcribe and tidy it

    The speech becomes punctuated, timestamped text with the language detected automatically — fix any stray word next to the player, and turn timestamps off if the document shouldn't carry them.

  3. Step 3

    Export to Word (DOCX)

    Download a .docx with one paragraph per spoken segment — open and edit it in Microsoft Word, Google Docs, or Pages.

When you need it in Word

Plenty of work has to land in a Word document — start from the recording and skip the typing.

Interview write-ups

Turn a recorded interview into a Word doc ready for editing and quoting.

Meeting minutes

Convert a recorded meeting into a DOCX of minutes you can format and share.

Drafts & dictation

Dictate a draft, then export a Word document to polish.

Research notes

Get field recordings into an editable doc for analysis.

No retyping

Skip transcribing by hand — the Word doc is generated for you.

99 languages

Transcribe in almost any language and export to Word.

Why export audio straight to Word

If your deliverable is a document, plain text isn't the finish line — you need it in Word. SlayScribe transcribes your audio and exports a clean, editable DOCX, so an interview, meeting, or dictated draft lands in Microsoft Word ready to format and share. No copy-paste, no reformatting, no manual typing.

What is actually inside the .docx

The document is assembled from the transcript in a fixed shape. Here is that shape, and where it stops.

The document is built in your browser

Nothing is rendered on our side. When you press Word, the document writer is fetched on demand — it stays out of the main bundle because most visits never export anything — and the .docx is assembled locally out of the transcript already on your screen, then handed straight to your download folder. The file takes the name of the recording with its extension swapped, so interview-03.m4a comes back as interview-03.docx. Two things follow from that: the export works on a machine with no Office installed, and what you see on screen is exactly what lands in the file, because the file is made out of it.

One paragraph per spoken segment

The transcript is not poured in as one block. Speech comes back split into phrase-level segments, and each segment becomes its own Word paragraph with a 6 pt gap after it — that is why the document reads as speech rather than as a wall of text. With timestamps on, a paragraph opens with a bold [mm:ss] run in front of the words: a bold run, not a field, a table cell, or a frame, so removing the marks is ordinary editing. The minutes run straight through instead of wrapping at sixty, so minute 80 of a recording reads [80:14] and not [1:20:14] — the mark is a position in the audio, not a clock time.

What the AI summary does to the document

When a summary exists and is not hidden, the file opens with an “AI Summary” heading, then the summary, then a “Transcript” heading and the segments. Both are real Word Heading 1 styles, so the document has a navigation pane and a table of contents can be built from it. The summary body is flattened on the way in, though, because it is written as Markdown and Word has no Markdown: bold marks are dropped to plain text, a heading line inside the summary becomes a bold paragraph rather than a heading style, and a bullet becomes a literal “•” character starting the line — text that looks like a list without being a Word list, so list indent and renumbering have nothing to act on.

What the file deliberately does not carry

No font, margin, or page style is set anywhere in the export. The document arrives in Word's own default body style and gives way to your template the moment you apply one — which is the reason not to style it in the first place. Also absent by design: speaker labels (the transcript is what was said, in order, not who said it), page numbers, headers and footers, a link back to the audio, and any table layout. And if a transcript has no segments at all, the whole text lands as a single paragraph — paragraphs come from the segment structure, not from the length of the text.

Translations, hidden summaries, and the free display limit

Export while the translated view is open and the translated text is what goes into the .docx; the original stays in the app rather than being overwritten, so both versions remain one click apart. Hide the summary before exporting and it is left out of the document entirely. On a free account a recording longer than the display limit is still transcribed all the way through, but the document carries the first 30 minutes, and your allowance is charged for what you were shown — never more than 30 minutes. That limit is stated before the upload starts, not discovered at the export button.

What the Word export contains

File you get
.docx — Office Open XML, the format Word 2007 and newer save natively
Where it is built
In your browser, from the transcript on screen — the document itself is never assembled on a server
Opens in
Microsoft Word (desktop and Word for the web), Google Docs, Apple Pages, LibreOffice Writer
Document structure
“AI Summary” section when one exists, then “Transcript”, then one paragraph per spoken segment
Headings
“AI Summary” and “Transcript” are real Word Heading 1 styles, so they appear in the navigation pane and feed a table of contents
Paragraph spacing
6 pt after each paragraph; no font, margin, or page style is set, so your own Word template applies cleanly
Timestamps
Bold [mm:ss] at the start of each paragraph, minutes counted straight through (minute 80 reads 80:14) — switch them off before exporting for a clean prose document
File name
The recording's name with its extension replaced: interview-03.m4a becomes interview-03.docx
Summary formatting
Markdown is flattened for Word — bullets arrive as “•” characters rather than as Word list items
Not in the file
Speaker labels, page numbers, headers and footers, and any link back to the audio
Audio you can start from
MP3, WAV, M4A, AAC, OGG, FLAC and other common files, video files, or a link to a public recording
Same transcript also exports to
TXT, Markdown, PDF, and SRT / VTT subtitles, all built from the same segments

Frequently asked questions

Can I export the transcript to Word?

Yes. Transcribe the audio, then download a .docx that opens in Microsoft Word, Google Docs, or Pages.

What is inside the Word file?

The Word file holds one paragraph per spoken segment, in order, with a 6 pt gap after each and a bold timestamp in front when timestamps are on. If a summary was generated, an “AI Summary” section sits above a “Transcript” heading; if not, the document starts at the first spoken line.

Does the Word file keep the timestamps?

The Word file keeps them: by default each paragraph opens with a bold [mm:ss] mark. Turn timestamps off before exporting and the DOCX comes out as plain prose.

Why does a timestamp read 80:14 instead of 1:20:14?

Because minutes are counted straight through the recording rather than wrapped into hours, so minute 80 stays 80:14. The mark answers “how far into the audio is this line”, which is what you need when checking a quote against the recording.

Does the document use real Word heading styles?

The two section titles do — “AI Summary” and “Transcript” are Heading 1, so Word's navigation pane lists them and a table of contents can be generated. The transcript paragraphs themselves carry no style, which is what lets a corporate template take them over untouched.

Are the summary's bullet points a real Word list?

The summary's bullet points reach the document as characters, not as a Word list. The summary is written as Markdown and flattened on its way into the document, so each bullet arrives as a “•” starting the line. Word's list tools have nothing to renumber or re-indent there, though turning the block into a real list is one click in Word.

Does the Word file say who is speaking?

Speaker labels are not part of the export today; the document is what was said, in order, with times. In an interview the turn-taking is usually visible from the segment breaks, and a name can be typed in once in Word.

Will the file fight my company template?

The export sets no font, no margins, and no page styles, so the paragraphs adopt whatever style you apply. Paste them into your .dotx template, or apply the template to the downloaded file, and the transcript takes its formatting from your side rather than ours.

Can I export a translated transcript to Word?

Yes — export while the translated view is open and the translation is what goes into the .docx. The original transcript is not overwritten, so you can switch back and export the source language as a second document.

Is the Word export free?

Yes — the DOCX download is included in the free monthly allowance of 60 minutes, with no card required. On a long recording a free account gets the first 30 minutes in the document and is charged at most 30 minutes for it; Pro lifts that and raises the allowance to 30 hours a month.

More ways to transcribe

Audio to Microsoft WordAudio to TextMP3 to Text

Long recordings and big files?

Go Pro to transcribe hours of audio and video, with more monthly minutes and unlimited uploads — accurate text in 99 languages.