Upload a WAV, the format recorders and editing software export, and get a transcript with timestamps and speakers. Click any line to hear it, then export.
MP3, WAV, M4A or OGG · up to 30 minutes
Upload a recording you have the rights to transcribe. The transcript is yours; the audio is not kept after processing.
Tap any line to hear it. The line under the playhead lights up as the audio plays. This is what your transcript will look like.
01—How it works
No software to install. A WAV goes in as it is, and if yours is over 50 MB the next section shows how to shrink it to an MP3 first without changing a word of the transcript. The result opens on its own page with the player at the top.
Drop a file of up to thirty minutes and 50 MB. The length is read before anything uploads, and the page tells you whether today's conversion is free or what it costs in credits.
Press Transcribe. The engine listens once, writes every word down with the second it was spoken, and separates the speakers. A ten-minute file usually comes back in under a minute.
The transcript opens beside a player. Click any line to hear it, search for a phrase, rename the speakers, and download plain text, text with timestamps, SRT or VTT.
About a tenth of the recording's length. A three-minute take is back in seconds and a thirty-minute interview in a few minutes. The page shows a timer while it works.
02—The format
A WAV is uncompressed PCM audio, the default export of field recorders, Audacity, every DAW and most phone systems. It stores every sample the microphone heard, one after another, which makes it the highest-quality file you can bring to a WAV to text converter and also by far the largest.
| Recording settings | Size per minute | Fits in 50 MB |
|---|---|---|
| 48 kHz · 24-bit · stereo (studio) | ≈ 17 MB | ≈ 3 minutes |
| 44.1 kHz · 16-bit · stereo (CD quality) | ≈ 10 MB | ≈ 5 minutes |
| 16 kHz · 16-bit · mono (speech) | ≈ 1.9 MB | ≈ 26 minutes |
| The same audio as MP3 at 128 kbps | ≈ 1 MB | ≈ 50 minutes |
Before it listens, the engine brings every file down to a speech sample rate, so a 48 kHz, 24-bit studio WAV carries no more usable words than a 16 kHz mono one. What changes the transcript is the room, the microphone distance and people talking over each other. The extra data in a big WAV is for mixing, not for transcription.
If a WAV is over 50 MB or thirty minutes, export it from Audacity as MP3 at 128 kbps, or set the recorder to MP3 mode for the next interview. The transcript comes back identical and the upload is about ten times faster. A WAV that already fits can be uploaded as it is.
The name your recorder or Audacity gave the file, often something like ZOOM0012.WAV, becomes the title of the transcript and the name of every export. Rename the file before you upload it, or rename the transcript afterwards with one click.
Mono is all transcription needs. A stereo interview with one voice on each channel is fine too, because the engine tells speakers apart by their voices rather than by channel.
03—Click a line
An hour on a field recorder is a long file to scrub through for one sentence. Here the transcript and the audio are one thing on one page, and every line knows the second it belongs to.
While the recording plays, the line being spoken is highlighted and kept in view. Click a different line and playback jumps there. Verifying a quote takes two seconds.
Type a phrase and the transcript filters to the lines that contain it. Click one and you are hearing that exact moment.
Each line carries the second it starts, measured from the file rather than estimated from reading speed, so a timestamp you paste into an edit log lands on the right moment.
Two voices come back as Speaker 1 and Speaker 2. Rename them once and every line and every export updates.
04—What you get
One conversion produces four files, all built from the same timed segments, so plain text, captions and the timestamped version never disagree with each other.
Paragraphs split where the speaker paused or changed, ready to paste into notes, a document or an email.
The same lines with their start times, or SRT and VTT files that drop straight onto a video track.
Interviews and calls come back with every turn attributed, and you decide what the speakers are called.
Each transcript stays in your history for thirty days with its audio, so you can reopen the page, play it again and export later. The WAV you uploaded is not kept anywhere else, and deleting the transcript removes both.
05—Where WAVs come from
WAV is the format professional gear and editing software write by default, so it tends to arrive from four places. Each one comes with a setting worth knowing before the next recording.
Zoom, Tascam and Sony recorders write WAV out of the box, which is right for music and heavy for interviews. Most of them also offer an MP3 mode; switch to it for speech and skip the conversion step entirely.
After trimming and cleaning a take, the export dialog defaults to WAV. Choose MP3 instead when the file is headed for transcription; the transcript is the same and the file is a tenth of the size.
Support lines and phone interviews usually export narrow 8 kHz WAVs, which are small and transcribe fine. Because they are often thin and noisy, a pass through the Voice Isolator first gives a noticeably better transcript.
Podcast studios and voice-over sessions leave big, clean WAVs. The transcript of a take is a script, so it can go straight to Text to Speech or a cloned voice for a pickup line without booking the studio again.
It will not transcribe words that were never audible, it does not translate the audio itself, and it is not live captioning. It works on WAV files you already have, up to 50 MB and thirty minutes each; shrink a bigger one to MP3 or split it at a natural pause.
06—Questions
Uncompressed PCM audio, the default export of recorders, Audacity, DAWs and phone systems. It keeps every sample the microphone captured, which is why it sounds best and takes the most space.
Yes. Upload the WAV, press Transcribe, and the transcript opens with timestamps and speaker labels. Click any line to hear it, then download it as text or subtitles.
Export it as MP3 at 128 kbps from Audacity or any editor, which shrinks a CD-quality WAV by about ten times, or split it at a natural pause into two files. The transcript is identical either way.
Only if the file is over the size or length limit. A WAV that fits can be uploaded as it is. When you do convert, nothing is lost for transcription, because the engine works at a speech sample rate anyway.
No. A 48 kHz, 24-bit file and a 16 kHz, 16-bit file of the same take produce the same transcript. Microphone distance, background noise and people talking over each other are what change it.
Yes. Every turn is assigned to a speaker you can rename, and it works on mono and stereo files alike. On a recording with one voice the labels are simply not shown.
Download the plain-text export and paste it into Word or Google Docs, or use the timestamped export if you want the times in the document. Paragraphs are already split where the speaker paused or changed.
The engine is built for spoken voices, so a mixed song comes back patchy. Split it with the Vocal Remover first and transcribe the isolated vocal track; that works far better.
07—More tools
WAV is the studio format; the rest of AnyVoice handles the other formats and everything you might do with the transcript.