Upload an MP3 and get a transcript with timestamps and speakers. Click any line to hear exactly where it was said, then export.
MP3, WAV, M4A or OGG · up to 30 minutes
Upload a recording you have the rights to transcribe. The transcript is yours; the audio is not kept after processing.
Tap any line to hear it. The line under the playhead lights up as the audio plays. This is what your transcript will look like.
01—How it works
No software to install and nothing to convert first. The MP3 goes in as it is, and the transcript opens on its own page with the player at the top and the export buttons where you expect them.
Drop a file of up to thirty minutes. The length is read before anything uploads, and the page tells you whether today's conversion is free or what it costs in credits.
Press Transcribe. The engine listens once, writes every word down with the second it was spoken, and separates the speakers. A ten-minute MP3 usually comes back in under a minute.
The transcript opens beside a player. Click any line to hear it, search for a phrase, rename the speakers, and download plain text, text with timestamps, SRT or VTT.
About a tenth of the recording's length. A three-minute voice note is back in seconds and a thirty-minute interview in a few minutes. The page shows a timer while it works.
02—The format
An MP3 is a compressed audio file. It drops the frequencies a human ear barely notices and keeps the rest, so a recording ends up around a tenth of its original size with the voice intact. That is why MP3 is the most common file people bring to an MP3 to text converter.
A spoken-word podcast at 64 kbps and a music file at 320 kbps sound very different, but to a transcription engine they carry almost the same words. Compression is not what loses words; noise in the room at recording time is. Only well below 32 kbps do consonants start to blur, and files that small are rare today.
Most speech MP3s are mono, which is exactly right for transcription. A stereo interview with one person on each channel is fine too. The engine tells speakers apart by their voices, not by which channel they sit on, so nothing needs to be re-mixed first.
The name of the MP3 you upload becomes the title of the transcript, and every export you download is named after it. Call the file something you will recognise before you upload it, or rename the transcript afterwards with one click.
No. An MP3 does not need to become a WAV before transcription, and a WAV or M4A does not need to become an MP3. Upload whatever you have. This page simply expects MP3 and accepts the rest.
03—Click a line
An MP3 podcast is often an hour long, and checking one quote should not mean dragging a progress bar back and forth. Here the transcript and the audio are one thing on one page, and every line knows the second it belongs to.
While the MP3 plays, the line being spoken is highlighted and kept in view. Click a different line and playback jumps there. Verifying a quote takes two seconds.
Type a phrase and the transcript filters to the lines that contain it. Click one and you are hearing that exact moment.
Each line carries the second it starts, measured from the file rather than estimated from reading speed, so a timestamp you paste into show notes lands on the right moment.
Two voices come back as Speaker 1 and Speaker 2. Rename them once and every line and every export updates.
04—What you get
One conversion produces four files, all built from the same timed segments, so plain text, captions and the timestamped version never disagree with each other.
Paragraphs split where the speaker paused or changed, ready to paste into notes, a document or an email.
The same lines with their start times, or SRT and VTT files that drop straight onto a video track.
Interviews and calls come back with every turn attributed, and you decide what the speakers are called.
Each transcript stays in your history for thirty days with its audio, so you can reopen the page, play it again and export later. The MP3 you uploaded is not kept anywhere else, and deleting the transcript removes both.
05—Where MP3s come from
MP3 is the format things get downloaded and exported in, which is why so much speech arrives in it. These are the six places an MP3 usually comes from, and what turning it into text gets you in each case.
Pull quotes for the show notes, find the minute a topic came up, and publish a transcript for the people who would rather read than listen.
Interviews and lectures straight off the device. One microphone close to the speaker beats a good one in the middle of the table.
A transcript with speakers is a record of who committed to what. Search for a name or a deadline instead of replaying an hour.
A recorded class becomes notes with timestamps, so a student can go back to the exact minute a concept was explained.
An interview saved as an MP3 ten years ago transcribes as well as one recorded today. Compression was never the problem; the room it was recorded in was.
The engine is built for spoken voices, not singing. For lyrics, split the song with the Vocal Remover first and transcribe the vocal track on its own; the result is noticeably cleaner.
It will not transcribe words that were never audible, it does not translate the audio itself, and it is not live captioning. It works on MP3 files you already have, up to thirty minutes each; split a longer one at a natural pause and convert it in two parts.
06—Questions
Yes. Upload the MP3, press Transcribe, and the transcript opens with timestamps and speaker labels. Click any line to hear it, then download it as text or subtitles.
Every signed-in account gets three free conversions a day of up to ten minutes each, with every export format. Paid plans remove the daily cap, allow files up to thirty minutes, and bill by the minute in credits. The rate is on the pricing page.
On a clear recording of one or two people it is close to what a careful listener would write down. MP3 compression barely matters; background noise, distant microphones and people talking over each other are what lower it. Cleaning the audio with the Voice Isolator first helps more than anything else.
Not in practice. Speech at 64 kbps transcribes as well as speech at 320 kbps. Only files well below 32 kbps start to lose consonants, and the distance to the microphone matters far more than the bitrate.
Yes. Each line carries its start time taken from the audio, and every turn is assigned to a speaker you can rename. On a recording with only one voice the labels are simply not shown.
The engine is designed for spoken voices, so a mixed song will come back patchy. Split the song with the Vocal Remover first and transcribe the isolated vocal track; that works far better.
Up to 50 MB and thirty minutes per file. A typical hour-long podcast MP3 fits within the size limit, so split it at a natural pause and convert it as two files.
Plain text, text with timestamps, SRT and VTT. Every export is named after the MP3 you uploaded, so a file called interview.mp3 comes back as interview.srt.
07—More tools
MP3 is one format among several, and the transcript is a script once it exists. The rest of AnyVoice handles both ends.