Video to transcript, free
Drop a video or audio file and get its whole transcript back, transcribed with WhisperX and timed word by word. No account, no email, no preview: the complete text, plus an SRT you can download. Capped at 20 minutes a file and three files a day so it stays free for people.
Given a video or audio file, produce its full transcript with word timings, in under one run.
Up to 20 minutes of audio or video per file (max 400 MB), 3 files a day. You get the whole transcript — not a preview, not the first minute.
Your file is uploaded to transcribe it and deleted as soon as the transcript is returned; the transcript itself is held for 30 minutes so your browser can collect it, then dropped. Nothing is stored, nothing trains a model, and no account or email is involved.
What a free video to transcript tool gives you, exactly
Plain text you can copy, and word-level timing underneath it — the structure that captioning, subtitling, quoting and search all need. This is the first stage of Never's own pipeline, exposed rather than described: every caption we render starts from this same transcription pass, corrected per customer for names and jargon before it becomes a locked caption style.
Word timings are the part that is easy to skip and expensive to be without. A transcript with sentence timings is enough to read; only word timings are enough to cut on, to caption on the beat, or to find the exact moment someone said the sentence you want to clip.
How to transcribe a video by hand
- Pull the audio out. Export a mono WAV or MP3. Video adds nothing to a transcription pass and makes every step slower.
- Type it, or dictate over it. Real-time typing runs about 4× the media length for a careful transcript. Dictation with a foot pedal gets that to roughly 2×.
- Timestamp as you go. Mark a timestamp every paragraph. Doing it afterwards means scrubbing the whole file a second time.
- Proof against the audio. Names, numbers and jargon are where transcripts fail, and they are also the words most likely to be quoted.
Where the manual way breaks
Transcription is the single most automatable step in video post-production, which is why it is the one worth giving away. Watching a 12-minute recording back and finding the good parts is 25 minutes in the Editing Tax dataset, and that is with a transcript in front of you; without one it is a scrubbing exercise measured in hours per week.
What a transcript still does not give you is the cut. Knowing what was said is not knowing which take to keep, and it is not knowing where the join goes — Whisper's word starts run 50 to 100 milliseconds late against the real acoustic attack, so a cut placed on the transcript's own timing lands inside the consonant and clips it. That gap is measurable, and the cut-point checker measures it on your own audio. Moving the boundary off the timestamp and onto the measured onset is what Never's caption and splice pass does on every join.
Read the transcript
Video to transcript, free Drop a video or audio file and get its whole transcript back, transcribed with WhisperX and timed word by word. No account, no email, no preview: the complete text, plus an SRT you can download. Capped at 20 minutes a file and three files a day so it stays free for people.
Frequently asked questions
Why is this free — what is the catch?
The marginal cost is about a cent per transcript, and a working transcriber demonstrates transcription quality better than a sentence claiming it does. There is no email wall and no watermark. The daily cap exists so the tool stays a tool for people rather than a free API for scripts.
Do I get the whole transcript or a preview?
The whole transcript. Every word, with its timings, plus a downloadable SRT. What is capped is scale — 20 minutes a file and three files a day — and both numbers are printed above the file picker before you choose anything.
How accurate is WhisperX?
Strong on clear speech, and its word-level timestamps are what make on-beat captions possible at all. Two honest caveats: word starts run 50 to 100 milliseconds late against the true acoustic attack, and an immediately repeated word is sometimes collapsed into one entry. Both are invisible in a reading transcript.
What happens to my file?
It is uploaded to transcribe it and deleted as soon as the transcript is returned. The transcript is held for 30 minutes so your browser can collect it, then dropped. Nothing is stored, nothing trains a model, and no account or email is involved.
What languages does it handle?
The major languages well — German, Spanish, French and Portuguese are first-class alongside English, and the language is detected rather than assumed. Accuracy tracks audio quality more than accent: a decent microphone matters more than where the speaker is from.