Video to transcript, free
Drop a video or audio file and get its transcript back — transcribed with WhisperX, the same word-timestamp engine that captions every Never video. Free, no account, rate-limited to stay free for humans. Journalists, students, researchers and accessibility teams are exactly who it's for.
Rate-limited to keep it free for humans. Same WhisperX engine that captions every Never video — word-timed, jargon-corrected.
What you get, exactly
Plain text you can copy, with word-level timing underneath — the structure captioning, subtitling and quoting all need. It's the first stage of Never's own pipeline, exposed: every caption we render starts from this same transcription pass, corrected per customer for names and jargon before it becomes a locked caption style.
Read the transcript
Video to transcript, free Drop a video or audio file and get its transcript back — transcribed with WhisperX, the same word-timestamp engine that captions every Never video. Free, no account, rate-limited to stay free for humans. Journalists, students, researchers and accessibility teams are exactly who it's for.
Frequently asked questions
Why is this free — what's the catch?
Marginal cost is about a cent per transcript, and the tool shows our transcription quality better than a claim would. No email wall, no watermark. The rate limit exists only so it stays a tool for people rather than a free API for scripts.
How accurate is WhisperX?
State of the art on clear speech, and its word-level timestamps are what make on-beat captions possible. Two honest caveats: word starts can run tens of milliseconds late against the true acoustic attack, and immediately repeated words sometimes merge — both are correctable downstream, and both are invisible in a reading transcript.
What languages does it handle?
The engine handles the major languages well — German, Spanish, French and Portuguese are first-class alongside English. Accuracy degrades gracefully with audio quality; a decent microphone matters more than an accent.