Skip to content
EN
English 简体中文 soon 日本語 soon

OpenAI Whisper

Open-weight speech-to-text that turns audio into accurate transcripts and subtitle files.

Visit official site

What OpenAI Whisper is

Whisper is OpenAI's speech-recognition model family, notable for being open-weight: you can download the models and run them on your own hardware. It converts spoken audio into text with strong multilingual accuracy, and it ships with tooling that outputs plain transcripts, word-timestamps, and subtitle files. In practice, Whisper is the default engine behind many transcription apps — and for people who care about privacy, it is the one you can run with no data leaving your machine.

This review reflects local runs of the large and medium models on interview audio, meeting recordings and short video clips, including Mandarin and accented English.

Who should use it

Whisper suits anyone who transcribes regularly and either has a bit of technical comfort or uses a wrapper around it. Video creators use it to generate SRT captions in one command. Journalists and researchers use it to turn interview recordings into searchable text. Privacy-conscious users — in law, medicine or finance — use self-hosted Whisper precisely because the audio never touches a third-party server. If you want zero setup and a graphical editor instead, Descript wraps similar transcription behind a polished timeline.

Pricing breakdown

Whisper has an unusual pricing story: the core product is free and open source. Running it yourself costs only compute — free on a friend's GPU or a few cents of cloud compute per hour of audio on a rented instance. OpenAI also sells a hosted API billed per minute, which is the right choice when you value convenience over control. For most individual users the practical question is not price but whether you have the hardware or are happy to use a hosted wrapper.

Hands-on notes

On clean interview audio the medium model produces transcripts that need only light cleanup. On noisy or accented audio, the large model is noticeably more reliable, at the cost of speed.The biggest practical trap is hallucination on near-silence: long pauses and background music can produce invented sentences, so always spot-check transcripts against the source rather than trusting the output blindly.

What we like

The combination of open weights, strong multilingual accuracy and subtitle-file output is unmatched at the price. Being able to transcribe an hour of audio on your own hardware — with no upload and no per-minute fee — makes Whisper the backbone of the privacy-friendly workflow. Command-line usage is simple once set up, and wrappers add a GUI when you want one.

What to watch out for

Setup friction is the main cost: you need Python, the right PyTorch build and ideally a CUDA GPU for long files. CPU-only transcription of a one-hour recording can take a long time. Hallucinations on silence and music are real and should be checked. And while the hosted API removes setup, it also removes the privacy guarantee of local processing — choose based on what your audio actually contains.

When Whisper does not fit

Whisper is the wrong tool when you need an interface, not an engine. If you want to click a button, see speaker labels and edit the transcript inside a polished timeline, a product such as Descript wraps transcription and editing together — at the cost of uploading your audio to its cloud. Whisper also struggles when your audio is mostly music or heavy background noise, where it can invent words to fill silence; for such files, clean the track first or budget for careful manual review.

The hardware question matters too. A short clip runs fine on almost anything, but a one-hour podcast on a CPU-only laptop means a long wait. If that is your situation, renting a GPU for a few minutes or using a hosted transcription API with a strong privacy policy may be more practical than a slow local run — just weigh the privacy trade-off against what your audio contains.

Practical setup tips

Start with the medium model for a quick quality test, then move to large for the files that matter. Keep a stable naming convention for outputs (episode-01.srt) so captions are easy to attach in your editor. For interviews, record in a quiet room and keep the microphone close; Whisper's accuracy drops far more with background noise than with accent variation. And always spot-check numbers, names and technical terms — the model has no way to know how your specific jargon is spelled.

Alternatives worth a look

  • Descript — a full editing studio where you fix the transcript and the video edits itself.
  • HandBrake — compress your finished, captioned video for a fast social upload.

Pros & cons

✓ What we like

  • Open-weight model you can run fully offline
  • Outputs structured transcript plus SRT/VTT/JSON
  • Free to self-host with a modest GPU

! What to watch out for

  • Local setup needs Python and a decent GPU for long files
  • Base model can hallucinate on silence and music-heavy audio
  • No built-in UI unless you add a wrapper
  • Large-model inference is slow on CPU-only machines

Alternatives

Similar tools worth a look, and why.

Descript

Pick Descript if you want a polished editor that also lets you cut video by editing the transcript.

Read our review

HandBrake

Pair Whisper transcripts with HandBrake to compress the video before uploading with captions.

Read our review

FAQ

Is Whisper free?

The model weights are open source, so self-hosting is free apart from your compute. OpenAI also offers a hosted API billed by audio minute, which avoids local setup.

Does it work in languages other than English?

Yes — Whisper was trained on multilingual data and handles languages such as Mandarin reliably, including code-switching, with the larger model sizes giving noticeably better results.

What output formats does it produce?

Plain text, segment-level JSON, and ready-to-use subtitle formats such as SRT and VTT, which is what makes it ideal for video captioning.

How accurate is it for meetings with several speakers?

Transcription accuracy is high, but Whisper does not identify speakers — use a tool such as Descript or a dedicated diarisation step if you need speaker labels.

Last reviewed: 2026-09-01

More Audio & Video tools

View all →

How we review