Free private local audio transcriber
Turn speech into timed subtitles directly in your browser. Fully private, WebGPU accelerated, and exports clean SRT, VTT, or plain text.
Drop an audio or video file here
or click to choose files from your device
Up to 200 MB · Demuxed in browser memory · Local Whisper, never uploaded
The default: multilingual, accurate enough for real use, and small enough to download once without thinking about it.
Free to use · No account · No audio upload · Offline after model download
What does it mean to transcribe audio locally?
To transcribe audio locally means your voice recording never touches a remote server. Instead of uploading megabytes of confidential interviews, internal meeting minutes, or unpublished podcast footage to a third-party cloud API, your browser downloads a Whisper neural model once and executes speech-to-text inference directly on your machine's processor.
Every spoken sentence is sliced into millisecond-accurate segments, decoded into phonetic tokens, and formatted as subtitles on your device. For an in-depth breakdown of in-tab speech recognition architectures, explore our guide on how browser speech recognition works.
How speech-to-text runs in three steps
Everything executes sequentially inside your browser with complete transparency.
- 01
Decode in the browser tab
The browser decodes your audio or video file into 16 kHz single-channel floating-point audio in memory. No network packets are sent, and no command-line tools are required.
- 02
Run Whisper on local hardware
The OpenAI Whisper model runs locally via WebAssembly or WebGPU. It segments speech, predicts phonetic tokens, and formats timestamped text cues right on your GPU or CPU.
- 03
Export SRT, VTT, or translate
Review timed cues, optionally translate them into another language, and export production-ready SRT, WebVTT, or clean plain text drafts with preserved sentence boundaries.
Speech models tuned for local devices
Speech models are downloaded once and cached permanently in IndexedDB. Choose the right balance between download size, processing speed, and phonetic fidelity.
| Model | Download | Languages | Best Application |
|---|---|---|---|
| Whisper tiny | 39 MB | Multilingual (99+) | The smallest multilingual option. Noticeably less accurate, but it runs on almost anything and is the sensible choice without a GPU. |
| Whisper base | 73 MB | Multilingual (99+) | The default: multilingual, accurate enough for real use, and small enough to download once without thinking about it. |
| Whisper small | 238 MB | Multilingual (99+) | Clearly better on accents, background noise and proper nouns. Large enough that it is only offered on a desktop with WebGPU. |
WebGPU Acceleration: Chrome, Edge, and Chromium browsers automatically accelerate Whisper on your device's graphics processor. If WebGPU is unavailable, Whisper falls back to optimised WebAssembly SIMD threads on your CPU.
What to know before you transcribe audio
Local machine learning delivers complete privacy, but hardware constraints shape what is realistic on user devices.
- Acoustic clarity drives output accuracy
- Single-speaker recordings captured with a close microphone transcribe with remarkable fidelity. In contrast, heavy background music, room echo, overlapping crosstalk, and phone-compressed audio reduce accuracy. Whisper produces a continuous transcript without automated speaker diarisation.
- Hallucination risk during prolonged silence
- Like all autoregressive encoder-decoder speech models, Whisper can occasionally hallucinate grammatically fluent sentences during extended silent pauses or unintelligible murmurs. Reviewing timestamps against your recording ensures drafts remain accurate.
- Subtitles rather than burnt-in video
- This page outputs SRT and WebVTT subtitle files rather than re-encoding video streams. Burning subtitles into high-definition video inside a browser requires immense memory and battery. Standalone subtitle files load instantly in any player. Read why zero-upload architecture matters in our essay on why private translation matters.
Transcribe speech by language
Each language presents distinct acoustic challenges — tonal inflection in Vietnamese, agglutination in Turkish, and dialectal diversity in Arabic. Choose a dedicated language guide to explore nuanced behavior.
Frequently asked questions
Everything you need to know about browser-based audio transcription, model downloads, and subtitle formatting.
How do I transcribe audio without uploading it?
To transcribe audio without uploading it to external servers, drop an audio or video file onto this page and press Transcribe. The browser decodes the audio track into raw audio buffers, an on-device Whisper speech recognition model downloaded to your device turns it into timed text, and nothing leaves your computer. You can watch DevTools → Network while it runs: the model weights are cached locally, but your recording never touches an external server.
Does it work on video files as well?
Yes. The browser decodes the audio track straight out of MP4, WebM, and MOV containers. You get timed subtitles without running ffmpeg or extracting the audio track first. Formats the browser cannot decode natively are cleanly reported.
How long does transcription take?
On modern machines with WebGPU acceleration, transcription runs significantly faster than real-time: a 10-minute speech clip transcribes in 2 to 3 minutes. On CPU-only devices, Whisper base processes 110 seconds of speech in about 84 seconds. Whisper tiny is roughly twice as fast.
Which Whisper speech model should I pick?
Whisper base (73 MB) is the balanced default for most speech recordings: multilingual, robust against background noise, and quick to load. Whisper tiny (39 MB) is ideal for older CPUs or rapid English dictation. Whisper small (238 MB) delivers higher fidelity on heavy accents when WebGPU is enabled.
Can I export subtitles rather than plain text?
Yes, timed cues are the native output format. You can download the result as an SRT or WebVTT subtitle file ready to load into VLC, mpv, YouTube, or Premiere, or download plain text without timecodes. If translated, both languages can be saved together.
How accurate is browser-based speech recognition?
Very high for clear single-speaker audio in well-supported languages. Accuracy degrades gracefully with heavy reverberation, background music, or phone-grade compression. Reviewing the timed cues before final export is always recommended.
Is there a file length or size limit?
Files up to 200 MB are supported. The practical constraint is device memory, as the browser decodes audio tracks into RAM before feeding chunks to Whisper. Half-hour and hour-long lectures transcribe comfortably on desktop devices.
Transcribe audio with this free audio transcriber
Drop in your recording above, inspect speech segments with millisecond timecodes, and export clean SRT subtitles with zero cloud tracking.