Convert Audio & Video to Text Privately in Your Browser
A free, offline-capable AI transcriber powered by WebAssembly and WebGPU. Your files never leave your device — all processing happens locally, right here in your browser.
Drag & drop your file here
or click to browse
How It Works
This transcriber runs an OpenAI Whisper speech-recognition model directly inside your web browser. There is no back-end server involved in the transcription itself. Here is the full journey your audio takes — none of which leaves your computer.
Select a File
Drag an MP3, WAV, M4A, or MP4 file onto the drop zone, or click to browse. The file is read into memory using the browser's File API and is never transmitted anywhere.
Decode & Resample
The native Web Audio API decodes your file and resamples it to a 16 kHz mono Float32 waveform — the exact format the Whisper model expects for accurate recognition.
Local AI Inference
Our local transformers loads a quantized Whisper model from files stored on this same site and runs inference using WebAssembly, accelerated by WebGPU where available.
Read, Copy, Download
The recognized text appears in the transcript box. Copy it to your clipboard or download it as a plain-text file. Everything stayed on your device the whole time.
Key Benefits
Browser-based, on-device AI changes what a free web utility can offer. Here is why local transcription is a meaningful upgrade over traditional cloud tools.
Genuine Privacy
Because inference happens locally, sensitive interviews, medical notes, legal recordings, or personal voice memos are never uploaded to a third-party server or stored in the cloud.
Works Offline
Once the model files are cached by your browser, the tool continues to work without an internet connection — ideal for travel, fieldwork, or restricted environments.
No Usage Limits or Fees
There are no API keys, subscriptions, per-minute charges, or account sign-ups. Transcribe as many files as your own hardware can handle, completely free.
Hardware Accelerated
On supported devices, WebGPU taps into your graphics hardware for faster inference, while WebAssembly provides a reliable fallback on every modern browser.
Multilingual Recognition
The underlying Whisper model was trained on many languages, so it can recognize speech well beyond English and automatically adapt to the audio it receives.
Nothing to Install
No desktop software, browser extension, or command-line setup. The entire application is a static web page that loads instantly and runs in any current browser.
Frequently Asked Questions
Answers to the questions people most often ask about local, in-browser transcription.
Are my files really private?
Yes. The transcription model runs entirely within your browser tab using WebAssembly and WebGPU. Your audio is read into local memory, decoded by your browser, and processed on your own device. No part of your file is uploaded to a server, and we never receive or store your recordings.
Which file formats are supported?
You can transcribe common audio and video formats including MP3, WAV, M4A, and MP4. The tool uses your browser's built-in audio decoder, so any format your browser can play should work. For best results, use clear recordings with minimal background noise.
Why does it take a moment to load the first time?
The first time you use the tool, the browser downloads the AI model files (roughly a few tens of megabytes for the compact "tiny" model). After that initial download, your browser caches the files, so subsequent visits and transcriptions start much faster — and can even work offline.
How accurate is the transcription?
Accuracy depends on audio quality, accent, background noise, and the model size. The "tiny" model prioritizes speed and small download size, which makes it excellent for quick drafts and clear speech. It may occasionally misinterpret difficult audio, so we recommend proofreading important transcripts.
Does it work on mobile phones?
It can, on modern mobile browsers, though performance depends on your device's memory and processor. Because AI inference is computationally intensive, a desktop or laptop with WebGPU support will generally deliver the fastest experience.
Is there a limit on how long a file can be?
There is no hard limit imposed by the tool, but longer recordings require more memory and time to process. The application automatically splits long audio into overlapping chunks so it can handle extended recordings gracefully. Very large files may still be constrained by your device's available memory.
Do you use cookies or advertising?
This site displays advertisements to keep the tool free, and our advertising partners — including Google — may use cookies to serve relevant ads. The transcription feature itself uses no tracking. Please review our Privacy Policy for full details on cookies and how to opt out.