On-device AI · WebGPU

Transcribe audio right in your browser.

Whisper speech-to-text, fully on-device. Upload audio files or record live — no servers, no limits.

Detecting GPU…
Upload or record
100% private
Why Speech-to-Text

Whisper-quality transcripts. Zero uploads.

OpenAI's Whisper model running locally in your browser via Transformers.js.

100% private

Your audio never leaves your device. All transcription happens locally.

WebGPU accelerated

Inference runs on your GPU when available, with WASM fallback.

Whisper model

Multilingual speech recognition with automatic language detection.

Upload or record

Drag & drop audio files (MP3, WAV, M4A) or record live from your microphone.

How it works

Three steps. Zero servers.

  1. 01

    Upload or record

    Drag & drop an audio file, or record live. MP3, WAV, M4A, and more.

  2. 02

    AI transcribes

    Whisper runs locally, transcribing your audio to text with timestamps.

  3. 03

    Copy or download

    Copy the transcript or download as a text file. The model caches for reuse.

Complete guide

About the speech-to-text tool

A free speech-to-text tool that transcribes audio and video locally in your browser, powered by a Whisper model running on your device. It is intended for meetings, interviews, lectures, podcasts, and voice memos — including sensitive recordings that should not be uploaded to a cloud transcription service.

How it works

Audio is decoded in the browser with the Web Audio API, then a Whisper model (transformer-based speech recognition) runs through Transformers.js. On WebGPU-enabled Chrome or Edge the model executes on your GPU and a one-hour file transcribes in a fraction of real time; on other browsers a WebAssembly path keeps the same features available. The model files download once and are cached, so you can transcribe offline afterward. Finished transcripts can be copied as plain text or exported with timestamps.

Limits & requirements

The browser ships the lightweight Whisper tier for fast, free inference; it is excellent for clear speech but less forgiving on heavy accents or noisy multi-speaker recordings than paid large-model APIs. Long multi-hour files may hit browser memory limits on low-RAM devices; splitting long recordings is the workaround.

Privacy

This is the key difference from cloud services: your recordings are never uploaded. Interviews, confidential calls, and personal memos are processed and stay on your machine, with no per-minute billing.

Support

Questions, answered.

No. Everything runs in your browser. The Whisper model downloads once from Hugging Face and is cached locally — after that, transcription is 100% on-device.