On-device TTS · Kokoro-82M

Natural text-to-speech in your browser.

On-device neural text-to-speech powered by Kokoro-82M. 128 voices across English and Mandarin Chinese, zero servers.

Detecting GPU…
128 neural voices
100% private
Why Voice Synth

Neural voices. Zero uploads.

Studio-quality TTS running entirely on your device via Transformers.js.

100% private

Your text never leaves your browser. All synthesis happens locally.

WebGPU accelerated

Inference runs on your GPU when available, with WASM fallback.

Kokoro-82M

Apache-2.0 open-weight neural TTS: an English engine plus a native Mandarin engine (v1.1-zh) that reads mixed Chinese-English text naturally.

128 voices

28 American & British English voices (A-F quality grades) and 100 Mandarin female/male voices.

How it works

Three steps. Zero servers.

  1. 01

    Type or paste text

    Enter the text you want to convert to speech.

  2. 02

    Pick a voice

    Choose from 128 neural voices - English or Mandarin. Adjust speed as needed.

  3. 03

    Play or download

    Listen in your browser or download the audio as a WAV file.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Support

Questions, answered.

No. Everything runs in your browser. The model downloads once from Hugging Face and is cached locally — after that, synthesis is 100% on-device.