On-device TTS · Kokoro-82M

Natuurlijke tekst-naar-spraak in je browser.

On-device neurale tekst-naar-spraak aangedreven door Kokoro-82M. 128 stemmen in Engels en Mandarijn, nul servers.

GPU detecteren…
128 neural voices
100% privé
Waarom Voice Synth

Neurale stemmen. Nul uploads.

TTS in studiokwaliteit die volledig op je apparaat draait via Transformers.js.

100% privé

Je tekst verlaat nooit je browser. Alle synthese gebeurt lokaal.

WebGPU-versneld

De inferentie draait op je GPU indien beschikbaar, met WASM-fallback.

Kokoro-82M

Een lichtgewicht open-weight TTS-model (Apache-2.0): een Engelse engine en een Mandarijn-engine (v1.1-zh) die gemengde Chinees-Engelse tekst natuurlijk leest.

128 stemmen

28 Engelse stemmen (VS/VK, met kwaliteitsniveaus A-F) en 100 Mandarijn-stemmen (vrouwelijk/mannelijk).

Hoe het werkt

Drie stappen. Nul servers.

  1. 01

    Typ of plak tekst

    Voer de tekst in die je naar spraak wilt omzetten.

  2. 02

    Kies een stem

    Kies uit 128 neurale stemmen - Engels of Mandarijn. Pas indien nodig de snelheid aan.

  3. 03

    Afspelen of downloaden

    Luister in je browser of download de audio als WAV-bestand.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Ondersteuning

Vragen, beantwoord.

Nee. Alles draait in je browser. Het model wordt één keer van Hugging Face gedownload en lokaal gecachet; daarna gebeurt de synthese 100% op je apparaat.