Natürliche Text-to-Speech in deinem Browser.
On-Device-neurale Text-to-Speech mit Kokoro-82M. 128 Stimmen für Englisch und Mandarin-Chinesisch, keine Server.
Neuronale Stimmen. Keine Uploads.
TTS in Studioqualität, läuft vollständig auf deinem Gerät via Transformers.js.
100 % privat
Dein Text verlässt deinen Browser nie. Die gesamte Synthese erfolgt lokal.
WebGPU-beschleunigt
Die Inferenz läuft auf deiner GPU, wenn verfügbar — mit WASM-Fallback.
Kokoro-82M
Leichtgewichtiges Open-Weight-TTS unter Apache-2.0: ein englischer und ein Mandarin-Build (v1.1-zh), der gemischte chinesisch-englische Texte natürlich liest.
128 Stimmen
28 englische Stimmen (US/UK, mit Qualitätsstufen A-F) plus 100 Mandarin-Stimmen (weiblich/männlich).
Drei Schritte. Keine Server.
- 01
Text eingeben oder einfügen
Gib den Text ein, der in Sprache umgewandelt werden soll.
- 02
Stimme wählen
Wähle aus 128 neuronalen Stimmen - Englisch oder Mandarin. Passe bei Bedarf das Tempo an.
- 03
Abspielen oder herunterladen
Im Browser anhören oder das Audio als WAV-Datei herunterladen.
About the neural text-to-speech tool
A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.
How it works
Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.
Limits & requirements
First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.
Privacy
Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.