On-device TTS · Kokoro-82M

Natürliche Text-to-Speech in deinem Browser.

On-Device-neurale Text-to-Speech mit Kokoro-82M. 128 Stimmen für Englisch und Mandarin-Chinesisch, keine Server.

GPU wird erkannt…
128 neural voices
100 % privat
Warum Voice Synth

Neuronale Stimmen. Keine Uploads.

TTS in Studioqualität, läuft vollständig auf deinem Gerät via Transformers.js.

100 % privat

Dein Text verlässt deinen Browser nie. Die gesamte Synthese erfolgt lokal.

WebGPU-beschleunigt

Die Inferenz läuft auf deiner GPU, wenn verfügbar — mit WASM-Fallback.

Kokoro-82M

Leichtgewichtiges Open-Weight-TTS unter Apache-2.0: ein englischer und ein Mandarin-Build (v1.1-zh), der gemischte chinesisch-englische Texte natürlich liest.

128 Stimmen

28 englische Stimmen (US/UK, mit Qualitätsstufen A-F) plus 100 Mandarin-Stimmen (weiblich/männlich).

So funktioniert's

Drei Schritte. Keine Server.

  1. 01

    Text eingeben oder einfügen

    Gib den Text ein, der in Sprache umgewandelt werden soll.

  2. 02

    Stimme wählen

    Wähle aus 128 neuronalen Stimmen - Englisch oder Mandarin. Passe bei Bedarf das Tempo an.

  3. 03

    Abspielen oder herunterladen

    Im Browser anhören oder das Audio als WAV-Datei herunterladen.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Support

Fragen, beantwortet.

Nein. Alles läuft in deinem Browser. Das Modell wird einmal von Hugging Face geladen und lokal zwischengespeichert – danach läuft die Synthese zu 100 % auf deinem Gerät.