Synthèse vocale naturelle dans votre navigateur.
Synthèse vocale neuronale sur l'appareil propulsée par Kokoro-82M. 128 voix en anglais et en mandarin, zéro serveur.
Voix neuronales. Zéro envoi.
Une TTS de qualité studio, exécutée entièrement sur votre appareil via Transformers.js.
100 % privé
Votre texte ne quitte jamais votre navigateur. Toute la synthèse se fait en local.
Accéléré par WebGPU
L'inférence s'exécute sur votre GPU quand il est disponible, avec repli sur WASM.
Kokoro-82M
Un TTS neuronal open-weight (Apache-2.0) : un moteur anglais et un moteur mandarin (v1.1-zh) qui lit naturellement les textes mixtes chinois-anglais.
128 voix
28 voix anglaises (États-Unis/Royaume-Uni, qualités A-F) et 100 voix mandarin (féminines/masculines).
Trois étapes. Zéro serveur.
- 01
Saisissez ou collez du texte
Entrez le texte que vous voulez convertir en voix.
- 02
Choisissez une voix
Choisissez parmi 128 voix neuronales - anglais ou mandarin. Ajustez la vitesse si besoin.
- 03
Écoutez ou téléchargez
Écoutez dans votre navigateur ou téléchargez l'audio en fichier WAV.
About the neural text-to-speech tool
A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.
How it works
Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.
Limits & requirements
First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.
Privacy
Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.