Texto a voz natural en tu navegador.
Texto a voz neuronal en el dispositivo impulsado por Kokoro-82M. 128 voces en inglés y chino mandarín, cero servidores.
Voces neuronales. Cero subidas.
TTS de calidad de studio ejecutándose por completo en tu dispositivo vía Transformers.js.
100 % privado
Tu texto nunca sale de tu navegador. Toda la síntesis ocurre en local.
Acelerado por WebGPU
La inferencia se ejecuta en tu GPU cuando está disponible, con respaldo de WASM.
Kokoro-82M
Un TTS neuronal de código abierto (Apache-2.0): un motor en inglés y un motor de mandarín (v1.1-zh) que lee con naturalidad textos mixtos chino-inglés.
128 voces
28 voces inglesas (EE. UU./Reino Unido, con calidades A-F) y 100 voces en mandarín (femeninas/masculinas).
Tres pasos. Cero servidores.
- 01
Escribe o pega texto
Introduce el texto que quieras convertir en voz.
- 02
Elige una voz
Elige entre 128 voces neuronales - inglés o mandarín. Ajusta la velocidad si lo necesitas.
- 03
Reproduce o descarga
Escúchalo en tu navegador o descarga el audio como archivo WAV.
About the neural text-to-speech tool
A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.
How it works
Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.
Limits & requirements
First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.
Privacy
Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.