On-device TTS · Kokoro-82M

Texto a voz natural en tu navegador.

Texto a voz neuronal en el dispositivo impulsado por Kokoro-82M. 128 voces en inglés y chino mandarín, cero servidores.

Detectando GPU…
128 neural voices
100 % privado
Por qué Voice Synth

Voces neuronales. Cero subidas.

TTS de calidad de studio ejecutándose por completo en tu dispositivo vía Transformers.js.

100 % privado

Tu texto nunca sale de tu navegador. Toda la síntesis ocurre en local.

Acelerado por WebGPU

La inferencia se ejecuta en tu GPU cuando está disponible, con respaldo de WASM.

Kokoro-82M

Un TTS neuronal de código abierto (Apache-2.0): un motor en inglés y un motor de mandarín (v1.1-zh) que lee con naturalidad textos mixtos chino-inglés.

128 voces

28 voces inglesas (EE. UU./Reino Unido, con calidades A-F) y 100 voces en mandarín (femeninas/masculinas).

Cómo funciona

Tres pasos. Cero servidores.

  1. 01

    Escribe o pega texto

    Introduce el texto que quieras convertir en voz.

  2. 02

    Elige una voz

    Elige entre 128 voces neuronales - inglés o mandarín. Ajusta la velocidad si lo necesitas.

  3. 03

    Reproduce o descarga

    Escúchalo en tu navegador o descarga el audio como archivo WAV.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Soporte

Preguntas, respondidas.

No. Todo se ejecuta en tu navegador. El modelo se descarga una sola vez desde Hugging Face y se guarda en caché local: después, la síntesis es 100 % en el dispositivo.