On-device TTS · Kokoro-82M

Texto em fala natural no seu navegador.

Síntese de fala neural on-device movida por Kokoro-82M. 128 vozes em inglês e mandarim, zero servidores.

Detectando GPU…
128 neural voices
100% privado
Por que Voice Synth

Vozes neurais. Zero uploads.

TTS com qualidade de estúdio rodando inteiramente no seu dispositivo via Transformers.js.

100% privado

Seu texto nunca sai do navegador. Toda a síntese acontece localmente.

Acelerado por WebGPU

A inferência roda na sua GPU quando disponível, com fallback para WASM.

Kokoro-82M

Um TTS neural open-weight (Apache-2.0): um motor em inglês e um motor de mandarim (v1.1-zh) que lê naturalmente textos mistos chinês-inglês.

128 vozes

28 vozes inglesas (EUA/Reino Unido, com qualidades A-F) e 100 vozes em mandarim (femininas/masculinas).

Como funciona

Três passos. Zero servidores.

  1. 01

    Digite ou cole o texto

    Insira o texto que deseja converter em fala.

  2. 02

    Escolha uma voz

    Escolha entre 128 vozes neurais - inglês ou mandarim. Ajuste a velocidade se necessário.

  3. 03

    Ouça ou baixe

    Ouça no navegador ou baixe o áudio como arquivo WAV.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Suporte

Perguntas respondidas.

Não. Tudo roda no seu navegador. O modelo é baixado uma única vez do Hugging Face e fica em cache local; depois disso, a síntese é 100% no dispositivo.