On-device TTS · Kokoro-82M

Text-to-speech naturale nel tuo browser.

Sintesi vocale neurale on-device basata su Kokoro-82M. 128 voci in inglese e cinese mandarino, zero server.

Rilevamento della GPU…
128 neural voices
100% privato
Perché Voice Synth

Voci neurali. Zero upload.

TTS di qualità da studio interamente sul tuo dispositivo tramite Transformers.js.

100% privato

Il tuo testo non lascia mai il browser. Tutta la sintesi avviene in locale.

Accelerato WebGPU

L'inferenza gira sulla tua GPU quando disponibile, con fallback WASM.

Kokoro-82M

Un TTS neurale open-weight (Apache-2.0): un motore inglese e un motore mandarino (v1.1-zh) che legge in modo naturale testi misti cinese-inglese.

128 voci

28 voci inglesi (US/UK, con qualità A-F) e 100 voci in mandarino (femminili/maschili).

Come funziona

Tre passaggi. Zero server.

  1. 01

    Digita o incolla il testo

    Inserisci il testo che vuoi convertire in voce.

  2. 02

    Scegli una voce

    Scegli tra 128 voci neurali - inglese o mandarino. Regola la velocità se necessario.

  3. 03

    Ascolta o scarica

    Ascolta nel browser o scarica l'audio come file WAV.

Complete guide

About the neural text-to-speech tool

A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.

How it works

Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.

Limits & requirements

First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.

Privacy

Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.

Supporto

Domande, con risposte.

No. Tutto gira nel tuo browser. Il modello viene scaricato una sola volta da Hugging Face e memorizzato nella cache locale: dopodiché la sintesi è al 100 % sul dispositivo.