자연스러운 음성 합성을 브라우저에서.
Kokoro-82M 기반 온디바이스 신경 음성 합성. 영어와 중국어(표준 중국어) 총 128가지 음성, 서버 제로.
신경 음성. 업로드 제로.
Transformers.js를 통해 기기에서 완전히 실행되는 스튜디오 품질의 TTS.
100% 프라이빗
텍스트는 절대 브라우저를 떠나지 않습니다. 모든 합성은 로컬에서 이루어집니다.
WebGPU 가속
GPU를 사용할 수 있으면 GPU에서 추론하고, WASM으로 폴백합니다.
Kokoro-82M
Apache-2.0 경량 신경 TTS. 영어 엔진과 중영 혼합 문장도 자연스럽게 읽는 표준 중국어(v1.1-zh) 엔진을 함께 제공합니다.
128가지 음성
미국/영국 영어 28가지 음성(A-F 품질 등급)과 표준 중국어 100가지 음성(여성/남성).
세 단계. 서버 제로.
- 01
텍스트 입력 또는 붙여넣기
음성으로 변환할 텍스트를 입력하세요.
- 02
음성 선택
영어 또는 표준 중국어 128가지 음성 중에서 선택. 필요에 따라 속도를 조절하세요.
- 03
재생 또는 다운로드
브라우저에서 듣거나 WAV 파일로 다운로드하세요.
About the neural text-to-speech tool
A free text-to-speech tool that generates natural-sounding speech entirely in your browser with the Kokoro neural voice model. Use it for narration drafts, voiceovers for video, language learning, accessibility, or previewing writing aloud — with no account, no API keys, and no per-word pricing.
How it works
Kokoro is a compact, high-quality neural TTS model running via Transformers.js, with WebGPU acceleration when the browser supports it. Two engines ship with the tool: 28 American and British English voices (graded A–F for naturalness) and 100 Mandarin voices, including a v1.1-zh build that reads mixed Chinese–English text naturally. Text is synthesized locally into audio you can play instantly and export as WAV. Voices and model weights download once and are then cached, after which synthesis works offline.
Limits & requirements
First run downloads model and voice data. Each generation accepts up to 500 characters, so long passages are best processed section by section; synthesis itself returns in seconds once the models are cached. Quality is strongest on English and Mandarin — the two engines shipped here.
Privacy
Cloud TTS services receive every word you type — including, routinely, confidential draft content. Here your text never leaves the browser. Scripts, notes, and documents you read aloud remain private by construction.