IA en el dispositivo · WebGPU

Transcribe audio directo en tu navegador.

Voz a texto con Whisper, totalmente en el dispositivo. Sube archivos de audio o graba en directo — sin servidores, sin límites.

Detectando GPU…
Subir o grabar
100 % privado
Por qué Speech-to-Text

Transcripciones con calidad Whisper. Cero subidas.

El modelo Whisper de OpenAI funcionando localmente en tu navegador vía Transformers.js.

100 % privado

Tu audio nunca sale de tu dispositivo. Toda la transcripción ocurre en local.

Acelerado por WebGPU

La inferencia se ejecuta en tu GPU cuando está disponible, con respaldo de WASM.

Modelo Whisper

Reconocimiento de voz multilingüe con detección automática de idioma.

Subir o grabar

Arrastra y suelta archivos de audio (MP3, WAV, M4A) o graba en directo desde tu micrófono.

Cómo funciona

Tres pasos. Cero servidores.

  1. 01

    Subir o grabar

    Arrastra y suelta un archivo de audio o graba en directo. MP3, WAV, M4A y más.

  2. 02

    La IA transcribe

    Whisper funciona localmente y transcribe tu audio a texto con marcas de tiempo.

  3. 03

    Copiar o descargar

    Copia la transcripción o descárgala como archivo de texto. El modelo se guarda en caché para reutilizarlo.

Complete guide

About the speech-to-text tool

A free speech-to-text tool that transcribes audio and video locally in your browser, powered by a Whisper model running on your device. It is intended for meetings, interviews, lectures, podcasts, and voice memos — including sensitive recordings that should not be uploaded to a cloud transcription service.

How it works

Audio is decoded in the browser with the Web Audio API, then a Whisper model (transformer-based speech recognition) runs through Transformers.js. On WebGPU-enabled Chrome or Edge the model executes on your GPU and a one-hour file transcribes in a fraction of real time; on other browsers a WebAssembly path keeps the same features available. The model files download once and are cached, so you can transcribe offline afterward. Finished transcripts can be copied as plain text or exported with timestamps.

Limits & requirements

The browser ships the lightweight Whisper tier for fast, free inference; it is excellent for clear speech but less forgiving on heavy accents or noisy multi-speaker recordings than paid large-model APIs. Long multi-hour files may hit browser memory limits on low-RAM devices; splitting long recordings is the workaround.

Privacy

This is the key difference from cloud services: your recordings are never uploaded. Interviews, confidential calls, and personal memos are processed and stay on your machine, with no per-minute billing.

Soporte

Preguntas, respondidas.

No. Todo se ejecuta en tu navegador. El modelo Whisper se descarga una vez desde Hugging Face y se guarda en caché localmente — a partir de ahí, la transcripción es 100 % en el dispositivo.