On-Device-KI · WebGPU

Audio transkribieren direkt im Browser.

Whisper-Spracherkennung, vollständig on-device. Audiodateien hochladen oder live aufnehmen — keine Server, keine Limits.

GPU wird erkannt…
Hochladen oder aufnehmen
100 % privat
Warum Speech-to-Text

Transkripte in Whisper-Qualität. Keine Uploads.

OpenAIs Whisper-Modell läuft lokal in deinem Browser via Transformers.js.

100 % privat

Dein Audio verlässt dein Gerät nie. Die gesamte Transkription erfolgt lokal.

WebGPU-beschleunigt

Die Inferenz läuft auf deiner GPU, wenn verfügbar — mit WASM-Fallback.

Whisper-Modell

Mehrsprachige Spracherkennung mit automatischer Sprachenerkennung.

Hochladen oder aufnehmen

Audiodateien per Drag & Drop (MP3, WAV, M4A) oder live über dein Mikrofon aufnehmen.

So funktioniert's

Drei Schritte. Keine Server.

  1. 01

    Hochladen oder aufnehmen

    Audiodatei per Drag & Drop oder live aufnehmen. MP3, WAV, M4A und mehr.

  2. 02

    KI transkribiert

    Whisper läuft lokal und transkribiert dein Audio mit Zeitstempeln in Text.

  3. 03

    Kopieren oder herunterladen

    Transkript kopieren oder als Textdatei herunterladen. Das Modell wird zur Wiederverwendung zwischengespeichert.

Complete guide

About the speech-to-text tool

A free speech-to-text tool that transcribes audio and video locally in your browser, powered by a Whisper model running on your device. It is intended for meetings, interviews, lectures, podcasts, and voice memos — including sensitive recordings that should not be uploaded to a cloud transcription service.

How it works

Audio is decoded in the browser with the Web Audio API, then a Whisper model (transformer-based speech recognition) runs through Transformers.js. On WebGPU-enabled Chrome or Edge the model executes on your GPU and a one-hour file transcribes in a fraction of real time; on other browsers a WebAssembly path keeps the same features available. The model files download once and are cached, so you can transcribe offline afterward. Finished transcripts can be copied as plain text or exported with timestamps.

Limits & requirements

The browser ships the lightweight Whisper tier for fast, free inference; it is excellent for clear speech but less forgiving on heavy accents or noisy multi-speaker recordings than paid large-model APIs. Long multi-hour files may hit browser memory limits on low-RAM devices; splitting long recordings is the workaround.

Privacy

This is the key difference from cloud services: your recordings are never uploaded. Interviews, confidential calls, and personal memos are processed and stay on your machine, with no per-minute billing.

Support

Fragen, beantwortet.

Nein. Alles läuft in deinem Browser. Das Whisper-Modell wird einmal von Hugging Face heruntergeladen und lokal zwischengespeichert – danach erfolgt die Transkription zu 100 % auf dem Gerät.