Skip to content
GenLovers

Try it: speech to text

Record or upload a short clip and get a transcript back from Parakeet, an open speech-to-text model running live on our own machine.

Try it: speech to text

Record or upload a short clip and get a transcript back from Parakeet, an open speech-to-text model running live on our own machine.

Checking server
or upload a file

Auto-detects English and Spanish best; other European languages work but less reliably, and Mandarin isn't supported.

Run speech to text on your own machine

The live demo runs Parakeet on one small server, so it is offline at times. Parakeet is an open speech-to-text model from NVIDIA that runs on a normal CPU. One script installs it and transcribes a WAV file on your computer.

What you need

Python
Version 3.10 or newer. Get it from python.org if you don't have it.
Hardware
CPU only. No graphics card needed.
Disk space
About 640 MB for the model, plus the Python packages.
Input
A WAV file. Record one with any voice recorder, or export one from your audio editor.
Internet
Only for the first run, to install the packages and download the model.
Tested on
Windows 11 with Python 3.13. Written for Windows, Mac and Linux, but only run on Windows so far.
Download setup_speech_to_text.py
  1. Install Python

    Skip this if you already have Python 3.10 or newer. Check with the command below, and install it from python.org if the command fails. On Mac and Linux the command may be python3.

    python --version
  2. Download the script

    Use the button above and save the file in the folder that holds your recording.

  3. Run one command

    Open a terminal in that folder and pass your recording. The first run installs onnx-asr into a private environment and downloads the Parakeet model (an int8 build), so give it a few minutes.

    python setup_speech_to_text.py recording.wav
  4. Read the transcript

    The transcript prints in the terminal and is saved next to the audio as recording.txt.

What Parakeet handles well, and what it does not

Parakeet TDT 0.6B v3 is NVIDIA's open speech-to-text model. The script runs an int8 build of it through the onnx-asr package, which works on a CPU and returns a transcript with punctuation.

It detects the language on its own and covers European languages. English and Spanish are where it is strongest. Other European languages work but less reliably, and Mandarin is not supported.

The script expects a WAV file. If your recording is an MP3, M4A or WebM file, convert it to WAV first with a free tool such as ffmpeg or Audacity.

Nothing leaves your machine after setup, which suits private recordings such as interviews or voice memos.

Common questions

Is this speech to text free?
Yes. Parakeet is open source and runs on your computer, so there is no account, subscription or minute limit.
Which languages does it support?
It detects the language automatically and covers European languages. English and Spanish work best, other European languages work less reliably, and Mandarin is not supported.
Is my audio uploaded anywhere?
No. Transcription runs locally. Only the first run uses the internet, to install packages and download the model.
How do I transcribe an MP3?
Convert it to WAV first, for example with ffmpeg (ffmpeg -i talk.mp3 talk.wav), then run the script on the WAV file.
How do I uninstall it?
Delete the genlovers-tools/speech-to-text folder in your home directory. The model itself is stored in your Hugging Face cache folder (.cache/huggingface).