Try it: speech to text
Record or upload a short clip and get a transcript back from Parakeet, an open speech-to-text model running live on our own machine.
Try it: speech to text
Record or upload a short clip and get a transcript back from Parakeet, an open speech-to-text model running live on our own machine.
Run speech to text on your own machine
The live demo runs Parakeet on one small server, so it is offline at times. Parakeet is an open speech-to-text model from NVIDIA that runs on a normal CPU. One script installs it and transcribes a WAV file on your computer.
What you need
- Python
- Version 3.10 or newer. Get it from python.org if you don't have it.
- Hardware
- CPU only. No graphics card needed.
- Disk space
- About 640 MB for the model, plus the Python packages.
- Input
- A WAV file. Record one with any voice recorder, or export one from your audio editor.
- Internet
- Only for the first run, to install the packages and download the model.
- Tested on
- Windows 11 with Python 3.13. Written for Windows, Mac and Linux, but only run on Windows so far.
Install Python
Skip this if you already have Python 3.10 or newer. Check with the command below, and install it from python.org if the command fails. On Mac and Linux the command may be python3.
python --versionDownload the script
Use the button above and save the file in the folder that holds your recording.
Run one command
Open a terminal in that folder and pass your recording. The first run installs onnx-asr into a private environment and downloads the Parakeet model (an int8 build), so give it a few minutes.
python setup_speech_to_text.py recording.wavRead the transcript
The transcript prints in the terminal and is saved next to the audio as recording.txt.
What Parakeet handles well, and what it does not
Parakeet TDT 0.6B v3 is NVIDIA's open speech-to-text model. The script runs an int8 build of it through the onnx-asr package, which works on a CPU and returns a transcript with punctuation.
It detects the language on its own and covers European languages. English and Spanish are where it is strongest. Other European languages work but less reliably, and Mandarin is not supported.
The script expects a WAV file. If your recording is an MP3, M4A or WebM file, convert it to WAV first with a free tool such as ffmpeg or Audacity.
Nothing leaves your machine after setup, which suits private recordings such as interviews or voice memos.
Common questions
- Is this speech to text free?
- Yes. Parakeet is open source and runs on your computer, so there is no account, subscription or minute limit.
- Which languages does it support?
- It detects the language automatically and covers European languages. English and Spanish work best, other European languages work less reliably, and Mandarin is not supported.
- Is my audio uploaded anywhere?
- No. Transcription runs locally. Only the first run uses the internet, to install packages and download the model.
- How do I transcribe an MP3?
- Convert it to WAV first, for example with ffmpeg (ffmpeg -i talk.mp3 talk.wav), then run the script on the WAV file.
- How do I uninstall it?
- Delete the genlovers-tools/speech-to-text folder in your home directory. The model itself is stored in your Hugging Face cache folder (.cache/huggingface).
