Try it: text to speech
Type a short line and hear it read aloud by Kokoro, an open TTS model. Runs free on our own machine, no account needed.
Try it: text to speech
Type a short line and hear it read aloud by Kokoro, an open TTS model. Runs free on our own machine, no account needed.
Run text to speech on your own machine
The live demo runs Kokoro on one small server, so it is offline at times. Kokoro is an open text-to-speech model that runs on a normal CPU. One script downloads it and turns a line of text into a WAV file on your computer.
What you need
- Python
- Version 3.10 or newer. Get it from python.org if you don't have it.
- Hardware
- CPU only. No graphics card needed.
- Disk space
- About 325 MB for the Kokoro model files, plus the Python packages.
- Internet
- Only for the first run, to install the packages and fetch the model files.
- Tested on
- Windows 11 with Python 3.13. Written for Windows, Mac and Linux, but only run on Windows so far.
Install Python
Skip this if you already have Python 3.10 or newer. Check with the command below, and install it from python.org if the command fails. On Mac and Linux the command may be python3.
python --versionDownload the script
Use the button above and save the file in any folder.
Run one command
Open a terminal in that folder and give it a line to speak. The first run installs kokoro-onnx into a private environment and downloads the model files, so give it a few minutes.
python setup_text_to_speech.py "Hello from my own machine"Play the result
The audio is saved as speech.wav in the same folder. Open it in any player.
Change the voice, language or speed
Add flags to pick a voice, a language code (en-us, en-gb, fr-fr, es, it, pt-br or hi), a speed from 0.5 to 2.0 and an output file name.
python setup_text_to_speech.py "Bonjour tout le monde" --lang fr-fr --voice ff_siwis --speed 1.1 --out bonjour.wav
What Kokoro gives you, and its limits
Kokoro is an open-weights text-to-speech model that runs on a CPU. The script uses its ONNX build through the kokoro-onnx package. It covers English (US and UK), Spanish, French, Hindi, Italian and Brazilian Portuguese, with a different set of voices for each.
Keep each run to a few sentences. The hosted demo caps input at 300 characters. For a long script, split it into paragraphs and run the command once per paragraph.
Japanese and Chinese are not covered by this script. They need extra language packages that it does not install. On some Linux setups Kokoro also needs the espeak-ng system package.
Common questions
- Is this text to speech free?
- Yes. Kokoro is open source and runs on your own computer, so there is no account, subscription or character quota.
- Does the audio leave my computer?
- No. The text is turned into speech locally. Only the first run uses the internet, to install packages and fetch the model files.
- Which voices and languages can I use?
- English (US and UK), Spanish, French, Hindi, Italian and Brazilian Portuguese. Pass a language with --lang and a voice with --voice. The live demo lists the voice names for each language.
- How do I uninstall it?
- Delete the genlovers-tools/text-to-speech folder in your home directory. It holds the private environment and the model files.
