Supertonic (on-device, ONNX, 99M)
Lightning-fast on-device multilingual TTS. 99M-parameter ONNX model, ~400 MB download, 31 languages, 10 preset voices. No GPU required.
Settings Add-ons Supertonic (on-device, ONNX, 99M) Install
Subtitld downloads it, checks its SHA-256, and keeps it updated.
- Languages
- 32
- Models
- 1
- Voices
- 10
- Add-on download
- 46–78 MB
- RAM, at least
- 2 GB
Models
| Model | Download |
|---|---|
| supertonic-3 |
Voices
10 voices in 0 languages.
Options
In Settings › Add-ons › Supertonic (on-device, ONNX, 99M) › Configure.
supertonic-3Larger model = broader language coverage, longer first-run download. Switching reloads the model on the next synthesis call.
cpu8Higher = cleaner audio, slower. 5 is fastest preview, 8 is the upstream default, 12 is the quality ceiling. Past 12 the model saturates — additional steps cost compute without measurable improvement.
autoPath to a Voice Builder export (https://supertonic.supertone.ai/voice_builder). When set, the addon exposes one additional voice id 'supertonic-custom' that loads this JSON instead of a preset. The 10 preset voices remain available.
Languages
arbghrcsdanlenetfifrdeelhihuiditjakolvltplptroruskslessvtrukvina
Platforms
| System | Download |
|---|---|
| Linux x86-64 | 78.1 MB |
| macOS (universal) | 49.1 MB |
| Windows x86-64 | 45.6 MB |
License
MIT (wrapper) + MIT (supertonic-py) + OpenRAIL-M (Supertonic-3 weights)
Versions
- 0.0.2
Automated release for supertonic v0.0.2.
Other Voices add-ons
-
Piper Offline
Lightweight offline neural TTS based on Piper. 158 voices across 49 languages; per-voice models range 20-110 M…
-
Kokoro Offline
Lightweight offline neural TTS. ~80 MB model, sub-second per line on CPU, 54 preset voices across 9 languages.…
-
sanoTTS Offline
Offline neural text-to-speech at 294k-2.3M parameters — pure numpy, no torch, no GPU. 28 voices across 14 languages,…
-
Parler-TTS Offline
Prompt-driven offline TTS. 34 named consistent speakers + free-text voice description. ~2 GB model, ~10-30s per…
-
Qwen3-TTS Clone
Multilingual neural TTS with 9 premium timbres + 3-second voice cloning. Heavy: ~3.5 GB model, 8+ GB RAM, CPU or…
-
Coqui XTTS Clone
Multilingual neural TTS with 6-second voice cloning. Heavy: ~2 GB model, CPU or CUDA.
-
F5-TTS Clone
Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh…