sanoTTS (ultra-light offline)
Offline neural text-to-speech at 294k-2.3M parameters — pure numpy, no torch, no GPU. 28 voices across 14 languages, each 0.3-9 MB, downloaded on first use and fully offline after that.
Settings Add-ons sanoTTS (ultra-light offline) Install
Subtitld downloads it, checks its SHA-256, and keeps it updated.
- Languages
- 14
- Voices
- 28
- Add-on download
- 38–63 MB
- RAM, at least
- 512 MB
Voices
28 voices in 14 languages.
ar-jo1cs-cz2de-de2en-us7es-es2fr-fr1id-id1it-it2pt-br2ro-ro2ru-ru2tr-tr2vi-vn1zh-cn1
Options
In Settings › Add-ons › sanoTTS (ultra-light offline) › Configure.
1.0Applied at synthesis time by stretching phoneme durations, which sounds better than resampling afterwards.
2Threads numpy/OpenBLAS may use. The default of 2 keeps synthesis from saturating every core while you are editing.
Where voice packs are fetched from on first use. Pin one if the other is blocked on your network.
Optional. Point at an already-extracted voice pack to synthesize with no network access at all.
Languages
ar-jocs-czde-deen-uses-esfr-frid-idit-itpt-brro-roru-rutr-trvi-vnzh-cn
Platforms
| System | Download |
|---|---|
| Linux x86-64 | 63.2 MB |
| macOS (Apple silicon) | 37.6 MB |
License
GPL-3.0-or-later (add-on + bundled espeak-ng phonemizer); sanoTTS runtime is MIT; voice weights download on first use and are GPL-3.0
Versions
- 0.1.0
Source for the GPL components
This add-on is distributed under the GNU General Public License v3.0 or later. Complete Corresponding Source for the GPL-licensed parts (espeak-ng, phonemizer-fork) is attached to this release as
*-corresponding-source.tar.gz, at no charge, per GPL-3.0 section 6(d).Each binary archive also contains a
THIRD_PARTY_LICENSES/directory with the full licence texts and the exact upstream revisions used.
Other Voices add-ons
-
Piper Offline
Lightweight offline neural TTS based on Piper. 158 voices across 49 languages; per-voice models range 20-110 M…
-
Kokoro Offline
Lightweight offline neural TTS. ~80 MB model, sub-second per line on CPU, 54 preset voices across 9 languages.…
-
Supertonic Offline
Lightning-fast on-device multilingual TTS. 99M-parameter ONNX model, ~400 MB download, 31 languages, 10 preset…
-
Parler-TTS Offline
Prompt-driven offline TTS. 34 named consistent speakers + free-text voice description. ~2 GB model, ~10-30s per…
-
Qwen3-TTS Clone
Multilingual neural TTS with 9 premium timbres + 3-second voice cloning. Heavy: ~3.5 GB model, 8+ GB RAM, CPU or…
-
Coqui XTTS Clone
Multilingual neural TTS with 6-second voice cloning. Heavy: ~2 GB model, CPU or CUDA.
-
F5-TTS Clone
Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh…