Qwen3-TTS (offline, 1.7B)
Multilingual neural TTS with 9 premium timbres + 3-second voice cloning. Heavy: ~3.5 GB model, 8+ GB RAM, CPU or CUDA.
Settings Add-ons Qwen3-TTS (offline, 1.7B) Install
Subtitld downloads it, checks its SHA-256, and keeps it updated.
- Languages
- 10
- Models
- 3
- Voices
- 10
- Add-on download
- 317–456 MB
- RAM, at least
- 8 GB
Models
| Model | Download |
|---|---|
| qwen3-tts-1_7b-customvoice | |
| qwen3-tts-1_7b-base | |
| qwen3-tts-tokenizer-12hz |
Voices
10 voices in 4 languages.
en2ja1ko1zh-cn5
Options
In Settings › Add-ons › Qwen3-TTS (offline, 1.7B) › Configure.
cpuautoUsed by the 'qwen3-clone' voice. Per-speaker overrides set in the speaker panel take precedence.
Optional. If provided alongside the reference clip, clone quality is noticeably better. Leave empty to use x-vector-only mode.
When the dubbing UI extracts a reference clip from the project audio, it normally also forwards the matching subtitle text as the transcript to improve clone accent fidelity. Enable this to suppress the transcript — qwen3 will fall back to its built-in ASR. Useful when the subtitle text paraphrases rather than transcribes verbatim, or when the source language is uncertain.
Languages
zh-cnenjakodefrruptesit
Platforms
| System | Download |
|---|---|
| Linux x86-64 | 455.7 MB |
| macOS (universal) | 345.8 MB |
| Windows x86-64 | 316.5 MB |
License
Apache-2.0 (wrapper) + Apache-2.0 (Qwen3-TTS weights)
Versions
- 1.0.6
Automated release for qwen3-tts v1.0.6.
Other Voices add-ons
-
Piper Offline
Lightweight offline neural TTS based on Piper. 158 voices across 49 languages; per-voice models range 20-110 M…
-
Kokoro Offline
Lightweight offline neural TTS. ~80 MB model, sub-second per line on CPU, 54 preset voices across 9 languages.…
-
Supertonic Offline
Lightning-fast on-device multilingual TTS. 99M-parameter ONNX model, ~400 MB download, 31 languages, 10 preset…
-
sanoTTS Offline
Offline neural text-to-speech at 294k-2.3M parameters — pure numpy, no torch, no GPU. 28 voices across 14 languages,…
-
Parler-TTS Offline
Prompt-driven offline TTS. 34 named consistent speakers + free-text voice description. ~2 GB model, ~10-30s per…
-
Coqui XTTS Clone
Multilingual neural TTS with 6-second voice cloning. Heavy: ~2 GB model, CPU or CUDA.
-
F5-TTS Clone
Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh…