F5-TTS (offline + voice clone)
Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh only.
Settings Add-ons F5-TTS (offline + voice clone) Install
Subtitld downloads it, checks its SHA-256, and keeps it updated.
- Languages
- 2
- Models
- 2
- Voices
- 1
- Add-on download
- 433–648 MB
- RAM, at least
- 4 GB
Models
| Model | Download |
|---|---|
| f5-tts-v1-base | |
| vocos-mel-24khz |
Voices
1 voices in 0 languages.
Options
In Settings › Add-ons › F5-TTS (offline + voice clone) › Configure.
cpuRequired by F5-TTS. Per-speaker overrides set in the speaker panel take precedence.
Optional. If empty, F5-TTS auto-transcribes via Whisper (slower first call).
When the dubbing UI extracts a reference clip from the project audio, it normally also forwards the matching subtitle text. Enable this to suppress the transcript — F5-TTS will fall back to Whisper auto-transcription (slower first call, cached afterwards). Useful when the subtitle text paraphrases rather than transcribes verbatim.
Languages
enzh-cn
Platforms
| System | Download |
|---|---|
| Linux x86-64 | 647.8 MB |
| macOS (universal) | 433.8 MB |
| Windows x86-64 | 433.2 MB |
License
MIT (wrapper) + CC-BY-NC-4.0 (F5TTS_v1_Base weights, non-commercial only)
Versions
- 1.0.6
Automated release for f5-tts v1.0.6.
Other Voices add-ons
-
Piper Offline
Lightweight offline neural TTS based on Piper. 158 voices across 49 languages; per-voice models range 20-110 M…
-
Kokoro Offline
Lightweight offline neural TTS. ~80 MB model, sub-second per line on CPU, 54 preset voices across 9 languages.…
-
Supertonic Offline
Lightning-fast on-device multilingual TTS. 99M-parameter ONNX model, ~400 MB download, 31 languages, 10 preset…
-
sanoTTS Offline
Offline neural text-to-speech at 294k-2.3M parameters — pure numpy, no torch, no GPU. 28 voices across 14 languages,…
-
Parler-TTS Offline
Prompt-driven offline TTS. 34 named consistent speakers + free-text voice description. ~2 GB model, ~10-30s per…
-
Qwen3-TTS Clone
Multilingual neural TTS with 9 premium timbres + 3-second voice cloning. Heavy: ~3.5 GB model, 8+ GB RAM, CPU or…
-
Coqui XTTS Clone
Multilingual neural TTS with 6-second voice cloning. Heavy: ~2 GB model, CPU or CUDA.