F5-TTS (offline + voice clone)

VoicesOfflineGPU optionalVoice cloningNon-commercial

Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh only.

Install it from Subtitld

Settings Add-ons F5-TTS (offline + voice clone) Install

Subtitld downloads it, checks its SHA-256, and keeps it updated.

Source on GitHub
Languages
2
Models
2
Voices
1
Add-on download
433–648 MB
RAM, at least
4 GB

Models

Model Download
f5-tts-v1-base
1.5 GB
vocos-mel-24khz
80 MB

Voices

1 voices in 0 languages.

Options

In Settings › Add-ons › F5-TTS (offline + voice clone) › Configure.

Inference device
Default: cpu
Default voice reference (6-15 s, mono WAV recommended)
Default: empty

Required by F5-TTS. Per-speaker overrides set in the speaker panel take precedence.

Reference clip transcript
Default: empty

Optional. If empty, F5-TTS auto-transcribes via Whisper (slower first call).

Skip auto-extracted reference transcript
Default: off

When the dubbing UI extracts a reference clip from the project audio, it normally also forwards the matching subtitle text. Enable this to suppress the transcript — F5-TTS will fall back to Whisper auto-transcription (slower first call, cached afterwards). Useful when the subtitle text paraphrases rather than transcribes verbatim.

Languages

  • en
  • zh-cn

Platforms

System Download
Linux x86-64 647.8 MB
macOS (universal) 433.8 MB
Windows x86-64 433.2 MB

License

Add-on
MIT (wrapper) + CC-BY-NC-4.0 (F5-TTS weights) — non-commercial only

MIT (wrapper) + CC-BY-NC-4.0 (F5TTS_v1_Base weights, non-commercial only)

Versions

  • 1.0.6

    Automated release for f5-tts v1.0.6.

Other Voices add-ons