Add-ons
Speech recognition, translation, voices, audio separation and lip-sync install from inside Subtitld, each with its own models. Most run on your own computer.
Speech to text
-
Whisper Offline
Offline speech to text with whisper.cpp, CPU only.
-
Vosk Offline
Light and CPU-only. Models from about 40 MB.
-
RealtimeSTT Offline
For Record mode, as you speak. Uses the GPU when there is one.
-
AssemblyAI Cloud
With speaker diarization. Not available yet.
Translation
Voices
-
Piper Offline
Lightweight offline neural TTS based on Piper. 158 voices across 49 languages; per-voice models range 20-110 MB.
-
Kokoro Offline
Lightweight offline neural TTS. ~80 MB model, sub-second per line on CPU, 54 preset voices across 9 languages. No voice …
-
Supertonic Offline
Lightning-fast on-device multilingual TTS. 99M-parameter ONNX model, ~400 MB download, 31 languages, 10 preset voices. No…
-
sanoTTS Offline
Offline neural text-to-speech at 294k-2.3M parameters — pure numpy, no torch, no GPU. 28 voices across 14 languages, each…
-
Parler-TTS Offline
Prompt-driven offline TTS. 34 named consistent speakers + free-text voice description. ~2 GB model, ~10-30s per line on CPU,…
-
Qwen3-TTS Clone
Multilingual neural TTS with 9 premium timbres + 3-second voice cloning. Heavy: ~3.5 GB model, 8+ GB RAM, CPU or CUDA.
-
Coqui XTTS Clone
Multilingual neural TTS with 6-second voice cloning. Heavy: ~2 GB model, CPU or CUDA.
-
F5-TTS Clone
Voice-clone neural TTS. Single 'f5-clone' voice; every request needs a 6-15s reference clip. ~1.5 GB model, en+zh only.