Whisper (offline)

Speech to textOfflineCPU onlyCommercial use OK

Turns the speech in your video into timed subtitles, on your own computer, with whisper.cpp. No cloud account, no graphics card, no Python to install. Pick a model once; after its download, nothing leaves your machine.

Install it from Subtitld

Settings Add-ons Whisper (offline) Install

Subtitld downloads it, checks its SHA-256, and keeps it updated.

Source on GitHub
Languages
99
Models
9
Add-on download
21–46 MB
RAM, at least
512 MB

What it does

One subtitle per sentence

Segments appear on the timeline while it decodes, then are regrouped so each subtitle holds one sentence instead of a fragment.

All, a range, or the selection

Transcribe the whole video, a stretch of it, or only the subtitles you selected.

Ready for Record mode

The model stays loaded between phrases, so only the first one waits for it to load.

Cancel any time

Stops within about a second, even mid-decode, and never leaves half a result behind.

Upgrading from 26.03? This is the engine that used to be built in. It keeps the same id, settings and model files, so your setup and downloaded models carry over without a new download.

Models

Bigger models make fewer mistakes and take longer. Base is the default and a good start; try Small or Medium for names, accents and noisy audio. The .en versions understand only English and are a little faster.

Model DownloadRAMGood for
Tiny tiny · tiny.en
75 MB
~0.3 GBQuick drafts, slow computers
Base Default base · base.en
142 MB
~0.4 GBClear speech, most videos
Small small · small.en
466 MB
~0.9 GBAccents, names, interviews
Medium medium · medium.en
1.5 GB
~2.1 GBNoisy audio, several languages
Large v3 large-v3
2.9 GB
~3.9 GBThe best result, when time allows

Models come from ggerganov/whisper.cpp on Hugging Face, once, the first time you use them.

Where models are kept

In Subtitld's shared model cache, never inside the add-on, so updates don't delete them. Uninstalling the add-on keeps them too; delete the folder yourself to free the space.

System Folder
Linux ~/.cache/subtitld/models/whispercpp
macOS ~/Library/Caches/subtitld/models/whispercpp
Windows %LOCALAPPDATA%\subtitld\subtitld\Cache\models\whispercpp

Options

In Settings › Add-ons › Whisper (offline) › Configure.

Model
Default: base

Which model to use. See the table above.

CPU threads
Default: 0

Auto uses up to 4 threads. Raise it on a computer with more cores to go faster.

Group output by phrases
Default: on

One subtitle per sentence. Turn it off to keep Whisper's own segments, which often break mid-sentence.

Models folder
Default: empty

Empty shares Subtitld's model cache. Point it elsewhere to keep models on another drive.

Languages

99 languages, or leave the language empty to detect it. Regional tags such as pt-br use their main language; every Chinese variant maps to zh. Of Subtitld's languages, only Zulu isn't supported.

  • af
  • am
  • ar
  • as
  • az
  • ba
  • be
  • bg
  • bn
  • bo
  • br
  • bs
  • ca
  • cs
  • cy
  • da
  • de
  • el
  • en
  • es
  • et
  • eu
  • fa
  • fi
  • fo
  • fr
  • gl
  • gu
  • ha
  • haw
  • he
  • hi
  • hr
  • ht
  • hu
  • hy
  • id
  • is
  • it
  • ja
  • jw
  • ka
  • kk
  • km
  • kn
  • ko
  • la
  • lb
  • ln
  • lo
  • lt
  • lv
  • mg
  • mi
  • mk
  • ml
  • mn
  • mr
  • ms
  • mt
  • my
  • ne
  • nl
  • nn
  • no
  • oc
  • pa
  • pl
  • ps
  • pt
  • ro
  • ru
  • sa
  • sd
  • si
  • sk
  • sl
  • sn
  • so
  • sq
  • sr
  • su
  • sv
  • sw
  • ta
  • te
  • tg
  • th
  • tk
  • tl
  • tr
  • tt
  • uk
  • ur
  • uz
  • vi
  • yi
  • yo
  • zh

Platforms

System DownloadTested
Linux x86-64 46.3 MBOn real machines and in CI
macOS (Apple silicon) 21.4 MBIn CI
Windows x86-64 21.7 MBIn CI

Before every release, each build runs a real transcription in CI, installed the way Subtitld installs it.

Troubleshooting

  • model_missingThe model file is missing or damaged. The message names the file; delete it and run again to download it fresh.
  • network_unavailableThe model couldn't download. Try again; behind a proxy that inspects HTTPS, set SSL_CERT_FILE to its certificate.
  • disk_fullNot enough space for the model. Free some, or move the models folder.
  • oomOut of memory. Pick a smaller model.
  • unsupported_languageWhisper doesn't know this language. Try Vosk, or auto-detect.

License

Add-on
Apache-2.0
Engine
MIT
whisper.cpp, pywhispercpp
Models
MIT
OpenAI Whisper weights in GGML

Fine for commercial work. Some other add-ons use models for non-commercial use only; each page says so at the top.

Versions

  • 0.0.1

    Offline speech-to-text for Subtitld with whisper.cpp (pywhispercpp 1.5.1).

    Models are not included: the selected GGML model downloads once on first use into Subtitld's shared model cache, so models downloaded by earlier Subtitld versions are reused.

    The add-on is Apache-2.0. Each archive contains THIRD_PARTY_LICENSES/ with the licences of the third-party components frozen into it: whisper.cpp and pywhispercpp are MIT, pybind11 and numpy are BSD-3-Clause, and the Python runtime and its libraries are under their own permissive licences.

    The Linux archive also ships, inside numpy's OpenBLAS, the GCC Fortran runtime (GPL-3.0 with the GCC Runtime Library Exception) and libquadmath (LGPL-2.1-or-later). The complete source of that libquadmath build is attached here as whispercpp-addon-<version>-libquadmath-source.tar.gz; the add-on's own source is whispercpp-addon-<version>-source.tar.gz.

Other Speech to text add-ons