MuseTalk (GPU lip-sync)

Lip-syncOfflineNeeds a GPUCommercial use OK

High-quality generative lip-sync (visual dubbing) via MuseTalk. Requires an NVIDIA GPU (CUDA) and Python 3.10+. Small download: the PyTorch runtime and multi-GB models install on first use. MIT-licensed and commercial-OK.

Install it from Subtitld

Settings Add-ons MuseTalk (GPU lip-sync) Install

Subtitld downloads it, checks its SHA-256, and keeps it updated.

Source on GitHub
Add-on download
9–18 MB

Options

In Settings › Add-ons › MuseTalk (GPU lip-sync) › Configure.

MuseTalk version
Default: v15

MuseTalk model version. 1.5 gives sharper mouths; 1.0 is the original.

Bbox shift
Default: 0

Shifts the face crop vertically to tune mouth openness. Positive opens the mouth more, negative closes it. MuseTalk 1.0 only.

Extra margin
Default: 10

Extra pixels added below the face crop (jaw region). MuseTalk 1.5 only.

Batch size
Default: 8

Frames generated per UNet forward pass. Higher is faster but needs more VRAM; lower it on 'oom' errors.

Output FPS
Default: 25

Frame rate of the generated video. MuseTalk is trained at 25 fps; leave at 25 for best results.

Use float16
Default: on

Run the UNet/VAE in fp16 to cut VRAM use. Disable if you see numerical artifacts.

Platforms

System Download
Linux x86-64 17.7 MB
Windows x86-64 8.6 MB

License

Add-on
MIT

MIT

Versions

  • 0.2.0

Other Lip-sync add-ons