Make video dubbing understandable, inspectable, and reusable.
Project Film · Interactive Web App · Methods · Playable Examples · Documentation
Watch the OpenDub project introduction
4 min 37 s · 1920 x 1080 · Chinese / English subtitles · delivery details
OpenDub is a local-first research platform for multimodal video dubbing. It turns a difficult research task into a clear, interactive experience: explain the inputs, inspect complete methods developed by the team, listen to authorized archived examples, relate hearing to observable acoustic evidence, and prepare a rights-aware local project.
OpenDub does not splice internal modules from different papers into a new, unverified model. HPMDubbing, StyleDubber, and EmoDubber remain independent, complete methods. OpenDub makes their task assumptions, evidence, and usage boundaries visible in one place.
OpenDub presents the team's original work as complete methods with distinct priorities, rather than treating them as interchangeable fragments.
| Method | Complete-method focus | Upstream source |
|---|---|---|
| HPMDubbing | Hierarchical visual prosody: lip motion, facial affect, and scene context guide duration, pitch, energy, and emotion. | Repository · paper |
| StyleDubber | Multi-scale style learning: visual frames, phonemes, and utterance-level context support clear pronunciation and character style. | Repository |
| EmoDubber | Emotion-controllable movie dubbing: lip-related alignment, pronunciation, speaker identity, and emotion-guided generation. | Repository · paper |
| Speaker2Dubber | From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency Learning. | Repository · paper |
| InstructDubber | Instruction-based Alignment for Zero-shot Movie Dubbing. | Repository · paper |
| HiCoDiT | Hierarchical Codec Diffusion for Video-to-Speech Generation. | Repository · paper |
| CoSyncDiT | CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing. | Repository · paper |
Silent Video + Text + Authorized Reference Speech
│
▼
One complete dubbing method
│
▼
Target Dubbed Speech + Dubbed Video
Video dubbing is more than reading a sentence aloud. The video carries lip motion, facial expression, scene context, and timing; text defines the intended content; authorized reference speech supplies an identity and style condition. OpenDub exposes these signals as an interactive, time-aware task rather than a black-box audio button.
The following clips are authorized, team-provided historical research examples. Select a method name to open its MP4 in GitHub's video viewer; run the local web app to inspect the same assets with synchronized playback and acoustic features. These are not fresh OpenDub runs, common-input replay, or rankings.
| Human portrait case · 3.0 s | Animated character case · 1.36 s |
|---|---|
![]() |
![]() |
| Reference performance · HPMDubbing · StyleDubber · EmoDubber | Reference performance · HPMDubbing · StyleDubber · EmoDubber |
| Case record · authorization record | Case record · authorization record |
| Animated cinematic scene · 1.56 s | Presenter and display scene · 7.8 s |
|---|---|
![]() |
![]() |
| Reference performance · HPMDubbing · StyleDubber · EmoDubber | Reference performance · HPMDubbing · StyleDubber · EmoDubber |
| Case record · authorization record | Case record · authorization record |
Archived research example — not a fresh OpenDub run or a common-input ranking.
| Available now | Evidence-gated by design |
|---|---|
| Interactive task explanation, method canvases, original-paper component views, local Studio preparation, evidence records, and authorized archived examples. | Fresh model execution, numerical comparison, replay, and live generation require a verified method runtime, licensed weights, authorized inputs, and a real smoke test. |
This distinction is deliberate. It prevents mechanism illustrations or historical media from being misrepresented as a new inference result. See the project overview and model admission policy.
The interactive experience runs entirely on your machine.
pnpm install
pnpm web:devOpen http://127.0.0.1:5173 and visit Task, Methods, Examples,
Compare, Evidence, and Studio. The Studio/API workflow is also
available through the local compose stack:
docker compose up --buildFor the full quality gate:
make check- Public project overview
- Project film, subtitles, and checksum
- Documentation index
- Platform architecture
- Method admission status
- Example-media rights records
- Contribution guide
Use only video, text, and reference speech that you own or are authorized to process. Do not impersonate people, misrepresent generated media, or redistribute restricted source material. OpenDub is designed for local-first workflows and keeps evidence, input authorization, and runtime admission explicit.
New OpenDub platform code is released under Apache-2.0. Upstream methods, model weights, datasets, and example media remain subject to their own licenses and permission records. See NOTICE and CITATION.cff.








