this repo beat Runway Gen-3 and Luma 1.6 in blind human evaluation. it also needs 60GB of VRAM for one video.
- text alignment: 61.8%, best of the 6
- motion quality: 66.5%, best of the 6
- overall preference score: 41.3%, ranked #
1#
a 13B parameter model, largest open-source video model at release, using a decoder-only LLM as the text encoder instead of CLIP+T5.
this is the original HunyuanVideo though.
Tencent's own HunyuanVideo-1.5 replaced it since, same lineage, far lighter to run.