MiniMax H3 image+audio to video
feed an audio into h3 instead of generating it. i was testing this flow, got shocked with how good the lipsync and movement is
implemented it on the i2v version in diffusers, made a demo for it. what do you think? ๐
โถ๏ธ