No set. No recapture. Same face, same mic. The vocal is already on the track.
P-Video-2 generates a whole new video from text, image, or audio-conditioned input. A jump in quality and efficiency for videos up to 20 seconds.
Standard from $0.025/s of output video · Draft mode from $0.015/s
What's the setup? 👀
- Model: P-Video-2
- Input image: Artist on stage with a microphone
- Input audio: Imported music track
- Input prompts: A young woman sings the song, natural mouth shapes, believable breathing, and facial expressions that reflect the emotion and intensity of the vocals. She maintains a confident stage presence while making subtle rhythmic movements.
👉 Free Playground:
📚 Documentation:
🧩 API:
显示更多