Register and share your invite link to earn from video plays and referrals.

Stav Zilbershtein
@mightyking
Building create winning ads with AI at scale
Joined April 2013
59 Following    1.1K Followers
Requiem for building in public Music video workflow + Prompts: One song, a simple story, singers across 4 locations, a full edit, captions burned in. Models used: Suno for the track, Seedance 2.5 for the story footage, MiniMax H3 for the lip synced singing, faster-whisper for word timing, Claude to edit, ffmpeg to burn captions. THE SONG (Suno) Short lyrics, fast beat, vocals on second 0. Long slow AI songs fall apart because the model has nothing to hide behind. Fire small batches, 2 clips at a time, and listen before firing more. Write every new attempt from zero. Stacking "less this, no that" onto the last try feeds the model your confusion and hands it back. One adjective moves everything. I put "soft" in a prompt once and the whole vocal switched to a woman. Remix prompt (paste into Suno style box): aggressive male rap, hard boom bap drums with fast energy, dark piano loop, deep male voice on every line including the hook, punchy mix, vocals start immediately at 0:00, no instrumental intro Lyrics: [Hook] It's just this thing I feel When I wanna steal It's just this thing I feel When I wanna steal [Verse 1] Yo, I see you on X, all over my feed You're building in public, I'm watching you build Your MRR chart looks like a hockey stick I screenshot it sometimes, that's normal right [Hook] It's just this thing I feel When I wanna steal [Verse 2] I learned a lot from you I think I deserve it too So I copied everything from you Same landing page, same pricing, same font And now you blocked me What happened bro I was your biggest fan [Hook] It's just this thing I feel When I wanna steal Tip: short lyrics, fast beat, vocals at 0:00, fresh prompt every round. THE STORY Think old MTV. The video is the movie this song is the soundtrack of. Keep the plot dead simple, something you can follow with the sound off. Mine: a broke founder copies a guy, dreams he is rich, wakes up, sees he got blocked, spits his cereal at the screen. That is all of it. Tip: if you cannot explain the story with zero words, cut it down. THE STORY FOOTAGE (Seedance 2.5) Seedance made the apartment story as one 30 second clip from reference images. Two things kill Seedance: Too many object interactions in one shot, and timestamps like "0 to 4 seconds" which it reads as a time lapse and speeds through. plain shots labelled "Shot 1, Shot 2" at natural speed will do the trick For the hard beat at the end, a guy waking up, eating cereal, seeing a screen, then spitting milk on the lens, text alone will not hold it. I built a 3 panel storyboard image and fed it as a reference. In the prompt you tag that image at the exact moment it happens, tell it the board reads left to right, describe it once, and move on. The rule that saved it: chronological order, tag each image where it belongs in time, say everything a single time, never repeat a thing. Repeat one detail twice and the model fixates on it and breaks the shot. Tip: for anything complex, hand it a storyboard picture and describe it once, in order. THE SINGING (MiniMax H3) H3 is the model that lip syncs to your actual track and keeps it. Seedance cannot, it regenerates its own audio. In H3 you attach your audio slice, set it to copy, and the mouth follows your real song. H3 caps around 15 seconds a clip and the song is 43. So I cut the song into 4 windows of about 11 seconds and generated a shot for each window. Then I did it across 4 locations, subway, warehouse, empty office, street. That is a 4 by 4 grid, 16 clips. I added 8 more where the whole crew sings and dances. Around 24 short singing clips to cover a 43 second song. You are building a bank of clips to cut from. Small H3 rules that matter: the audio slice must be a touch shorter than the clip, name every speaker, and compress the slice so there are no silent gaps for the model to fill with invented sound. Tip: chop the song into sub 15 second windows, shoot each shot per window, build a clip bank. THE EDIT (Claude) This is where most people lose hours. I made editing fast by doing the prep once. Every clip gets normalized to the same size and frame rate up front. After that each edit is a single ffmpeg pass with no re-encoding loops. The base layer is the song. Every clip's own audio is thrown out. To keep mouths in sync I gave the editor the math: each clip knows which second of the song its first frame belongs to, so to place it at song second S you trim it to start at S minus that offset. I also handed over word level timing from a whisper pass so cuts could land on real lyric moments. The rules I gave: nothing stays on screen too long, pace every cut to the lyric and the beat and what is on screen, never put two shots from the same location back to back, keep it heavy on story B-roll, never reuse a frame. If the cut feels like a metronome you failed. If it feels random you failed. Then the actual move. I did not ask for one perfect edit. I gave 7 agents the same rules and the same clip bank and told each to cut the whole thing its own way with a different emphasis. 6 came out flat. 1 landed around 90 percent. I finished that one by hand in CapCut. Tip: give strict rules plus the timing data, generate many full edits, keep the best and finish it yourself. CAPTIONS (ffmpeg) Burned straight from a styled subtitle file with ffmpeg. Seconds, not the long render a motion tool costs. The words come from the real lyrics, the timing comes from a whisper pass on the audio, and it highlights the word being sung. Big, thick, one pop color on the active word. Tip: real lyrics for the words, whisper for the timing, burn with ffmpeg. The AI did not make this video. I directed it, generated in volume, and kept the best takes. That is the whole game right now.
Show more