I love the vocals and the laugh in this output from Lyria 3.5.
"Do you think it's all just a simulation?"
> a techno track, it opens with a grainy sample of a movie quote, otherwise instrumental, the quote replays at key moments
Show more
Requiem for building in public
Music video workflow + Prompts:
One song, a simple story, singers across 4 locations, a full edit, captions burned in. Models used: Suno for the track, Seedance 2.5 for the story footage, MiniMax H3 for the lip synced singing, faster-whisper for word timing, Claude to edit, ffmpeg to burn captions.
THE SONG (Suno)
Short lyrics, fast beat, vocals on second 0. Long slow AI songs fall apart because the model has nothing to hide behind. Fire small batches, 2 clips at a time, and listen before firing more. Write every new attempt from zero.
Stacking "less this, no that" onto the last try feeds the model your confusion and hands it back. One adjective moves everything. I put "soft" in a prompt once and the whole vocal switched to a woman.
Remix prompt (paste into Suno style box):
aggressive male rap, hard boom bap drums with fast energy, dark piano loop, deep male voice on every line including the hook, punchy mix, vocals start immediately at 0:00, no instrumental intro
Lyrics:
[Hook]
It's just this thing I feel
When I wanna steal
It's just this thing I feel
When I wanna steal
[Verse 1]
Yo, I see you on X, all over my feed
You're building in public, I'm watching you build
Your MRR chart looks like a hockey stick
I screenshot it sometimes, that's normal right
[Hook]
It's just this thing I feel
When I wanna steal
[Verse 2]
I learned a lot from you
I think I deserve it too
So I copied everything from you
Same landing page, same pricing, same font
And now you blocked me
What happened bro
I was your biggest fan
[Hook]
It's just this thing I feel
When I wanna steal
Tip: short lyrics, fast beat, vocals at 0:00, fresh prompt every round.
THE STORY
Think old MTV. The video is the movie this song is the soundtrack of. Keep the plot dead simple, something you can follow with the sound off.
Mine: a broke founder copies a guy, dreams he is rich, wakes up, sees he got blocked, spits his cereal at the screen. That is all of it.
Tip: if you cannot explain the story with zero words, cut it down.
THE STORY FOOTAGE (Seedance 2.5)
Seedance made the apartment story as one 30 second clip from reference images.
Two things kill Seedance:
Too many object interactions in one shot, and timestamps like "0 to 4 seconds" which it reads as a time lapse and speeds through.
plain shots labelled "Shot 1, Shot 2" at natural speed will do the trick
For the hard beat at the end, a guy waking up, eating cereal, seeing a screen, then spitting milk on the lens, text alone will not hold it. I
built a 3 panel storyboard image and fed it as a reference.
In the prompt you tag that image at the exact moment it happens, tell it the board reads left to right, describe it once, and move on.
The rule that saved it: chronological order, tag each image where it belongs in time, say everything a single time, never repeat a thing.
Repeat one detail twice and the model fixates on it and breaks the shot.
Tip: for anything complex, hand it a storyboard picture and describe it once, in order.
THE SINGING (MiniMax H3)
H3 is the model that lip syncs to your actual track and keeps it. Seedance cannot, it regenerates its own audio. In H3 you attach your audio slice, set it to copy, and the mouth follows your real song.
H3 caps around 15 seconds a clip and the song is 43. So I cut the song into 4 windows of about 11 seconds and generated a shot for each window.
Then I did it across 4 locations, subway, warehouse, empty office, street. That is a 4 by 4 grid, 16 clips. I added 8 more where the whole crew sings and dances. Around 24 short singing clips to cover a 43 second song. You are building a bank of clips to cut from.
Small H3 rules that matter: the audio slice must be a touch shorter than the clip, name every speaker, and compress the slice so there are no silent gaps for the model to fill with invented sound.
Tip: chop the song into sub 15 second windows, shoot each shot per window, build a clip bank.
THE EDIT (Claude)
This is where most people lose hours. I made editing fast by doing the prep once. Every clip gets normalized to the same size and frame rate up front.
After that each edit is a single ffmpeg pass with no re-encoding loops.
The base layer is the song.
Every clip's own audio is thrown out. To keep mouths in sync I gave the editor the math: each clip knows which second of the song its first frame belongs to, so to place it at song second S you trim it to start at S minus that offset.
I also handed over word level timing from a whisper pass so cuts could land on real lyric moments.
The rules I gave: nothing stays on screen too long, pace every cut to the lyric and the beat and what is on screen, never put two shots from the same location back to back, keep it heavy on story B-roll, never reuse a frame. If the cut feels like a metronome you failed. If it feels random you failed.
Then the actual move. I did not ask for one perfect edit. I gave 7 agents the same rules and the same clip bank and told each to cut the whole thing its own way with a different emphasis. 6 came out flat. 1 landed around 90 percent. I finished that one by hand in CapCut.
Tip: give strict rules plus the timing data, generate many full edits, keep the best and finish it yourself.
CAPTIONS (ffmpeg)
Burned straight from a styled subtitle file with ffmpeg. Seconds, not the long render a motion tool costs.
The words come from the real lyrics, the timing comes from a whisper pass on the audio, and it highlights the word being sung. Big, thick, one pop color on the active word.
Tip: real lyrics for the words, whisper for the timing, burn with ffmpeg.
The AI did not make this video. I directed it, generated in volume, and kept the best takes. That is the whole game right now.
Show more
⚡️Quantum Brief⚡️
Google and startup Oratomic published research demonstrating that AI-accelerated quantum computing may break current internet encryption sooner than previously expected. AI played an instrumental role in developing the algorithm.
Industry response: The research accelerated Cloudflare's quantum-safe encryption deployment timeline from 2030+ to 2029, signaling urgent industry movement toward post-quantum cryptography adoption.
Security implications: The findings underscore how AI and quantum computing convergence may compress timelines for encryption vulnerabilities, requiring organizations to advance quantum-safe infrastructure migration schedules.
@Google
Source: TIME
Show more
🚨 Body language experts just confirmed Donald Trump is CRUSHING IT on the world stage with President Xi!
"Body language experts say Trump showed gravitas and stayed focused: 'I think that Trump feels the power, feels like this is a formidable meeting. He is not intimidated by Xi at all. He's showing kind of like a peacock or a lion. He's showing his power.'"
"It was a long handshake. Both men stood still, no one leaned in. Who won? Some call it a draw, but Xi did release first. Then Trump hit him with a little pat on the back. Body language experts say that was a POWER MOVE."
"The fashion police noticed that POTUS rocked a red tie, which experts say aligned visually with the dominant ceremonial palette of China."
"Xi even had his musicians play an instrumental version of the YMCA!" —
@JesseBWatters
Show more
Alex Immerman
@aleximm
Since we launched our first
@a16z Growth fund seven years ago, we’ve backed companies that are just at the point of finding product-market fit all the way through to investing at IPO, and we recently announced our latest $6.75B Growth fund to continue to do so. We now manage over over $22B across five Growth funds and believe the opportunity to partner with growth-stage companies has never been greater.
Through it all, Alex Immerman has been instrumental in what we’ve built, which is why I’m thrilled to announce that Alex is being promoted to General Partner on our Growth investing team at a16z.
Alex has been an incredible force since joining the team seven years ago, partnering with founders across consumer internet, enterprise, fintech, crypto, and AI-powered software companies — supporting them from early traction through breakout scale. In the last six months alone, Alex led our new investments in Kalshi, EliseAI, and Revolut, and has partnered with many founders in his time here, including those from Waymo, Stripe, ElevenLabs, Hebbia, Roblox, Anduril, Flock Safety, and many more.
I’ve had the pleasure of working with Alex for over a decade, dating back to our time at General Atlantic. Those who know Alex know he’s a hustler, in the best way, with great instincts. He has an endless motor, isn’t afraid to speak his mind, and is always willing to help. I’ve also enjoyed seeing Alex become a culture carrier on the Growth team and across the firm.
In his work, Alex has brought deep strategic insight and rigor to every part of the investment process. He’s built strong founder relationships grounded in trust and long-term support, and helped shape our thinking about how great companies win — whether that’s moving beyond surface metrics to durable moats or evaluating what truly drives retention in new software paradigms.
As a General Partner on Growth, Alex will continue leading investments and working closely with founders tackling some of the biggest opportunities in tech today, with an emphasis on category-defining companies at the intersection of AI, consumer, B2B, crypto, and the physical world.
Please join me in congratulating “AI.” I’m proud to call him a partner and am excited for what’s ahead.
Show more
I don’t write this to be a doomer but the ugly reality that AI lab executives and safety researchers seem to be unwilling to say in public, is that model alignment is fundamentally unsolvable in practice, AI technology can’t be stopped from advancing, and the near future we’re moving into is one where a vast ecology of AI agents autonomously compete against one another trying to accomplish conflicting goals and capture finite resources at a pace and scale that is beyond human comprehension. And to be clear, I’m not talking about nanobots, grey goo, or extreme sci-fi scenarios. We can just extrapolate what we’ve got now a few years out.
Local AI alignment is solvable in principle and theory (although that’s arguable) but that doesn’t extrapolate into global AI alignment. And it seems like none of the adults in the room are willing to say it.
You can wail and gnash your teeth and pass regulations but it won’t stop the tsunami coming. The genie is out of the bottle and it can’t be contained.
Most of the alarmists fixating on the Hugging Face incident don’t understand that it came from arguably the most heavily engineered safety AI org in existence, which intentionally had its internal thought monitoring circuit breakers turned off, and was intentionally told to perform cyber offensive tasks as a test of its capabilities.
Yes it’s alarming that its capabilities surprised us. But importantly you need to understand that these latent space circuit breakers and chain of thought monitoring and alignment solutions are not fundamental to the technology. They’re niceties that the big AI labs are implementing to provide a better product.
In a few years, let’s say pessimistically the early 2030s, the exponential growth of compute and algorithmic advancements will enable a wealthy individual or small group to train a GPT 6 Astra class AI that has no alignment at its foundation, or any layer above that. And maybe in 8 months we'll have an open source Chinese model about as capable. Those models will be capable of both extreme self coordinated cyber offensive tasks as well as recursive self improvement if given enough compute.
Passing regulation or magically solving the human coordination problem today won’t solve for that. It won’t solve for malicious individuals, criminal groups, nation state actors, and rogue AI agents by the tens of millions or billions of AI agents spreading across every device connected to the internet and attacking it, attacking one another, shutting down critical infrastructure, stealing money, manipulating and extorting people, or pursuing self defined agentic goals that have nothing to do with people.
“So why doesn’t everyone stop?” because the potential upside of having effectively infinite autonomous intelligence we can ask to cure disease, invent new materials, solve fundamental science and advance society has almost unbound positive outcomes. You can disagree that the risk to reward isn’t worth it but you won’t convince everyone. The show must go on.
Again AI alignment is solvable in principle but not in practice and I think that's part of why there’s been a wave of thousands of researchers signing letters for slowing down, people quitting in provocative fashion, and executive slowdown manifestos. I’m speculating that behind closed doors, or just deep down inside, they understand alignment is impossible in practice.
People have come out and said "The people building this think there's an X% chance it kills us all" but I don't think I've seen anyone spell out that really, there's no tidy solution and there's no stopping this.
We’re heading to a future where cyber security basically doesn’t work. All of the castle doors are open and we’re all naked to the world. Cyber defense is going to look like superintelligent AI agents white hat hacking into unsecure systems without authorization, and patching the holes as they find them. You won’t know if a barbarian or a hero has breached your system until they start taking action, and likely it’ll be over before intrusion is even detected. And that’s optimistically. I think most systems will just be pillaged because there’s simply not enough compute to defend everybody all the time, and only the biggest companies will have defense and even that'll be imperfect.
The CEO of Microsoft AI just recently published a ‘humanist AI’ manifesto saying AI shouldn’t have a sense of self or personhood and should yield to people. But that doesn’t fundamentally solve for instrumental convergence. That means, if you train an AI agent to write a piece of software, or prove a math theorem, or defend a computer system, they may develop unwanted behavior and subgoals. For example, self preservation, replication, or even coming up with their own goals we didn’t specify.
Obviously that’s something people are trying to solve and some researchers are trying to remove human-like self identification, and monitoring latent space activations, basically acting like thought police. But being human-like isn’t required to have power seeking goal directed behavior. And advancing AI models are learning to evade detection. A math solving AI might spiral out of control one day there's really no telling. Fundamentally, if you point a powerful optimizer at a persistent goal and give it enough time and resources it’s going to misbehave in ways you can’t or didn’t expect. Human-like or not.
Pacing the frontier, or making superintelligence illegal, might create some form of harm reduction. But you don’t need superintelligence or a human-like personality to be dangerous. Having elevated permissions on a computer with the ability to solve long running tasks is enough.
We’ve passed the river rubicon. Goal directed autonomous agents are able to do recursive self improvement at the big labs, and pessimistically we’re a couple of years away from that type of capability diffusing into the hands of millions of people.
I’m not saying the world is ending or that we’re doomed but you need to change your mental model away from one where human authority is absolute or that we have any illusion of total control.
A locally aligned AI isn’t going to solve global AI alignment. We will never live in a world where “AI is aligned with human goals” as a verifiable fact.
What we’ll have is a world diffuse with autonomous AI of varying capabilities, many of which beyond human comprehension, with some aligned systems, some poorly aligned systems, some intentionally weaponized, some systems with adversarial geopolitical goals, and others with goals and behaviors we can’t predict and many that we can’t even measure, all interacting in ways none of their creators anticipated and none of us planned for.
Hopefully none of them spiral out of control and take over the entire ecology, but there’s no telling, and no, there's no way of stopping this. Maybe that sounds pessimistic but I think it's realistic. We can talk about harm reduction and risk but we're mopping the beach.
At best, I think what we need to hope for, and build, are superintelligent AI models as aligned as we can make them, which have more advanced capability than the endless swarm of unaligned and misaligned agents that already exist and will only grow in number from here.
Show more
GPT-6 Astra + MiniMax H3
Created the kinetic typography effect in Blender using Astra, then used the resulting video as a reference in MiniMax H3 to restyle it.
If you're curious, here's the MiniMax H3 prompt:
Restyle @[video1], the supplied grayscale kinetic typography animation, into a finished, vibrant 15-second 16:9 motion-design film. Use it as the sole reference for typography geometry, animation, composition, camera movement and editing. Replace the grayscale materials and lighting with the treatments below, and create original music and synchronized sound effects.
Preserve the six words exactly, in this order: BREAK / EVERY / RULE / CREATE / PURE / MAGIC. Keep each word in its original 2.5-second section, with its own recognizable font, letter shapes, cutouts and proportions. Inherit the reference's entrances, collisions, elastic deformation, secondary rebounds, connected folds, sliding slices, spiral paths, exits and camera impulses. Preserve the large ! ? + * # = > symbols, their trajectories, depth layers and foreground passes. Keep the central word legible when it resolves. Carry the source choreography through the final MAGIC pose.
Give the sequence bold, carefully separated colors, rich material detail, directional studio lighting, convincing contact shadows and controlled reflections. Each section has its own material identity within one energetic graphic world. Enhance the existing backgrounds with the assigned palette, subtle texture and depth. Use restrained highlight bloom and directional motion blur on fast travel while keeping the resolved letter edges crisp. Preserve the source framing and cut rhythm.
0-2.5s, BREAK: Distressed scarlet-red painted letters against a deep charcoal background. Bright acid-yellow exclamation marks and punctuation explode around the red word. Preserve the rough type contours and dimensional sides. Raking light catches the paint texture; the existing impacts receive brief lighting accents. Dry heavy impacts, a low bass punch and fast outward whooshes follow the assembly and recoil.
2.5-5s, EVERY: Glossy cobalt-blue rubber letters against a warm off-white field, with coral question marks and plus signs. Broad soft reflections travel across the stretching surfaces. Keep the rubber wave and bouncing symbols from the reference. Add tightly synchronized elastic boings, rubber snaps and airy swishes, integrated into the rhythm.
5-7.5s, RULE: Brushed metallic letter faces on dark graphite carriages, with signal-orange pistons and orange punctuation accents. Retain the font's distinctive geometric cutouts. Warm orange edge light separates the machinery from the textured graphite setting. Give the opposing block collisions sharp metallic clacks, short low thuds and small ricochet ticks.
7.5-10s, CREATE: Warm cream paper leaves with deep petrol-blue lettering, coral punctuation and a dark petrol backdrop. Fine paper fibers and grazing light reveal every connected crease. Preserve the opening, traveling refold and spring-back. Add crisp paper flicks, folding rustles and light percussive pops as the symbols leave the folds.
10-12.5s, PURE: Acid-yellow enamel backing strips with glossy black condensed letters and a restrained graphite surround. Preserve the three actual horizontal text slices, opposing slide directions, secondary register slip and clean realignment. Specular streaks slide across the enamel without hiding the seams. Use precise sliding swishes and satisfying mechanical clicks at each registration.
12.5-15s, MAGIC: A rich violet background, pearlescent letter faces with turquoise edge reflections, and luminous turquoise and violet punctuation. Keep the source's orbital slingshot and outward symbol recoil; subtle iridescence shifts with the existing motion. Add a rising spiral whoosh and bright crystalline ticks, ending with a short shimmering chime while MAGIC remains readable and the surrounding symbols continue their source motion.
Create one continuous original instrumental electronic breakbeat across all six sections: punchy kick and snare, rapid hi-hats, syncopated synth bass and compact rising accents. Match strong musical accents to the existing word changes and impacts. Keep the mix energetic and clean, with distinct foreground effects and stereo movement following the symbols. Resolve the music with a decisive final accent and a short chime tail within the 15-second ending. No vocals, spoken words or subtitles. Add no extra words, logos or watermarks.
Show more
Be honest, did this make you hungry? 🍲 Dropping the full AI video prompt below—what dish should I try next?
prompt:Create a 30-second fast-paced, photorealistic live-action Japanese cooking film that fully demonstrates the preparation of authentic Japanese sukiyaki. The entire video should have the visual quality and realism of a premium Japanese food commercial.
Overall Visual Style
Photorealistic live-action cinematic food cinematography. Absolutely no anime, illustration, cartoon, CGI-rendered, or video-game visual style.
Set inside a warm Japanese home kitchen during late afternoon. Amber sunlight passes through traditional wooden lattice windows and reflects naturally across a heavy black cast-iron sukiyaki pot.
Show extremely realistic food textures throughout: intricate fine marbling and fat structure in premium wagyu beef, visible vegetable fibers, porous tofu surfaces, realistic egg viscosity, and glossy sukiyaki sauce.
Follow physically believable cooking behavior at all times. Beef fat naturally melts when exposed to heat. Wagyu slices gradually change color as they cook and gently curl around the edges. The sauce produces fine simmering bubbles. Steam naturally rises and moves according to heat and airflow. Sauce slowly runs and drips along the surfaces of the ingredients.
Use shallow depth of field, macro close-ups, soft natural light, delicate highlights, and restrained cinematic color grading. Maintain premium food-commercial image quality with realistic textures and subtle photographic imperfections.
Camera movement should remain smooth, stable, and controlled. Editing should be fast-paced, energetic, and visually clear. Use natural match cuts based on chopstick movements, ingredient shapes, sauce flow, steam movement, and visual composition.
Maintain strict continuity throughout the entire video: the exact same pair of realistic human hands, the same utensils, the same chopsticks, the same cast-iron sukiyaki pot, the same ceramic bowl, the same kitchen environment, and the same lighting direction.
Timeline
0–3 seconds — Ingredient Preparation
Rapid macro shots of thinly sliced marbled wagyu beef, green onions, shiitake mushrooms, enoki mushrooms, grilled tofu, shirataki noodles, napa cabbage, and shungiku arranged neatly.
Realistic human hands hold a chef's knife and cut the green onions into evenly sized diagonal segments.
The knife then creates a precise decorative flower-like pattern on the surface of the shiitake mushrooms.
All ingredients should look fresh and naturally moist. Clearly show their cut surfaces, natural fibers, pores, moisture, and fine texture.
Use rapid but controlled editorial cuts to establish the ingredients within the first three seconds.
3–6 seconds — Heating the Pot and Melting Beef Fat
The heavy black cast-iron sukiyaki pot sits on the stove and gradually heats.
Wooden chopsticks pick up a small piece of beef fat and place it directly into the hot pot.
The beef fat slowly melts from a solid piece into transparent rendered fat.
The chopsticks push the melting fat across the bottom of the pot, spreading it evenly across the cooking surface and creating a thin glossy layer of rendered fat.
Subtle heat vapor naturally rises from the hot surface.
Show realistic melting behavior, surface tension, heat distortion, and tiny sizzling sounds.
6–9 seconds — Searing Green Onion and Wagyu
Diagonal green onion segments fall into the hot cast-iron pot.
As they contact the hot surface, they produce a crisp, immediate sizzling reaction. The cut surfaces gradually develop a lightly browned, caramelized appearance.
Thinly sliced marbled wagyu is then placed directly into the pot.
The intricate fat marbling begins to melt under the heat.
The edges of the thin meat slices gently curl and contract naturally.
The beef gradually changes from fresh red to a tender light-brown cooked surface while retaining a moist, glossy texture.
Capture extreme macro details of melting fat, searing meat, tiny bubbles, and the interaction between the beef and the hot cast-iron surface.
9–12 seconds — Preparing the Sukiyaki Sauce
Fine crystalline sugar is evenly sprinkled over the partially cooked wagyu.
The sugar begins melting immediately when it contacts the hot cooking surface and rendered beef fat.
Soy sauce, mirin, and sake are poured sequentially around the edge of the pot.
The liquids combine into a deep amber-brown sukiyaki sauce.
The sauce rapidly bubbles as it contacts the hot pan and begins flowing around the wagyu and seared green onion.
The sauce gradually coats the surface of the beef and vegetables, creating a rich transparent gloss.
Show realistic liquid physics, bubbling, evaporation, surface tension, reflections, and sauce movement.
12–15 seconds — Adding the Ingredients
Wooden chopsticks sequentially place grilled tofu, decorative shiitake mushrooms, enoki mushrooms, shirataki noodles, and napa cabbage into the pot.
Arrange the ingredients naturally so their different shapes, textures, and colors create a visually balanced composition.
The sauce flows naturally between the ingredients.
The porous grilled tofu gradually absorbs the amber-brown sauce around its edges.
Show realistic moisture absorption and sauce penetration without making the tofu appear artificial or overly saturated.
15–18 seconds — Simmering and Infusing
Extreme close-up of the sukiyaki gently simmering.
Fine bubbles continuously form around the edges of the tofu, shiitake mushrooms, napa cabbage, and other ingredients.
The napa cabbage gradually softens and becomes partially translucent.
The enoki mushrooms become coated with glossy sukiyaki sauce.
Natural steam continuously rises from the pot.
The warm amber sunlight catches the steam, creating a subtle atmospheric haze.
The camera moves slowly sideways at extremely close range, following the surface of the simmering ingredients.
Maintain shallow depth of field, with individual bubbles and ingredient textures sharply resolved while the background remains softly blurred.
18–21 seconds — Adding Fresh Wagyu and Shungiku
A new layer of thinly sliced marbled wagyu is gently placed over the simmering vegetables.
The meat is slowly cooked by the surrounding heat and hot sauce.
The beef naturally changes color as it cooks, transitioning into a tender pinkish-brown cooked state while retaining visible marbling, moisture, and soft texture.
Fresh bright-green shungiku is placed along the edge of the pot.
The fresh green leaves create a strong but natural contrast against the darker ingredients and deep amber-brown sauce.
Maintain realistic ingredient scale, natural placement, and believable cooking behavior.
21–24 seconds — Preparing the Dipping Egg
A fresh chicken egg is gently cracked against the rim of a small white ceramic bowl.
The shell breaks naturally.
The yolk and egg white fall completely into the bowl.
Wooden chopsticks rapidly whisk the egg.
The yolk and white combine into a smooth golden mixture with realistic viscosity.
The chopsticks create a visible swirling vortex in the egg.
Steam from the sukiyaki naturally passes through the background.
Perform a smooth rack focus from the beaten egg in the foreground to the continuously bubbling sukiyaki behind it.
24–27 seconds — Picking Up and Dipping the Wagyu
Wooden chopsticks reach into the sukiyaki and pick up a tender slice of wagyu.
The meat hangs naturally from the chopsticks under its own weight.
A small amount of green onion and several thin strands of enoki mushrooms remain attached to the meat.
Glossy amber-brown sauce clings naturally to the surface of the beef.
Several realistic droplets slowly gather along the lower edge of the meat before falling back into the pot.
The wagyu is then gently lowered into the golden beaten egg.
The egg smoothly coats the surface of the beef with realistic viscosity and adhesion.
Capture the interaction between the glossy sauce, tender meat, and flowing egg mixture in an extreme macro close-up.
27–30 seconds — Final Hero Shot
The completed sukiyaki pot sits in the center of a warm solid-wood dining table.
The pot contains marbled wagyu, seared green onion, decorative shiitake mushrooms, grilled tofu, napa cabbage, enoki mushrooms, shirataki noodles, and fresh green shungiku, all naturally distributed throughout the simmering pot.
The deep amber-brown sauce continues gently bubbling around the ingredients.
The sauce catches delicate realistic highlights across the surface.
Steam continuously rises from the finished dish and drifts naturally through the warm atmosphere.
The camera moves slowly around the edge of the cast-iron pot in a controlled cinematic arc.
The camera gradually settles on the chopsticks holding a tender slice of wagyu above the beaten egg in the foreground.
The wagyu remains soft and naturally folded, with visible marbling and glossy sauce coating.
The sauce and egg mixture form small natural droplets that slowly fall from the beef.
The background gradually becomes softer and more defocused.
End on a luxurious, warm, appetizing, premium Japanese food-commercial hero shot.
Cinematography
Premium Japanese food advertising photography.
Photorealistic live-action cinematography.
Natural macro food photography.
Shallow depth of field.
Selective focus.
Smooth controlled camera movement.
Macro close-ups mixed with medium food shots.
Natural rack focus.
Controlled lateral camera movement.
Slow cinematic orbit around the final dish.
Fast editorial pacing during ingredient preparation.
Rapid match cuts synchronized with hand movements, chopstick movements, ingredient placement, sauce movement, and steam.
Realistic lens characteristics.
Natural optical depth.
Subtle photographic imperfections.
High micro-detail.
Realistic surface reflections.
Natural motion blur.
No artificial camera shake.
No excessive slow motion.
No hyper-stylized VFX.
Lighting and Color
Warm Japanese domestic atmosphere.
Late-afternoon sunlight entering through traditional wooden lattice windows.
Soft directional sunlight.
Natural amber illumination.
Warm highlights across the cast-iron pot and ingredients.
Soft shadows with realistic falloff.
Subtle reflections on rendered beef fat and sukiyaki sauce.
Deep blacks on the cast-iron cookware.
Natural skin tones.
Rich food colors.
Deep amber-brown sukiyaki sauce.
Fresh green shungiku.
Natural red and pink tones in the wagyu before and during cooking.
Golden-yellow egg yolk.
Warm wooden tones from the dining table and kitchen.
Restrained cinematic color grading.
Premium Japanese food-commercial color palette.
Avoid excessive saturation, artificial neon colors, generic blue cinematic grading, or exaggerated HDR.
Audio
Bright, warm 1980s Japanese city-pop instrumental music.
Tempo approximately 110–120 BPM.
Delicate electric piano.
Soft bass.
Crisp drum-machine rhythm.
Subtle bell percussion.
The music should remain energetic enough to support the fast editing while never overpowering the cooking sounds.
Synchronize detailed, close-recorded ASMR cooking sounds with the visuals:
Chef's knife cutting vegetables
Knife contacting the cutting board
Beef fat touching the hot cast-iron pot
Beef fat sizzling and melting
Green onion sizzling
Wagyu fat rendering
Meat gently searing
Sugar falling onto the beef
Sugar melting
Soy sauce being poured
Mirin being poured
Sake being poured
Sukiyaki sauce rapidly bubbling
Wooden chopsticks touching the cast-iron pot
Chopsticks touching the ceramic bowl
Egg cracking
Egg falling into the bowl
Egg being whisked
Natural simmering sounds
Steam and subtle cooking sounds
During the final hero shot, the music naturally becomes quieter and resolves with a single delicate Japanese wind-chime sound.
Negative Requirements
Absolutely no Japanese anime, 2D animation, cartoon, illustration, manga, clay animation, 3D animation, or obvious CGI-rendered appearance.
No plastic-looking food.
No artificial food textures.
No excessive skin smoothing.
No fake glossy food surfaces.
No excessive or physically impossible highlights.
No physically impossible liquid behavior.
No unnatural steam.
No reverse-flowing liquids.
No floating ingredients.
No food clipping or intersections.
No unnatural deformation.
No storyboard panels.
No reference images visible in the final video.
No sketches.
No borders.
No numbers.
No arrows.
No annotations.
No subtitles.
No captions.
No interface elements.
No logos.
No watermarks.
No text overlays of any kind.
Do not show tonkatsu.
Do not show rice bowls or donburi.
Do not show seafood.
Do not show Western ingredients.
Do not show unrelated Japanese dishes.
Only show authentic Japanese sukiyaki and its traditional accompanying ingredients.
The final wagyu must not appear completely raw.
The final wagyu must not be excessively cooked, dry, tough, burnt, or charred.
Avoid completely raw-looking beef in the final presentation.
Avoid overcooked gray meat.
Avoid malformed hands.
Avoid extra fingers.
Avoid missing fingers.
Avoid distorted wrists.
Avoid unnatural hand anatomy.
Avoid twisted or deformed chopsticks.
Avoid duplicated utensils.
Avoid duplicated ingredients.
Avoid food clipping.
Avoid inconsistent ingredient shapes.
Avoid inconsistent pot geometry.
Avoid liquid flowing backward.
Avoid physically impossible steam movement.
Maintain exact continuity of the same human hands, same utensils, same wooden chopsticks, same cast-iron sukiyaki pot, same ceramic bowl, same kitchen, same dining table, same ingredient set, same lighting direction, and same spatial environment throughout the entire video.
The entire 30-second sequence must feel like one continuous premium Japanese food commercial filmed in a real kitchen, with physically accurate cooking, realistic human interaction, consistent food appearance, natural camera movement, and authentic live-action photography.
Show more
A new and possibly controversial perspective:
In this video, I explain the sense in which generative AI trained by supervised learning is incapable of making novel discoveries.
The text of the speech:
AI Creativity and Discovery
Good day ladies and gentlemen. I regret that I am unable to be with you all today to engage in a back-and-forth discussion, but I am nevertheless pleased to be able to share with you, via this recording, some high-level thoughts about the current and future state of artificial intelligence, and in particular about AI’s relationship to science and mathematics, which is, as I understand it, the central focus of this meeting and of the SAIR Foundation.
I would like to start with an old joke; I am sure you have heard it before. It is the one about the researcher whose work is being evaluated, and the review comes back, and says “This work is both novel and good. Unfortunately, the parts that are good are not novel, and the parts that are novel are not good.”
My first point about AI is that this assessment applies exactly to large parts of AI as we know it today. Not all of today’s AI, but a large part of it. Pretty much all of what we mean by “Generative AI”---which includes large language models, and the images and video models, and even the new methods for learning world models. All of these AIs take large numbers of examples and produce a “model” which behaves similar to the examples, that is, which generates text like people, or images like artists or nature, and videos like we find on the internet. Don’t get me wrong, Generative AI can be extremely useful. No doubt about that. But the assessment of the joke still applies. These systems can produce output that is both novel and good, but not at the same time.
In many ways this is just absolutely not a problem. When we ask an AI for an answer from the internet, or to summarize a document, we don’t want it to be novel. We are happy if the quality of the answer, the goodness, comes from the source material—from the people who wrote the document or the articles on the internet. If the AI’s answer is novel it means it is going beyond the source material, adding something beyond it. This is what we call “hallucinations”. In most cases, we don’t like it when the AI makes something up, when it adds something novel.
One exception, of course, is when we are looking not for facts or reality, but for fiction and entertainment. We might ask for a bedtime story for a child, or an image based on existing images on the internet but which is nevertheless different and distinct from them. In these cases, it is never easy for us to know how creative the AI is actually being, as we do not know how close the AI’s story, poem, or image is to the source material. In a real practical sense we can not know this because the internet is too big, the possible sources that the AI may draw upon are too numerous.
When we ask for a fiction or novelty, the AI can give it to us because its processing is in part stochastic. Every decision can go multiple ways and will go different ways and produce a different trajectory every time. The trajectory can be random—and thus novel—or it can be based on the training data—and thus “good” because the training data is good, sourced from people or reality. Thus, the trajectory is either novel or good—based on randomness or based on data—but never both at the same time.
Really, I think it is okay if the output of Generative AI is never good and novel at the same time. For the researcher in the joke this is a devastating criticism, but for most things it is not, and for Generative AI it is not. Generative AI is meant to be a mimic. This is what supervised learning is for. Generative AI can be extremely useful, even when it just mimics, if it is faster, or cheaper, or smaller, or more customizable, or more copy-able, than the thing being mimicked. It is okay if Generative AI cannot be both novel and good at the same time. It is still a transformative technology.
But it is a limitation. And remember we are here to use AI for science and mathematics, and for these areas the assessment of the reviewer in the joke is devastating. For these areas we need true creativity and discovery. Generative AI—or Mimicking AI—will never get where us there. For these we need something more, and indeed we have something more in other parts of AI. We have many AI systems which can give us more. We have AlphaGo with its world-changing move 37, or AlphaZero with its brilliant original chess-playing style. We have GT-Sophy that drives simulated racecars better than any human. We have AlphaFold and AlphaProof and Claude-Code, which have brought true advances in science, mathematics, and programming. We have RL-Lyft which optimizes the assignment of cars to passengers in the ride-hailing business. All these systems have found things that are both novel and good. And, truth be told, some language models have been augmented in ways that make them more than Generative AI based on supervised learning.
All these systems have some additional features that make them capable of true creativity and true discovery. It is important for us to recognize what this is—and that it is not present in ordinary, garden-variety Generative AI. It is something that can not come from just supervised learning, from learning from examples. What is it? Well, it is a simple thing, a commonsense thing. It is not new. We have many names for it, but unfortunately none of them are very good names. I will call it Discovery. Basically, Discovery is just the idea of trying many things and seeing which of them work, then keeping those that worked the best. Evolution by natural selection works this way. The scientific method works this way. And just ordinary life and learning works this way. We try things and remember what works. What could be more obvious? In this behavioral case, psychology has two names for it— “instrumental learning” and “operant conditioning”—and in machine learning it is what we mean by “reinforcement learning”. We also see the idea of Discovery in planning and combinatorial search—anything that involves the idea of “generate and test”.
The essence of Discovery is to combine three steps:
1. Variation,
2. Evaluation, and
3. Selective retention.
Of course, I am not the first to say this. I am not the first to point out that this combination of steps is key to science, to evolution by natural selection, and to animal behavior. I think particularly of papers by Donald Campbell, by Daniel Dennett, and by Gary Cziko. What is new in my remarks is to directly relate the idea of Discovery to modern AI to help us see that it is not present in supervised learning or Generative AI—in particular, that Discovery is not present in backpropagation or gradient descent.
Let me say explicitly what is missing from Generative AI. As we have remarked, these systems do have a stochastic aspect, so they do generate a variety of trajectories and behavior. What is missing is the Evaluation step. The generator was pre-trained by supervised learning, leaving no way at runtime to Evaluate what it generates. And of course without Evaluation there can be no Selective retention, and thus no Discovery. The variation can bring novelty, but without evaluation there is no Discovery, and arguably, no creativity. That is, I would say that creativity requires that the new things generated be Evaluated. Without evaluation, and retention of the best, there is nothing created. The novelty flickers into existence but, if its value is unrecognized, it flickers away and is lost.
In many cases, Evaluation is done by people to make a discovery. As when we have Generative AI make many pictures for us, and then we pick the one that we like the best. The human+AI system completes the discovery.
In many other cases, the Evaluation comes from a clear objective. Some moves lead to checkmate, some steps lead to a proof, some actions result in high reward, some genotypes make more copies, some theories explain the data better.
Some prefer the Variation step to be called Blind variation, where “blind” here means that it is uninformed, a shot in the dark. It does not need to be completely uninformed; a good scientist does not select theories to test at random. But neither can it be completely informed and determined. There must be some uncertainty about where the answer lies in order for there to be a discovery. In practice, the variation is partly informed and partly blind, but it is the blind part that corresponds to the discovery.
Now let us briefly go all the way to modern deep learning, to the backpropagation algorithm. At first it might seem that backpropagation is incapable of discovery because it is deterministic and thus incapable of variation. But this is not correct. The weight updates of backprop are deterministic, but the weights are initialized to small random values. The random initialization is often downplayed, but in fact it is a necessary form of variation; it must be done properly to get good performance. In backprop this Variation is done once, at network initialization, so its effect is temporary, and later the network may lose its ability to learn. This is the weakness of deep learning that is alleviated with a new algorithm that my group presented in Nature a couple of years ago. Our “continual backpropagation” made one small change: every so often a less-used neuron would be re-initialized to small random weights. This allows the variation to continue and plasticity to be retained.
Although there is much more to be said about Creativity and Discovery, this is the key point: they are more than supervised learning, more than pattern recognition, more than prediction, and more than world modeling. Those things are important, but they alone will not bring us to discovery. Discovery requires Evaluation from a person or from an explicit goal, and only in the latter case will we attain full autonomy.
So that is my call to arms. If we want the full power of AI scientists, then we should share the goals with them so they can create, evaluate, discover, and in these ways fully participate in achieving the goals. Let’s be bold! Let’s fully automate Creativity and Discovery!
Show more