Register and share your invite link to earn from video plays and referrals.

Enze Xie
@xieenze_jr
Tech Lead & Staff Research Scientist @ NVIDIA, Efficient VideoGen / SANA / Sol-Engine , CS PhD from HKU MMLab.
336 Following    3.1K Followers
Day-0 Sol Engine + Sol-Attn for @video_rebirth HyperFlow. 8-step, data-free distill of MiniMax-H3. One LoRA: T2V / FL / Ref. Sol-Attn is the training-free sparse path — near-lossless vs dense attention, no extra fine-tune. On 8×B200, same HyperFlow recipe before vs after Sol: T2V 5s 9.89s → 2.75s (3.6×) T2V 15s 36.93s → 11.85s (3.1×) Ref 5s 15.27s → 2.97s (5.1×)
Show more
We're open-sourcing H3 HyperFlow. A data-free flow self-distillation technology built on @MiniMax_AI H3. No external training data. Significantly reduces inference cost while preserving frontier model quality. Full demo:
Show more
Sol-engine helps accelerate H3 video generation 😄
Open weights. Shared progress. MiniMax H3 is moving fast. We built H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: 15 seconds of 768p video + audio in 6.6 seconds on 8×B300, in the team’s warm-inference benchmark.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sharing. Special thanks to: @haoailab @nuvalab @NVIDIAAI @xieenze_jr @HaochengXiUCB @ArashVahdat @julberner @LightX2V @ComfyUI And to the individual contributors pushing the work forward: @haozhangml @cxlcl1 @lawrence_cjs @yitongli165665 @haopengl33 @songhan_mit @shanasaimoe Thank you for building with H3 and helping make it faster, more accessible, and more useful for the community. Powerful models go further when we build together. Keep pushing H3. Excited to see what comes next. 🚀 Explore the ecosystem:
Show more
🚀 Sol-H3 on DGX Spark: 768p in Under a Minute 🤩 Monday: 8×B300, 5s 768p in 1.65s — faster than playback. Today: the same stack on one desktop Spark — about 56s hot E2E. Five seconds of 1344×768 video at 24 FPS with stereo audio, on a single NVIDIA DGX Spark (GB10). Two-stage, not the datacenter profile: 384p H3 draft → latent ×2 → H3-to-LTX VAE adapter → 768p LTX refine → VAE decode No decode/re-encode between stages. Stage 2 is conditioned on the draft latent, so Gemma stays off the box. Quantized weights stay resident; sparse attention cuts the rest. Stage 1 takes any MiniMax-H3 few-step LoRA. Timing is hot E2E (encode → both stages → video/audio VAE); cold start and MP4 mux are separate. Apache 2.0. Server was realtime. Edge is one box, under a minute. 🔗 Amazing team effort—full credits in the blog. @haopengl33 @lawrence_cjs @yitongli165665 @shanasaimoe ,Jingyu Xin, @HaochengXiUCB @songhan_mit
Show more
RSI for efficiency🚀
🚀 Introducing SoL-Pi from NVIDIA. ⚡️ RSI for Efficiency. Efficiency for Swarm Intelligence. Before scaling to thousands of parallel agents, can we first make each agent waste fewer tokens? SoL-Pi targets the harness layer for Agent ! 1/8
Show more
RSI for efficiency 🚀🚀
🚀 Introducing SoL-Pi! ⚡️ RSI for Efficiency. Efficiency for RSI 🔮 We observe scaling laws in RSI loops. 🌐 Official Blog Read the blog → 💻 Official Code Code →
Show more
Thanks Minimax for broadcasting our work!We will also release a spark version of Sol-H3 pls stay tuned 😁
New video generation acceleration for MiniMax-H3 from the NVIDIA Sol team! Faster than playback, fully open-sourced, and available via the Reactor API so you can try it before committing your GPU😜
Show more
🚀 Sol-H3: @MiniMax_AI H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with @reactorworld to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: 🔗 Amazing team effort—full credits in the blog. @shanasaimoe @lawrence_cjs @yitongli165665 @haopengl33 @HaochengXiUCB @songhan_mit
Show more
Check out Sol-H3 served on @reactorworld ~ Thank you guys for the continued support for our Sol-Engine!
We're proud to partner with @nvidia's SANA team to ship FastH3 with Sol. For the first time, generate clips 3x faster than realtime. Available day 0 via the Reactor API. Try it here:
Show more
Great work!
(1/6) Open Weight @MiniMax_AI FastH3 v1: Generate 15s 768p video in 13s 🚀 - FastVideo collab w/ @nuvalab + FastGen - Up to 14x speedup on @NVIDIAAI Blackwell GPU - Fully open so community can run and improve the acceleration recipe. The era of open weight video models just started, that calls for post training. With great community effort such as Minimax H3 Max based on @MiniMax_AI , we have more to share about how post training can improve speed, quality and most importantly, how it’s done, with open weights and recipes.
Show more
pls try this and give us feedback! Will release Super Acceleratior v1.1 too~😄
The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3! By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s (27.7x speedup). Replacing heavy VAE decodes with TAEH3/TAEHV while holding the latents stable for refinement is a masterclass in co-designing sampling topology with hardware kernel acceleration. When inference latency collapses this dramatically, unit economics fundamentally shift: a single node can suddenly serve 378K videos a month at 97%+ GPU margins. This is how high-fidelity AI video moves from asynchronous batch rendering to near-instant, interactive infrastructure. Huge respect to the team for setting a new engineering bar for our open-weights ecosystem! 🫡🩵
Show more
The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3! By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s (27.7x speedup). Replacing heavy VAE decodes with TAEH3/TAEHV while holding the latents stable for refinement is a masterclass in co-designing sampling topology with hardware kernel acceleration. When inference latency collapses this dramatically, unit economics fundamentally shift: a single node can suddenly serve 378K videos a month at 97%+ GPU margins. This is how high-fidelity AI video moves from asynchronous batch rendering to near-instant, interactive infrastructure. Huge respect to the team for setting a new engineering bar for our open-weights ecosystem! 🫡🩵
Show more
🚀 MiniMax H3 Super Acceleration in Sol-Engine🤩 We pushed H3 far beyond our previous 3–4× optimization regime — reaching 22.2× speedup for 5s video and 27.7× for 10s video vs. the published SGLang baseline. On a single NVIDIA GB200: • 5s 768p: 152.3s → 6.85s (22.2×) • 10s 768p: 414.1s → 14.93s (27.7×) But the more interesting part may be what this means economically. Using MiniMax’s published H3 API price as a reference, we translate inference speed directly into production economics. Under an ideal fully utilized GB200 scenario, Sol-Super can serve about 525 five-second videos/hour, corresponding to roughly $210/hour of output value at the reference API price. Assuming $5.50/GPU-hour, that implies a 97%+ GPU-only gross margin in the idealized model. At full utilization, one GB200 could produce: • 12.6K 5s videos/day • 378K videos/month • equivalent to 525 hours of finished video per month For us, this is the bigger point of inference optimization: a 20×+ speedup does not just reduce latency — it can fundamentally change the unit economics, serving capacity, and viable business models of video generation. 🔗
Show more
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!
Show more
0
211
3.8K
489
Forward to community
🔥MiniMax-H3 on GeForce RTX 4090 — reaches 4.44× end-to-end on RTX 4090. Sol-Attn now includes an optimized SM89 CuTe DSL kernel for RTX 4090.
In one week, we also support @MiniMax_AI H3 in H100 and A100, feel free to check what hardware sol-engine supports 🎉
🚀 LTX-2.5 × Sol-Engine — from data center to desktop. Just one day after the release of the new open-source LTX-2.5, we’ve brought it into Sol-Engine and optimized it across a wide range of deployment settings: ⚡ GB200 — high-performance data center inference 💻 DGX Spark — compact desktop AI deployment 🎮 GeForce RTX 5090 — consumer Blackwell GPU LTX-2.5 comes with a more complex multi-stage inference pipeline and multiple generation configurations. Sol-Engine automatically adapts its full-stack optimization recipe across these different workloads and hardware targets — combining kernel optimization, attention acceleration, caching, and system-level optimization where they matter most. Really excited to see another strong open video generation model released to the community — and to make it faster and easier to run everywhere from GB200 to desktop GPUs. 🚀 🔗
Show more
awesome, this pipeline is similar to NVIDIA dlss, which is generarive rendering, the video dit only need to focus on rendering the draft
you don't need Seedance 2.5, here's cheaper and more efficient workflow we made this entire game shooting scene in one shot for $1.97 with DeepSeek V4 Flash 0731 + MiniMax H3 we use DeepSeek to code the raw Three.js scene first: exact camera movement, character motion, timing, everything. keep tweaking it there for basically pennies. then feed the final footage into H3 and let it cook the realism 48 mins, $1.97 total, one H3 generation way cheaper than rerolling a video model 10 times
Show more
In one week, we also support @MiniMax_AI H3 in H100 and A100, feel free to check what hardware sol-engine supports 🎉
🚀 MiniMax H3, accelerated on Day 1 with Sol Engine! Within just 4.5 hours, our agent-native Sol Video Inference Engine achieved: ⚡ 3.95× end-to-end speedup over Diffusers ⚡ 2.80× speedup over SGLang 🎬 8× NVIDIA GB200, 1344×768, 24 FPS, 124 frames The acceleration combines kernel fusion and graph capture, cross-step caching, and training-free sparse attention powered by Sol-Attn—with no distillation, fine-tuning, LoRA, or offline calibration. We are especially excited to see a powerful open-weight model like MiniMax H3 released to the community. Open models are essential for pushing video generation research, systems optimization, and real-world deployment forward. We hope this is the beginning of a much more vibrant open-source video generation ecosystem—and Sol Engine will keep working to make the latest models faster and easier to deploy from day one. 🔗
Show more
🚀 Sol Video Inference Engine is here! An agent-native, training-free full-stack accelerator for video diffusion. It auto-tunes cache + sparse attn + token pruning + quant + kernel fusion for any model/hardware/config. >2× end-to-end speedup on 64B Cosmos3-Super, 22B LTX-2.3 and 2B SANA-Video — near-lossless VBench quality, minimal human effort. Practical acceleration for real video gen deployment. 📄 Paper: 🌐 Project: 💻 Code: Proud of the team! 🎉
Show more