๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Enze Xie
@xieenze_jr
Tech Lead & Staff Research Scientist @ NVIDIA, Efficient VideoGen / SANA / Sol-Engine , CS PhD from HKU MMLab.
๊ฐ€์ž… November 2023
336 ํŒ”๋กœ์ž‰ ์ค‘    3.1K ํŒฌ
๐Ÿš€ Sol-H3 on DGX Spark: 768p in Under a Minute ๐Ÿคฉ Monday: 8ร—B300, 5s 768p in 1.65s โ€” faster than playback. Today: the same stack on one desktop Spark โ€” about 56s hot E2E. Five seconds of 1344ร—768 video at 24 FPS with stereo audio, on a single NVIDIA DGX Spark (GB10). Two-stage, not the datacenter profile: 384p H3 draft โ†’ latent ร—2 โ†’ H3-to-LTX VAE adapter โ†’ 768p LTX refine โ†’ VAE decode No decode/re-encode between stages. Stage 2 is conditioned on the draft latent, so Gemma stays off the box. Quantized weights stay resident; sparse attention cuts the rest. Stage 1 takes any MiniMax-H3 few-step LoRA. Timing is hot E2E (encode โ†’ both stages โ†’ video/audio VAE); cold start and MP4 mux are separate. Apache 2.0. Server was realtime. Edge is one box, under a minute. ๐Ÿ”— Amazing team effortโ€”full credits in the blog. @haopengl33 @lawrence_cjs @yitongli165665 @shanasaimoe ,Jingyu Xin, @HaochengXiUCB @songhan_mit
๋” ๋ณด๊ธฐ