Register and share your invite link to earn from video plays and referrals.

Zbigniew Majewski
@majewskizby
47 Following    56 Followers
This doesn’t bode well for hardware prices 🤣
@minchoi Colossus 1 is 150k H100, 50k H200 and 30k GB200. Colossus 2 is 110k GB200 and 440k GB300. Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
Show more
Wow, Agora-2 doesn’t load a world. It invents one in real time - every frame generated on the fly, no map, no engine, just a model simulating the fight as you play 🫣
Introducing Agora-2, our next-generation multi-agent world model. Agora-2 supports up to 20 humans and agents interacting inside a shared environment, all simulated in real time. Our multiplayer research preview is available to try right now!
Show more
More DS4.1 TP4 improvements on 4x DGX Spark 🚀 - Prose decode: 85 tok/s at c1 (was 69) - Boot: ~3 min - Prefill: ~4.8k tok/s at 32k–128k, 4.3k at 262k - KV cache: ~6.5M tokens, 1M context - Same weights, same quality (qeval unchanged)
Show more
Pushed DeepSeek 4.1 Flash TP=4 (4 sparks) to 69 t/s on prose!
We'll probably be seeing more and more of this 😍
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more:
Show more
GPT 6 Sol worse at coding than 5.6 Sol
MiMo V2.6 Flash looks like the sweet spot for local inference. It delivers roughly 90-97% of Pro's practical coding and agentic performance while using just 15B active parameters vs 42B. On 4x DGX Spark, rough estimates are around 50-60 tok/s for Flash vs 20-25 tok/s for a heavily quantized Pro.
Show more
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:
Show more
Huge!
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: - GitHub: - Model Scope: - Hugging Face:
Show more
Tokens/sec lied to me. Same 30 tasks, 4x DGX Spark: DeepSeek-V4.1-Flash: 61 tok/s, thinks 39k tokens → 671 s GLM-5.3-Flash NVFP4 (effort high): 37 tok/s, thinks 5k tokens → 324 s Same answers. GLM done 2.1x sooner. At effort max GLM thinks like DeepSeek and loses. Agentic coding (6 repo tasks via OpenCode): DeepSeek 172 s, GLM 243 s. Little thinking there, raw speed wins. Rule: GLM on high for chat, DeepSeek for tool loops. Full numbers, launcher, harness:
Show more
Lmao, this is Mythos:
JUST IN: Some Anthropic engineers are reportedly "worshipping" Claude as a God.
Heads up for DGX Spark owners: capping clocks with `nvidia-smi -lgc 0,2200` keeps the GB10 cooler and the clocks stable but it does NOT persist across reboots or driver resets. Wrap it in a systemd timer (OnCalendar=*:0/10) or your "capped" box quietly drifts back toward 3 GHz.
Show more
I've published this DeepSeek 4.1 Flash recipe for 4x DGX Spark (mainly clone of Mia's one with tweaks and experimental libs etc). Prose C1: 45-> 55 tps. C16: 134 -> 277 tps.
Show more
A new model.. Let the dopamine loop continue
Union Alpha (stealth model) is free for the next week - no data training - built for agentic coding - supports images let's see what you can do
Hugging Face banned its first model. This is probably just the beginning. The community is moving to torrents instead.
When working with local models please don’t focus on tokens per second. Focus on time per completed task ⏱️ including testing and fixes. That’s the metric that actually matters 🎯
This is insane. Mouse next?
A fly that is a fly, its own connectome brain, perceiving the real room through its own senses with power of AR now have digital 🪰 Flow: Spectacles(retina + rays + hands) → (room inventory) → sense channels → MaleCNS brain → DNs + motor pools → Specstacles → wings, legs, head @specsfordevs #fly#
Show more
Created 2 PRs for the @MiaAI_lab DeepSeek 4.1 Flash recipe on 4× DGX Spark. Combined: +16% single-stream, up to +41% at high concurrency 😎
Thank you Mia! 🔥
Run DeepSeek v4.1 Flash on your 3x DGX Sparks ✨ Conservative settings by default for stability: - 500k context, 750k KV cache - prefill is ~1600-1700 tok/s in all ctx ~38 tok/s prose in single stream ~79 tok/s prose in 4 concurrent streams Full prose decode tok/s numbers in below post. Get it here:
Show more