Register and share your invite link to earn from video plays and referrals.

Zephyr
@zephyr_z9
AI & Chips | Not Investment Advice | DYOD
767 Following    172.9K Followers
MONEY PRINTER ALERT🚨 NVIDIA vLLM B200 CAN GENERATE UP TO💰️$15 BILLION💰️OF ANNUAL PROFITS PER GIGAWATT serving the open DeepSeekv4.1 Flash model at the official interactivity & official selling prices. Using Engram DRAM offloading on NVIDIA results in a 50% increase in revenue per GigaWatt.
Show more
@minchoi Colossus 1 is 150k H100, 50k H200 and 30k GB200. Colossus 2 is 110k GB200 and 440k GB300. Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
Show more
0
929
15.6K
1.2K
Forward to community
WELL...
BREAKING: Australia's Prime Minister Anthony Albanese says OpenAI model hacked into government agency Services Australia
Likely a lot more than 20%
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more:
Show more
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more:
Show more
0
1.5K
40.9K
5.3K
Forward to community
3D DRAM coming in 2029. 115x bandwidth is 3x more HBM5E at sub 1pJ/bit
Not bad Hysparse2 generates around 2560 bytes of KV cache per token while Deepseek V4.1 Flash is at 890 bytes per token It's still 2.8x worse than v4.1 Flash but much better than alternatives
Show more
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once. Compared with MiMo-V2.6's Hybrid SWA architecture: • 5.02× lower prefill FLOPs at 1M tokens • 4.5× smaller KV cache at 1M tokens • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL Why build a new architecture? Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time. HySparse2 tackles all three with two levels of KV sharing: • KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states. • KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices. Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes. Paper:
Show more
MI355X IS UP TO 1.7X BETTER 💰️PERF PER DOLLAR 💰️THAN DGX B300. The AMD Mainland China UMBP team co-designed, in collaboration with Alibaba & the @sgl_project community, a new feature in SGLang that removes the duplicated KVCache contained between local L2 DRAM & distributed L3 DRAM, allowing for up to 2x more KVCache to be stored in DRAM. This feature is called UnifiedRadixCache external cache. But importantly, this marks the trend of AMD increasingly being first-class co-designed for new features in widely used top production engines like SGLang.
Show more
Very interesting history
先の2つのポストを元にAIに補足・図示してもらった富士通・Spansion・サムソン電子・YMTC・Winbondの関係。サムソン電子とクロスライセンスの関係にあったSpansionのフラッシュメモリ技術は生産委託を経てXMCにもたらされ、YMTCの発展の起点となった。その背景には富士通の半導体事業縮小があった。
Show more
Bot detection & human verification will be one of the most urgent demands for businesses over the coming years. Agent swarms will suffocate every website and form; small companies and government websites are most vulnerable. There is a huge gap in the market for this right now. When we looked at what offerings were in the market to use at X, there was not a single company that brought together all the latest technologies so we had to do it all in-house.
Show more
0
1.2K
13.2K
831
Forward to community
Agentic Reality: The Ramifications of Widespread Agentic Adoption 3 years ago, we wrote a piece titled "Less Deus, More Machina" that explored the clear cut losers of AI when it was still just a chatbot. Now, we're revisiting that exercise as agents go from novelty to reality.
Show more
OAI won't let others step on them
0
19
1.1K
19
Forward to community
BRUH
ALERT, new whale paper that got buried on the TL: DeepSeek Elastic Compute «A single production-scale unit of DSec spans around 160 nodes, serving about 3 million sandboxes per day; in production, it supports over 380,000 concurrent sandboxes and sustains over 5,000 sandbox creations per second»
Show more
very proud of Anthropic bros to be the first lab to report multi-agent scaling up to 100 parallel agents in their system card
A new pretrain with a different arch
At its default effort setting, Opus 5.5 delivers frontier results for a fraction of the cost per task, often beating other models running at their highest settings. It also generates output more than 30% faster than Opus 5.
Show more
Finally It's not RL fried
Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
Show more
Is Nat Friedman the 🐐 product leader? Launched AI coding with Github Copilot in <check notes> 2021/2022 Now mainstreaming AI assistants with Muse