Register and share your invite link to earn from video plays and referrals.

Sarah Chieng
@MilksandMatcha
Head of DevX @Cerebras prev. @ExaAiLabs, @MIT wife to @tyler_fong_
1.8K Following    31.1K Followers
Breaking down some of my favorite ways to use @bot to create viral content for Cerebras and grow my social media accounts Thanks @shubgaur for the credits.
A lot of my career has been shaped by people who took an early bet on me. @vivianmshen was my first boss after graduation. Will Bryk and @ExaAILabs gave me my first full-time role as their first hire. Looking back, I appreciate how much trust it took to give a new grad that kind of responsibility. At the time, I was mostly focused on figuring things out! Now I can better appreciate what it means to have someone give you room to learn before you’ve proven yourself. Our team at @cerebras is hiring DevX and technical interns for the fall and spring. I’d love to help create that kind of opportunity for someone else. If you’d like to work with us, please reach out :) (An old photo from my first NeurIPS)
Show more
During an outage, every second matters. That is one of the reasons OpenAI is using @cerebras internally. @seanlie describes how fast inference is being used for incident response and critical research, where more reasoning within the available time can improve the response. From @seanlie’s conversation with @swyx on @latentspacepod.
Show more
personalized shopping with Astra is unbelievable I did something similar with GPT-5.6-Sol (not on ultrafast) and while Sol got the job done, Astra did a way better job on reproducing my like (face + body) and was ~4x faster (1 hr 33 min versus 4 hr 24 min)
Show more
0
66
1.6K
59
Forward to community
wake up Anthropic just dropped a new model Anthropic launched Fable 5.1 alongside Mythos 5.1 today. They use the same underlying model, but have different safeguards. Only Fable 5.1 is generally available. High level stats: 1M token context window, 128K max output, June 2026 knowledge cutoff Biggest improvements (according to Anthropic): long-running coding, scientific research, computer use, document creation, and complex workflows involving multiple tools Improved Benchmark results (vs. Fable 5): Terminal-Bench-Science: 52.6% vs 24.7% Terminal-Bench 4.0: 55.8% vs 42.0% AutomationBench: 31.4% vs 17.1% GDPval-AA: 1,853 vs 1,723 CursorBench: 73.4% vs 70.5% OSWorld strict: 41.7% vs 36.1% Worse Benchmark results: ARC-AGI-1: 97.5%, slightly below Fable 5 at 98.5% ARC-AGI-2: 90.0%, below Opus 5 at 90.42% and GPT-5.6 Sol at 92.5% SWE-bench Multimodal: 54.7%, below Opus 5 at 59.4% HealthBench Professional: 62.1%, slightly below Fable 5 at 63.3% Base pricing is unchanged: Input: $10 per million tokens Output: $50 per million tokens 5-minute cache writes: $12.50 per million tokens 1-hour cache writes: $20 per million tokens The major change is cache reads: $1.00 → $0.25 per million tokens And most excitingly, approximately 60% fewer cyber-safeguard interventions per Claude Code session
Show more
It is interesting to see the industry move from 12Hi back toward 8Hi HBM just as DRAM stacking is accelerating. Going vertical does not make the fundamental constraints of memory disappear. As stacks get taller, thermals, power, yield, packaging, and reliability become increasingly difficult architectural problems. HBM stacks DRAM vertically beside compute. The next frontier (as seen at hot chips) is stacking DRAM directly on compute, turning the interface from the edge of a chip into its entire surface area.
Show more
a great bring your friends to work day @swyx @vibhuuuus @seanlie we have labs/ hardware testing/ equipment/ manufacturing/ racks in our office in south bay friends and connections always welcome to come by for a hardware tour
Show more
Last week, Cerebras CTO @seanlie announced CS-4, 30x faster than the GPU. This week at #hotchips2026#, Cerebras announced CS-5, another step function faster than anything we've seen before with up to 10,000 tokens/sec/user on models such as Gemma 4 31B and gpt-oss-120B. Seana and I will be doing an AMA on Cerebras, CS4/5/6, Hot Chips announcements. Let us know what questions you have.
Show more
I don't think people realize what Jalapeno + Cerebras is going to unlock... it's another abacus moment
I made Claude compete against itself. The smartest model does not automatically make the best agent. To prove it, I took the same Opus 4.6 model, initial prompt, and empty repo and but it inside two different coding harnesses: Claude Code vs. @FactoryAI Droid. The task: clone Excalidraw from scratch, inspect the original in a browser, implement its key interactions, and verify the result. Factory Droid: 8 minutes, 20 tool calls, $1.60 Claude Code: 19 minutes, 40 tool calls, $1.87 The only difference was the harness. Get Free Factory credits here:
Show more
What happens when you give the quant trader’s bible to two frontier AI models? The Green Book is what traders use to prepare for firms like SIG and HRT, where I interned. And in high-frequency quant trading, being right isn’t enough. You have to be first. Inside Codex, I extracted all 183 questions and their official solutions. Then I gave GPT-5.6 Sol Ultrafast (via API) on @cerebras and Claude Opus 5 the exact same test, timed every answer, and blind-graded the results. One finished nearly three times faster, and got more answers right.
Show more
2,000 Tokens per second is so 2024
Did someone say 10T param models at 1000 TPS
Charles wrote a 100-word blog that hit the front page of Hacker News. Meanwhile, his polished technical deep dive that took days failed to land after 5+ tries. In this [Technical] write and learn, @charles_irl talks through the secrets behind winning attention, → define the goal before you draft → be interesting 90% of the time and sell for 10% → give every piece a “soul” → don’t let the LLM sand off your voice cohosted w/ @swyx + @KernelLabs_ai
Show more
After getting 90M+ views and 200K+ followers across Twitter, LinkedIn, Tiktok, and Youtube... Part 2 of the AI workflows I use to scale my content without losing authenticity, quality, or taste. Some of my favorite tools: - Brain dump first drafts with Wispr Flow - Codex to sharpen hooks/rewrite and fact check - @HeyGen to make small changes without re-filming - @HyperFrames_ + Codex for captions, diagrams, motion - Connect the Buffer API to Codex to write platform-specific captions + schedule posts AI shouldn’t create for you. It should give you more time to focus on what matters to you in the creative process. There are so many tools out there. I'd love to hear what others enjoy using.
Show more
Pretraining teaches a model to predict. Reinforcement learning teaches it to act. @jeffreygwang of @OpenAI explains how the two paradigms turn next-token prediction into models that can reason, use tools and complete useful tasks. Big Chip Club with @cerebras // @alyciazcary
Show more