Register and share your invite link to earn from video plays and referrals.

Cerebras
@cerebras
The world's fastest AI inference and training. Try the latest open models at:
271 Following    76.3K Followers
CS4 uses the same wafer as CS3. So where does the extra speed come from? You asked on Twitter and Discord. Join Angela Yeung, SVP of product and @MilksandMatcha, on the behind-the-scenes on the engineering behind CS4: connecting wafers, recirculating cooling water, and getting compute deployed at scale. 1:08 Where does the extra speed come from? 2:33 Connecting wafers for 10-trillion-parameter models 4:15 What breaks first when you scale? 6:07 Breaking down cooling 7:30 Easy CS4 deployment
Show more
We built Money Agent, a personal finance assistant powered by @Alibaba_Qwen 3.8 27B on Cerebras. Now you can turn a home-buying question into a real-time conversation, then into a financial goal, faster.
Show more
Connecting the dies and working around defects were fundamental challenges for wafer-scale computing. Making it work also came down to power, cooling, and reliability. @seanlie explains how those lessons are shaping @cerebras’ approach to stacking DRAM. From @seanlie’s conversation with @swyx on @latentspacepod.
Show more
Qwen3.8-27B is now live at Cerebras speed. The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.
Show more
0
88
1.4K
82
Forward to community
Running GPT-5.6 Sol Ultrafast on Cerebras, @MatthewBerman saw his workflow flip: fewer parallel agents, same productivity, lower mental load. When tokens stop being the bottleneck, the rest of the stack becomes the next layer to optimize. Faster inference opens that door.
Show more
It is interesting to see the industry move from 12Hi back toward 8Hi HBM just as DRAM stacking is accelerating. Going vertical does not make the fundamental constraints of memory disappear. As stacks get taller, thermals, power, yield, packaging, and reliability become increasingly difficult architectural problems. HBM stacks DRAM vertically beside compute. The next frontier (as seen at hot chips) is stacking DRAM directly on compute, turning the interface from the edge of a chip into its entire surface area.
Show more
Today, we announced a new 165 MW data center in Mikkeli, Finland, in partnership with Compute Nordic Finland. Construction of the first 50 MW is already under way. Two things make this project special to me: First, we designed the Mikkeli data center to give back. The facility is designed to capture the heat our computers generate so the community can use it, and the water that cools our systems gets recirculated instead of drawn from the town's supply. Second, it creates long-term jobs. An estimated €1.0–1.7 billion will be invested into the region and it will create up to 250 permanent technical jobs in South Savo. Thank you to Compute Nordic Finland, the City of Mikkeli, and our network of Finnish partners for helping bring this to life.
Show more
a great bring your friends to work day @swyx @vibhuuuus @seanlie we have labs/ hardware testing/ equipment/ manufacturing/ racks in our office in south bay friends and connections always welcome to come by for a hardware tour
Show more
Ten years ago, we started @cerebras around an approach many believed was impossible. As a computer architect, it is hard for me to imagine a more exciting time. Model releases are accelerating, and hardware tapeout is compressing from multi-year roadmaps to annual launches. Hot Chips is my favorite conference, and it’s where I launched Cerebras 7 years ago. This year’s conference was especially exciting, and so much innovation was shared. I am watching the industry recreate itself: SRAM is mainstream, DRAM is moving into the third dimension, networks are being fundamentally redesigned, and AI is helping design and program the chips themselves. The industry has never moved faster and some of the hardest architectural questions are still wide open.
Show more
0
74
1.7K
157
Forward to community
where in the world is cafe compute this week? 🇯🇵 Tokyo: 🇲🇾 Kuala Lumpur: 🇮🇳 Hyderabad: big thanks to our @cerebras ambassadors for hosting @asahiXXXXXXXXX @hazlijohar @jayant99acharya
Show more
At Cafe Compute Fal.Con, we’re giving one lucky winner the opportunity to experience Cerebras-like speed from behind the wheel of a supercar at SPEEDVEGAS. The winner take a supercar onto a purpose-built Las Vegas racetrack with a professional instructor riding alongside. Ferrari, Lamborghini, Porsche, McLaren, Aston Martin, Mercedes-AMG, Nissan GT-R, Lotus. Which car will you drive? RSVP: Come see us at Booth # 1039 at Fal.Con with @CrowdStrike *** No purchase necessary. One winner will be selected at random at Cafe Compute on Sept 1. Must be in attendance at time of drawing to win. Must be 18+ and have a valid driver's license. Prize is subject to vehicle and reservation availability. Participation may require acceptance of SPEEDVEGAS’s terms and liability waiver. Cerebras is not responsible for personal injury, death, property damage, or other loss arising from participation in or use of the prize.
Show more
Congratulations @andrewdfeldman on being included in @TIME's list of TIME100 AI. 🟧🚀
Last week, Cerebras CTO @seanlie announced CS-4, 30x faster than the GPU. This week at #hotchips2026#, Cerebras announced CS-5, another step function faster than anything we've seen before with up to 10,000 tokens/sec/user on models such as Gemma 4 31B and gpt-oss-120B. Seana and I will be doing an AMA on Cerebras, CS4/5/6, Hot Chips announcements. Let us know what questions you have.
Show more
Congratulations to our partners at @OpenAI today. Taking a new chip to real lab results at great performance is a huge achievement, and improving speed across the board pushes our whole industry forward. We look forward to having Jalapeno and the next-gen Cerebras solution together in production in 2027, driving OpenAI’s comprehensive fast and ultrafast inference product portfolio to continue evolving what's possible with fast AI.
Show more
Excited about our Jalapeno results today. An incredible achievement from the team, taking a new chip from concept all the way to very impressive performance on real workloads in the lab. Tomorrow’s fast will feel like today’s ultrafast. As I’ve mentioned before, we’re pushing to bring this to as many people as possible. We’ve seen a TON of demand for /ultrafast, made possible by our deep partnership with Cerebras and their unique hardware, which will push the frontiers of tomorrow's ultrafast even further. I’m very much looking forward for the collaboration to continue pushing the absolute limits of how fast we can run our most capable models on Cerebras and bring this to the most demanding customers.
Show more
0
25
1.4K
71
Forward to community
Fast inference makes all things possible. Thanks for building this use case @thegeomaster
Meet the fastest AI design canvas in the world, running on Cerebras at 1000+ tokens/sec. In this demo of Rivermark, I use my voice to create beautiful, on-brand social content and tweak it to perfection in real time, powered by Gemma 4 on Cerebras Inference. This means you can go through dozens of ideas for your visuals and tune them to pixel perfection in the same time it takes other tools to respond to your first prompt. We’re opening early access in the coming weeks, so if you’re a busy designer, founder, or marketer with a high bar for content - reply or DM for a spot. Next up: Qwen3.8-27B is coming to @Cerebras on September 3, and we'll offer day 0 support for it. Excited to build!
Show more
The CS-4 is the fastest AI accelerator in the world, establishing a new roofline for Frontier AI. CS-4 is built to accelerate every part of the AI ecosystem: - Developers who want faster tokens for interactive reasoning and agentic applications. - AI factories require infrastructure that is simpler to deploy and built for hyperscale. - Hyperscalers and Neoclouds need ultra-fast, high-throughput inference to serve the most intelligent AI models in real-time.
Show more
.@cerebras CS-4 expands the frontier of AI compute. Up to 2x faster than CS-3. Up to 10x more token capacity. More speed for every user. More capacity for every operator. CS-4 solutions deliver both, in the same power budget. The fast just got faster.
Show more
What happens when you give the quant trader’s bible to two frontier AI models? The Green Book is what traders use to prepare for firms like SIG and HRT, where I interned. And in high-frequency quant trading, being right isn’t enough. You have to be first. Inside Codex, I extracted all 183 questions and their official solutions. Then I gave GPT-5.6 Sol Ultrafast (via API) on @cerebras and Claude Opus 5 the exact same test, timed every answer, and blind-graded the results. One finished nearly three times faster, and got more answers right.
Show more
Faster agent loops. New systems that were not practical before. Cerebras inference is coming to @CallosumAI APIs. 🟧
We are proud to announce our first major partnership with @CerebrasSystems. Their wafer-scale engines are delivering inference at speeds that are hard to comprehend. Composed within Callosum's platform, that speed becomes something entirely new: heterogeneous multi-agent systems operating in regimes of intelligence that were not possible before.
Show more
Excited to be collaborating on @Cerebras × @OpenAI community events in EMEA. Come join us! 🇩🇪19/8 BER 🇵🇹10/9 LIS 🇧🇪12/9 BXL 🇫🇷17/9 PAR 🇸🇪24/9 STO 🇬🇧1/10 LON 🇮🇪7/10 DUB
Show more