Register and share your invite link to earn from video plays and referrals.

Andrew Feldman
@andrewdfeldman
CEO and Founder @Cerebras (NASDAQ: CBRS) where we build the fastest AI infrastructure in the world.
217 Following    29.8K Followers
More data centers underway for @cerebras (Who doesn't love massive construction projects?) Build the fastest AI infrastructure. Build data centers around the world to house it. Deliver blisteringly fast tokens to leading customers including @OpenAI . Rinse. Repeat.
Show more
Everyday bravery. There are few things more commendable.
In 2007, a man had a seizure and fell onto the subway tracks in NYC. A train was already pulling in. Autrey, a construction worker, was waiting at a Harlem station with his two little girls, aged 4 and 6. When the young man tumbled off the platform, there was zero time to lift him back up before the train arrived. So, Autrey made a split-second decision. He left his daughters with total strangers on the platform and jumped onto the tracks. He shoved the man into the shallow drainage trench between the rails and threw his own body on top to hold him down. The train barreled into the station and rolled completely over both of them. The clearance was so terrifyingly tight that grease from the train's undercarriage actually smudged Autrey's hat. Miraculously, neither man was hurt. When the media swarmed him afterward, Autrey completely brushed off the praise, simply saying: "I only did what anyone should do." He was later awarded NYC's Bronze Medallion (the city's highest civilian honor) and received a massive standing ovation at the State of the Union. A true, real-life superhero.
Show more
Talk about talent density! Last night, Avenir hosted an AI dinner in Paris with some of our favorite founders and friends @ScottWu46 (Cognition), @andrewdfeldman (Cerebras) @oliveur (Datadog), Chris Clark (OpenRouter), @paraga (Parallel), @nikhilbenesch (TurboPuffer), @dylan522p (the self described shit poster and SemiAnalysis legend), @robertwachen (Etched), @goodwin_ml (Fractile), @annadgoldie (Ricursive), @agermanidis (Runway), @PhilipJohnston (Starcloud), @dan_lahav (Irregular), @G_Princen (Anthropic), @JacobWallenberg (Ramp) and Rohit Iragavarapu @graceisford , @DelahayeHenri @DavidPrilutsky. Great discussion on open vs closed source models, intelligence saturation, the inference explosion, the geopolitics of AI and the future of work
Show more
@andrewdfeldman I had a garden and I'm missing my tomatoes and herbs right now as much as the grounding. Are they also old varieties? I wish a productive season!
.@cerebras is proud to expand our partnership with @Flexintl. Together, we're scaling US production of the CS-3 by an anticipated 7x through 2026. Multiple new assembly and integration lines coming online in Milpitas, California. Liquid cooling installation. High-power integration. Full-rack qualification. Assembled and tested in Silicon Valley. Two highly technical teams building side by side. The fastest AI in the world. Designed and manufactured in the U.S. More coming soon.
Show more
We're expanding our partnership with @Cerebras to scale production of the CS-3, one of the world's most advanced AI accelerator systems, in the heart of Silicon Valley. Together, we're helping power the next wave of AI. 🔗 #AIInfrastructure# $FLEX
Show more
.@cerebras is expanding in Europe. We have plans for 200 megawatts of data center capacity, by the end of 2027. Norway. Finland. France. And we are pursuing more capacity across Europe. We work hard to meet customers where they are. The fastest AI in the world, delivered across the world. More coming soon.
Show more
Fast AI chips enable more and more interesting questions. At @cerebras our Wafer Scale Engine is the fastest AI processor in the world by 15 X. To ask interesting questions, use fast AI.
@cerebras @sarahookr Chips don't just constrain speed - they select which research questions get asked at all. Six years later that filter is still invisible to most people funding the stack.
Most Dad’s fail to play child support because they’re greedy bums. It is your obligation as a father to contribute financially to upbringing of your children. This whole, whether you’re in the household or not.
Show more
This is peak single-mother delusion theater. Your “superhero” mom didn’t protect you. she brainwashed you with 14 years of lies so she could play victim martyr while probably chasing deadbeats and collecting welfare. Real dads leave because the mom is unbearable, and faking child support just teaches kids to simp for frauds instead of facing reality. You handing her $20k? Pathetic l move rewarding her for turning you into another broken man who worships toxic female “sacrifice.” Bet dad dodged a bullet. Single moms create weak sons who glorify their dysfunction. change my mind.
Show more
A man pays his debts.
“There is no Dad money,” she said. “It’s me. Overtime. Second job. Plasma. Everything.” $400/month x 14 years = $67,200. She funded my whole childhood and blamed a ghost. Dad never sent a dollar. She just didn’t want me to hate him. Or feel like we had nothing. I graduated. $74,000 job now. First check? I gave her $20,000. “Child support,” I said. “From your son.” She cried. “I wasn’t owed that.” Yes, you were. For 14 years.
Show more
@cerebras is now running @GoogleDeepMind's Gemma 4 - the leading open-weight multimodal model - at 1,851 tokens per second in public preview. This is 35x faster than a typical GPU endpoint. Cerebras speed also translates into world class latency - Gemma 4 on Cerebras returns its first answer token inclusive of reasoning in 1.5 seconds, making Cerebras the only provider that lets Gemma 4 be used in real-time settings. This is the power of wafer scale.
Show more
True
Moving bits to and from memory. Because of parasitic capacitance+resistance of the wires. The bigger the memory, the longer the wires. The main trick is to organize the memory hierarchically: registers, small on-chip SRAM, caches of various types, and external RAM. It's all because we have to use hardware multiplexing: reusing the same multiply-accumulate unit for multiple parts of the network.
Show more
In hardware, pioneering innovation often leaves you without a supply chain. There are no suppliers, no vendors at the ready, for radically different architectures. We spent 3 years building a heat sink and cooling system. Our wafer is about 58 times larger than the largest GPU. There were no components in any catalogues ready for a wafer scale chip. The heat sinks on the market were built to cool chips 1/58th our size drawing 1/50th the power. We went to every major heat sink vendor. The answers were always the same. “You need a 46,000^m2 heat sink for 15KW of power for a chip?" "We've got nothing for you." So we built it ourselves. It took 3 years. 2016 to 2019. We went through dozens of iterations. JP one of our co-founders and chief system architect insisted we go to water cooling. I thought it was a bad idea. At the time, we knew nothing about liquid cooling. One of the hyperscalers told us categorically that they would never use water to cool. JP convinced me. We delivered one of the first production AI systems to use water cooling. Google announced its TPUs would be water-cooled around the same time. The nay-saying hyperscaler is now nearly 100% water cooled. NVIDIAs leading GPUs are now, 7 years later, water cooled. We became a world leader in liquid cooling - from knowing nothing about it. And we are still the only people who can cool a chip our size. We can run our wafer scale engine cooler than GPUs. And they are therefore more reliable. As heat is the enemy of electronics. When you do radical innovation, there are no vendors waiting to sell you a part. You have to inspire them to build something that's never existed, or do it yourself.
Show more
Yesterday, @OpenAI announced GPT-5.6 Sol - their most powerful model ever. It launches on @Cerebras in July. Frontier intelligence at unprecedented speed. Read about it here:
Show more
Most of the limits we accept aren't real. "Can't be done" just means "hasn't been done yet." @elonmusk set out to build data centers at a pace the whole industry called impossible. Turns out it was "impossible" for everyone but him. We closed a >$20 billion deal with @OpenAI in 4 weeks. Between Thanksgiving and Christmas. The night before Thanksgiving, we signed a term sheet with OpenAI for more than $20 billion. The deal was huge. But it took us just 27 days to close legal, due diligence, and commercial terms. On Christmas Eve, we signed the master agreement. 2 years ago, you would have said a deal like this couldn't be done in that time. 6 weeks after signing, we were in production. One of the best parts about the AI market. Is watching the “it can’t be done-s” crumble in the face of extraordinary execution.
Show more
Today, we reported @cerebras's financial results for Q1 FY 2026. This was our first quarterly earnings report as a public company. We remain laser focused on execution and on delivering the speed and performance our customers need. Read the report:
Show more
One of the biggest constraints in delivering AI today is memory. GPUs rely on a type of DRAM memory called HBM. There are only three primary makers of HBM, @Samsung, @SKhynix, and @MicronTech. They are all sold out. Prices are through the roof. Lead times are now measured in years. And Cerebras doesn’t use it. We avoided this entire tar pit. Its one of the benefits of our wafers scale approach. We use SRAM. It's faster. And it is etched onto the logic wafer. It's not a separate chip. So, there is approximately infinite supply. This is a fundamental architectural advantage. And right now, it allows us to completely avoid the number one constraint in the AI supply chain.
Show more
When you are CEO, turning off is hard. I grow tomatoes. What do you do?
Yesterday, we launched @GoogleDeepMind's Gemma 4 model on @cerebras. The first multimodal model on Cerebras. 1,500 tokens per second. 15x faster than the nearest comparable model. Multimodal agents can see, reason, act, and retry. At 1,500 tokens per second, that loop is becoming near instant. Fast tokens are the most valuable tokens. And Cerebras delivers the fastest tokens in the world.
Show more
Fast tokens are the most valuable tokens. Because inference speed is productivity. And @cerebras makes the fastest tokens in the world. @GoogleDeepMind's Gemma 4 runs at 1,500+ tokens per second on Cerebras — more than an order of magnitude faster than on GPUs and 15x faster than Claude Haiku 4.5. Private preview today. GA later this month. Build blazing fast multimodal apps.
Show more