Register and share your invite link to earn from video plays and referrals.

Search results for FastInference
FastInference community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including FastInference
SN50 is REALLY FAST. And you don't have to just take our word for it. ⚡️ @SemiAnalysis independently benchmarked SN50 running @MiniMax_AI M2.7 at ~800 t/s. That's consistent with the 800+ t/s we demonstrated at @RaiseSummit, validated by @ArtificialAnlys. Fast inference just got another third-party stamp of approval.
Show more
Grok 4.1 Fast combines frontier tool-calling performance with blazing-fast inference and cost effectiveness.
Korea is building the future of AI—fast. During ICML, we joined our partners at Upstage in Seoul to talk about what ultra-fast inference unlocks. Solar 31B runs at up to 2,000 tokens/sec on the Cerebras Wafer-Scale Engine. Thanks to everyone who joined us!
Show more
$AMD and CEREBRAS $CBRS TO OFFER COMBINED PRODUCT FOR FAST INFERENCE
Qwen 3.8 27B is an extremely good small model It’s perfect for small classifiers and fast inference A drop in replacement for Luna and a significant win for open source
i think someone should build an opposite of @cerebras (sarberec?) instead of crazy fast inference at an even higher price point, make it run the best models more slowly but dirt cheap let's do a poll to test demand. which would you use more?
Show more
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha. When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units. Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents. Skills you'll gain: - Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck - Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals - Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
Show more
0
76
1.1K
113
Forward to community
"If you knew you could get that many tokens, you would build different products." Logan Kilpatrick (@OfficialLoganK ,@GoogleDeepMind) on why fast inference doesn't just make AI faster. It changes what is possible to build. @googlegemma's Gemma 4 is now on Cerebras, running at over 1800 TPS.
Show more
We are entering an extremely exciting era for open-weight models. Kimi K2.6 now feels like a top agentic model. I took it for a spin via @FireworksAI_HQ fast inference APIs. Kimi K2.6 has impressive agentic capabilities, design skills, and the ability to synthesize large amounts of information. I built a little Skill that produces survey papers on any AI research topic you want. (see example in the clip) You can use the skill to tell your agent to generate a survey on whatever topic and watch it go to work. The artifact was fully generated by @Kimi_Moonshot's Kimi K2.6. It's cheap and fast. Next step for me is to explore ways to continue integrating the capabilities of these models on use cases like automating my LLM knowledge bases and augmenting my agent memory capabilities. Stay tuned for more.
Show more
$NVDA incremental upside is coming from a $200B market it has never competed in. (Save this) Jensen's $1T Blackwell and Rubin opportunity is now the base case rather than the bull case. Vera, Nvidia's new CPU, sits outside of that number and management has visibility to ~$20B of CPU revenue this year. Adoption is real too: OpenAI, Anthropic, CoreWeave and Oracle all plan to use Vera, and SpaceXAI announced deployments for Grok. LPX takes this even further, pushing Nvidia into ultra fast inference and the data layer that keeps AI agents running. So, the deeper thinking is: Nvidia can lose some GPU share and still grow its share of the total AI spend through CPUs, networking and storage. That makes the business less dependent on defending one product and strengthens the long term earnings story. So beyond the Q3 guide, I want proof that products beyond the GPU (Vera, LPX, storage, ...) are moving toward incremental revenue. Will be covering the broader implications of NVDA earnings inside Milk Road PRO. Only a couple of hours left before the price goes up:
Show more