Register and share your invite link to earn from video plays and referrals.

Etched
@Etched
Frontier Inference Clusters
3 Following    51.1K Followers
We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix. Our mission is to run the world's inference. This round accelerates production of our inference clusters. We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.
Show more
0
177
5.5K
348
Forward to community
Vertically integrating the full inference hardware stack is extremely hard. Hear @UbertiGavin and @robertwachen share three years of battle stories leading into today's launch
Show more
We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.
Show more
Introducing Cluster-Scale Memory (CSM) for low latency workloads. Today's AI chips using HBM can’t achieve SRAM-level decode speeds due to memory subsystem and interconnect bottlenecks. SRAM-only chips have lower FLOPs density and memory capacity, sacrificing throughput. You’re forced to make a tradeoff: serve at much slower speeds, or run at low batch sizes and suffer from higher costs. When running large MoE models, token routing across experts requires sending data through a deep memory hierarchy and a networking switch to reach a destination expert. Each memory layer inherently adds latency; thus, the best layer is no layer. We’ve designed a new architecture that creates a shared low-latency memory pool across the entire scale-up domain. We use a proprietary ultra-low-latency, high-bandwidth interconnect to enable dramatically faster memory access across chips. Our HBM/SRAM hybrid design solves both memory capacity and mem2mem latency, enabling high throughput and interactivity simultaneously. CSM improves latency and avoids today's cost, reliability, yield, thermal, and compute tradeoffs of SRAM-only chips, 3D DRAM chips, or optics.
Show more
We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.
Show more
0
654
9.8K
942
Forward to community
Thousands of people in the line 😅 adding compute to handle the load! Go to to try it yourself
You can play Oasis here: Learn more about the model and partnership here: