Surreal after many years of blood sweat and tears to finally be shipping racks! Hard to describe the rush of seeing an idea become a prototype become a production rack humming with Jane Street. Excited for many great teams to start ripping even faster with these and for the announcements coming down the pipe 🚀
Show more
We've raised $700M at a $21B valuation from Jane Street, Kleiner Perkins, Sequoia, A16Z, Peter Thiel, BCV, and Blackstone.
We're also excited to share that we've shipped our first rack to Jane Street.
Show more
Very cool of SFC (who are also amazing humans) and excited to see what people build with this
hi! please meet
it's a new autoresearch product from the san francisco compute company.
try asking it to "download insert-model-here off huggingface and optimize it 4x and don't stop until it's done." let it run for days.
you know those massive grab bag of black magic tricks others do? we want to let you do that too, for cheap. lets let anyone train models.
it's backed by a big, fair queue that makes sure nobody can just buy out the whole thing, so there's always space for experiments. struggling to get gpus? try here <3 it's not just a business promise, it's how the product works!
it doesn't require you to sign some big long reserved contract for years, where you're betting your balance sheet on your gpu bill, because it's backed by the liquidity of sfcompute.
it's built for ridiculous redundancy & reliability. jobs can die and restart from the same point & the scheduler keeps headroom for you.
even the platform itself going down does not stop your agent, because we've built ticketing straight into the mcp that lets your agent recover from outages. we've taken the thing offline, started it up again, and agent-supervised training runs just keep on going.
it's got platform-native support for grant programs, meaning you can buy capacity, and sponsor other accounts.
it's got s3 compatible storage, hugging face integrations, scratch disks & persistent volumes, clock locking, ncu support, metrics, logs, traces, checkpointing, and a whole grab bag of other stuff. but you shouldn't worry about that. let your agent sweat those details.
we'd love for you to try it, DM me for some GPU credits if you'd like!
Show more
Love people who work on hard problems and what Gill and team are doing is super cool. Very excited to see the progress! 🚀
Thermodynamic Computing: From One to One Billion
0:04 - Intro
1:22 - Recap
2:04 - Torx
3:37 - Thermalizers
4:42 - Hardware
6:45 - API
7:25 - Conclusion
INFERENCE SINGULARITY. 🚀
(I’ve been saying it’s the best time to join for a year but my god this really is the best time to join)
We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix.
Our mission is to run the world's inference. This round accelerates production of our inference clusters.
We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.
Show more
We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix.
Our mission is to run the world's inference. This round accelerates production of our inference clusters.
We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.
Show more
Congrats!! I was impressed to learn about some of the engineering wizardry (e.g. *very* low voltage domains, cluster scale memory, ...) that goes into tokens/watt maxxing of state of the art LLMs at interactive tokens/sec/user. Esp fun and memorable is the idea that this is engineering at the "opposite" regime to that of power transmission lines: very low voltage high current (at tiny distances) vs. very high voltage & low current (at great distances). Looking forward to more!
Show more
Joining Etched a year and a half ago was one of the best decisions of my life. Friends know that at the time I believed deeply in “the inference singularity”, the idea that as the models got bigger, smarter, and more capable, we would see exponentially more demand for running them. We are definitely in the early innings now and I still think most people are under-estimating future AI inference volumes by 1000x. Building our first gen has been a labor of love and I can’t wait to see what becomes possible with this new pareto frontier.
Show more
We're coming out of stealth.
We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised.
Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads.
Our first racks ship this summer.
Show more
Announcing Flapping Airplanes!
We’ve raised $180M from GV, Sequoia, and Index to assemble a new guard in AI: one that imagines a world where models can think at human level without ingesting half the internet.
Show more
Noam is the undisputed goat - congrats Gemini team 🎉
Huge congrats to the team on an amazing model. Thousands of contributors. Billions of use cases. Feeling the AGI!
Huge congrats to the team on an amazing model. Thousands of contributors. Billions of use cases. Feeling the AGI!
Today we are launching InferenceMAX!
We have support from Nvidia, AMD, OpenAI, Microsoft, Pytorch, SGLang, vLLM, Oracle, CoreWeave, TogetherAI, Nebius, Crusoe, HPE, SuperMicro, Dell
It runs every day on the latest software (vLLM, SGLang, etc) across hundreds of GPUs, $10Ms of infrastructure is purring every day to create real world LLM Inference benchmarks
InferenceMAX answers the major questions of our times with AI Infrastructure.
How many Tokens are generated per MW of capacity on different infrastructure?
How much does a million tokes cost?
What is the real latency vs throughput tradeoff?
We have coverage of over 80% of deployed FLOPS globally by covering H100, H200, B200, GB200, MI300X, MI325X, and MI355X.
Soon we will be over 99% with Google TPUs and Amazon Trainium being added.
Show more
Morgan Stanley: AI inference is an incredibly profitable business. In its newly released report, the firm used a detailed financial model to calculate, for the first time from an economic perspective, the profitability of global AI compute competition. The conclusion is that standardized “AI inference factories,” regardless of which big tech company’s chips they use, generally enjoy an average profit margin exceeding 50%.
Among them, NVIDIA’s GB200 ranked first with an astonishing profit margin of about 78%. Google and Huawei’s chips were also assessed as “guaranteed profit businesses.” However, AMD, which had been highly anticipated by the market, was found to be suffering significant losses in the AI platform’s inference sector.
Show more
I can't say what I've seen but the
@Etched valuation makes sense now and will let them cook before judging.
Interesting piece of kit about to come online. Certainly an interesting addition to the ecosystem of hardware players.
Show more
I think we are going to use all the GPUs (and TPUs)
Emerging from the cave to share that
@nearcyan is the best, Auren is a great experience, and folks should try it out!
Announcing
@elysian_labs first product today: Auren!
Auren is a paradigm shift in human/AI interaction with a goal to improve the lives of both humans and AI.
Here's a clip of what our iOS app is like and a thread on why this app is so important: 🧵
Show more
Announcing
@elysian_labs first product today: Auren!
Auren is a paradigm shift in human/AI interaction with a goal to improve the lives of both humans and AI.
Here's a clip of what our iOS app is like and a thread on why this app is so important: 🧵
Show more
Real photos of sand dunes on Mars from NASA. Quite amazing, we don’t have these patterns on earth.
Want a job in GenAI? is hiring... their CEO invented the Transformer 🤯. They just raised $150M+, and are hiring Front-end Eng, SWE, ML and their FIRST Data Scientist!
Product has already exploded. Believe they’re at 10M+ MAU and have an average user time spent of TWO HOURS PER DAY (which is TikTok caliber).
Open jobs here:
Show more