Korea is building the future of AI—fast.
During ICML, we joined our partners at Upstage in Seoul to talk about what ultra-fast inference unlocks. Solar 31B runs at up to 2,000 tokens/sec on the Cerebras Wafer-Scale Engine.
Thanks to everyone who joined us!