注册并分享邀请链接,可获得视频播放与邀请奖励。

Jacky Kwok
@jackyk02
Stanford CS PhD | Berkeley EECS
加入 June 2025
1.1K 正在关注    6.2K 粉丝
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: 💻 Code: 🗣️ Discord: 🤗 Data & Models: More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
显示更多
0
181
5.4K
623
转发到社区