๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Jacky Kwok
@jackyk02
Stanford CS PhD | Berkeley EECS
๊ฐ€์ž… June 2025
1.1K ํŒ”๋กœ์ž‰ ์ค‘    6.2K ํŒฌ
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9ร— faster inference than Jev โšก while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. ๐Ÿ“„ Blog: ๐Ÿ’ป Code: ๐Ÿ—ฃ๏ธ Discord: ๐Ÿค— Data & Models: More details on CLMโ€™s architecture, data recipe, and scaling laws in the thread below ๐Ÿงต
๋” ๋ณด๊ธฐ