注册并分享邀请链接,可获得视频播放与邀请奖励。

alignedai
@alignedai
applied ai @meta ⁕ cs @ucla ⁕ for the love of the game
加入 June 2026
4.3K 正在关注    585 粉丝
I suspect there are multiple instances of benchmarks missing the mark due to some esoteric config issue.
We implemented the harness with the Responses API and turned on: → Retained reasoning → Context compaction On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.
显示更多