one thing i think people dont appreciate enough about
@poolsideai is their unusual degree of openness — not only have they shipped an excellent Small model that somehow beat
@thinkymachines at coding, but most people (like
@eliebakouch) have been shouting out their excellent papers, but also they're among a rare few to actually expose their full eval dataset as well - beautifully published, with across 6 public benchmarks with 4 runs each and hundreds of turns per run. you can satisfy for yourself if they rewardhack. brilliant.