注册并分享邀请链接,可获得视频播放与邀请奖励。

Jeremy
@1jehuang
building Jcode age 21 Yc s26 hacked github top 500 monkey type
加入 January 2025
100 正在关注    681 粉丝
Today I'm fixing benchmarks with Jcode bench! Jcode bench is the first open and uncontaminatable benchmark. There are a few problems in the coding benchmarks of today: >Private benchmarks are hard to trust >Public benchmarks are easy to benchmaxx >Benchmark task grading can be too coarse or wrong >Benchmarks give little signal to distinguish between the frontier and mid models >Benchmarks are easy to saturate >Benchmarks don't represent the real world coding work. Jcode bench fixes these problems. The layout is like this: Given one reference implementation, optimize it as much as you can. This approach produces a high signal, continuous score over time. Because there is not a known optimal implementation for these tasks, there is no solution answer to train on. If frontier model task transcripts have been trained on, then generate new tasks to spec, and rerun. Transcripts and tasks are able to be audited to do it's open nature. When it's believed that a set of SOTA transcripts have been trained, simply generate a new set of tasks to spec and rerun. spec: results: Individual tasks:
显示更多
0
11
34
4
转发到社区