註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jeremy
@1jehuang
building Jcode age 21 Yc s26 hacked github top 500 monkey type
加入 January 2025
100 正在關注    681 粉絲
Today I'm fixing benchmarks with Jcode bench! Jcode bench is the first open and uncontaminatable benchmark. There are a few problems in the coding benchmarks of today: >Private benchmarks are hard to trust >Public benchmarks are easy to benchmaxx >Benchmark task grading can be too coarse or wrong >Benchmarks give little signal to distinguish between the frontier and mid models >Benchmarks are easy to saturate >Benchmarks don't represent the real world coding work. Jcode bench fixes these problems. The layout is like this: Given one reference implementation, optimize it as much as you can. This approach produces a high signal, continuous score over time. Because there is not a known optimal implementation for these tasks, there is no solution answer to train on. If frontier model task transcripts have been trained on, then generate new tasks to spec, and rerun. Transcripts and tasks are able to be audited to do it's open nature. When it's believed that a set of SOTA transcripts have been trained, simply generate a new set of tasks to spec and rerun. spec: results: Individual tasks:
顯示更多
0
11
34
4
轉發到社區