We’re growing our evaluation team at
@thinkymachines. We care a lot about whether models are actually useful and whether our measurements are good enough to serve as the y-axis for scaling research.
We’ll work across a wide range of problems and support several workstreams, including building signal-bearing internal evals for predictive scaling, turning real user feedbacks and product workflows into evals that close usability gaps, auditing graders and harnesses, and developing new benchmarks for customizability.
If you want to work on this, apply here: