9/
This is a joint effort across Columbia, UCLA, Tufts, UC Berkeley, and @ValsAI, with a lot of support from @DAP__Lab.
Our current paper draft is here (and will likely continue to evolve -- since we are actively evaluating new models):
Something to share: we built a benchmark for evaluating agentic reverse engineering.
- instead of grading intermediate outputs, e.g., recovered types/names
- we evaluate agents e2e on deterministic goals, e.g., JTAG a firmware to enable its debug mode.