가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Nikil Ravi
@nikilravi
가입 September 2021
4.3K 팔로잉 중    644 팬
To avoid all doubt about whether this is related to Site Reliability Engineering, we have changed the name to ReverseEngBench:
New cybersecurity benchmark: SRE-Bench🧵 Models are starting to become competent at finding and patching security vulnerabilities in source code, but can they reverse engineer a binary and understand its behavior? Many existing cybersecurity benchmarks test model capabilities given a source codebase. But most of the software that actually matters for security- the systems defenders protect and the malware they inspect- only exists as binaries. This is true on both sides of the threat landscape. Proprietary enterprise software, security appliances, and firmware are frequent attack targets, yet are almost always shipped as binaries. Malware, meanwhile, is deliberately obfuscated to resist inspection. As agents get more capable, binary reverse engineering becomes a critical, under-tested skill. We worked with collaborators at @Columbia, @ucla, @ucberkeley and @tufts to test this, and we're excited to share SRE Bench: a realistic, contamination-free software reverse engineering benchmark for AI agents. It measures whether an agent can RE a binary well enough to actually understand the underlying code's behavior. Initial results show real separation between frontier models, and how far there is to go:
더 보기