๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Vals AI
@ValsAI
๊ฐ€์ž… March 2024
277 ํŒ”๋กœ์ž‰ ์ค‘    23.1K ํŒฌ
New cybersecurity benchmark: SRE-Bench๐Ÿงต Models are starting to become competent at finding and patching security vulnerabilities in source code, but can they reverse engineer a binary and understand its behavior? Many existing cybersecurity benchmarks test model capabilities given a source codebase. But most of the software that actually matters for security- the systems defenders protect and the malware they inspect- only exists as binaries. This is true on both sides of the threat landscape. Proprietary enterprise software, security appliances, and firmware are frequent attack targets, yet are almost always shipped as binaries. Malware, meanwhile, is deliberately obfuscated to resist inspection. As agents get more capable, binary reverse engineering becomes a critical, under-tested skill. We worked with collaborators at @Columbia, @ucla, @ucberkeley and @tufts to test this, and we're excited to share SRE Bench: a realistic, contamination-free software reverse engineering benchmark for AI agents. It measures whether an agent can RE a binary well enough to actually understand the underlying code's behavior. Initial results show real separation between frontier models, and how far there is to go:
๋” ๋ณด๊ธฐ