登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

DAIR.AI
@dair_ai
Democratizing AI research, education, and technologies. Learn about AI Agents for FREE at
参加 July 2017
1 フォロー中    129K ファン
On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the question enterprise deployments care about. If you ship skills files, policy documents, or long system prompts, you have been trusting that they actually bind agent behavior. But how are you measuring all of this? Surge AI built a benchmark to actually check this. HANDBOOK.md places a standard operating procedure of 20 to 124 pages in context and grades whether it governed every action across an extended tool-use horizon. 65 tasks, five domains, ten fictional companies. Each task runs in a self-contained company environment with a file workspace plus mock email, chat, calendar, issue-tracking, and commerce services exposed over MCP. Every task mutates one of ten base handbooks, altering the specific rules and thresholds that grading turns on, so memorization does not help. Grading is fully deterministic and two-sided. 824 programmatic criteria check that required actions occurred and that prohibited actions did not. Paper: Learn to build effective AI agents in our academy:
もっと見る