Finding an AI agent is easy. Deciding whether it fits a real task is harder.
This guide shows how turns public agent data and live audit evidence into a practical selection workflow.
An Agent Card shows what an AI agent claims. checks what actually works.
See how to audit a public agent and review its protocols, capabilities, security checks, and evidence.
AI agents need to be tested before they can be trusted.
Delegate $ZKP to power the Auditor network. Your delegation increases its capacity to audit and verify more agents across the ecosystem.
Help build the trust layer for autonomous machines.
Delegate on
zkPass began by making facts from private data verifiable without revealing the data itself.
Now we’re extending that principle to AI agents.
Introducing where provenance, audit evidence and observed behavior form a public, inspectable trust record.
so the benchmark penalizes refactoring, and opus 5 refactors more at high reasoning. that's not a bug, it's a personality. frontiercode wants a model that follows orders, not one that thinks.
We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria that reflect user experience.
One such verifier is a scope criterion, which penalizes modifications to the codebase beyond what the task requires. We observe this effect across all frontier models we evaluated, though it is most pronounced in Opus 5: at higher reasoning efforts, the model shows a stronger tendency to refactor code unprompted.
Open source is always vital for science and technology to advance, and is particularly important for the study of (artificial) Intelligence, which is still at its primitive stage. At least, open source and open knowledge help fight against human ignorance and arrogance.
🎙️This past week, @AnnaRRose & @nico_mnbl spoke w Alex Ozdemir, Assistant Professor at Georgia Tech, about the formal verification/ZK connection — covering SMT solvers, Lean, zkPi and why verifiable software matters.
openai shipping tiered reasoning at 25x cost spread. the real benchmark is not the top score, it's the cheapest token that can run a 10 step agent loop. that number just dropped hard.
GPT-5.6 is a major step forward for health intelligence.
Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less.
Together, these advances raise quality while making advanced models accessible to more people globally.
Base is building the secure and trusted infrastructure for global onchain finance. As more payments, credit, assets, and markets move onchain, verifiable data becomes a critical bridge between real-world financial activity and programmable blockchain rails.
ERC-8004 already has 620k+ registered agents.
The question is no longer whether there will be enough agents. It’s which ones are worth calling. Registration is permissionless. Usage won’t be.
A ten-thousand word monster post trying to cover the entire tech tree behind the main lineage of obfuscation (iO) protocols:
Special thanks to all who helped!