Register and share your invite link to earn from video plays and referrals.

Search results for AgentEvals
AgentEvals community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including AgentEvals
TL;DR Swapping an LLM judge for a purpose-built model called Jev in agent evals reportedly delivers 100% accuracy at less than 1/80th the cost. Title: Jev-as-a-Judge for Agent Evals URL: Points ⚖️ It's proposed as a fix for a real dilemma: code-based evals are too rigid, LLM-as-a-judge is too non-deterministic 🎯 Across 500 repeated judgments, Jev hit 100% agreement, while Claude only reached 80.0% 📉 Jev also had the lowest variance in quality scores; other judges were up to 913x more variable 💰 Evaluating 5 requests cost $0.34 with Jev versus $28.17 with Claude — a massive gap ⚡ Average response time was just 0.44 seconds, making it fast as well as cheap 🔍 It's framed as a genuinely "third form" of agent evaluator, alongside code-based checks and LLM-as-a-judge Being able to run judgments cheaply and repeatedly could reshape the whole feedback loop of building agents. #AgentEvals# #LangChain#
Show more
we need agent evals that are really consistent with real world usages. otherwise people are optimizing foundation models for the wrong direction. the problem of targeting is even bigger than benchmaxxing.
Show more
With the latest results from the @rails agent evals, it's clear to see that @openai still has a solid lead, despite Opus 5.5 making a good jump. The dark horse here remains Luna Max. 18% completion at just $11! Not that far off GPT-6 Sol!
Show more
In 13 minutes, @jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs a month. Topics covered: ✅ Their four quadrant eval framework ✅ Why closing the production-to-eval loop is the hardest part ✅ How a data lake and long context changed what agents can do with data
Show more
We finished the Training Agents series. Six live sessions over six months, from evaluating agents to training them inside real environments. All of it is on the Hugging Face YouTube channel and all of the code is open. Here's what we did and who made it happen: 1. Agentic Evaluations WorkshopWhere agent evals actually stand, and why benchmark scores don't match what people see in use. With Avijit Ghosh and Nathan Habib (Hugging Face), Arvind Narayanan (Princeton), Pierre Andrews (Meta), J.J. Allaire (UK AI Security Institute) and Mahesh Sathiamoorthy (Bespoke Labs). 2. RL for Agents Workshop Environments, rollouts, reward design and the inference bottlenecks that appear when you move from RL for LLMs to RL for agents. With Lewis Tunstall (Hugging Face), Will Brown (Prime Intellect), Ofir Press (Princeton) and Alex Zhang (MIT CSAIL). 3. Training Agents 1: SFT on agent traces Public coding-agent traces turned into prompt/completion data, a TRL + LoRA fine-tune on Hugging Face Jobs, metrics in Trackio, and an honest look at what the first eval numbers can and cannot tell you. Joined by Sergio Paniego and Quentin Gallouédec. 4. Training Agents 2: Distillation Off-policy, on-policy and self-distillation for moving capability from a teacher into a smaller coding agent. 5. Training Agents 3: Reinforcement learning GRPO after SFT: group sampling, verifiable reward functions, reading the reward/KL/length curves, and three experiments, one of them with a deliberately gameable reward so we could watch the hacking happen. 6. Training Agents 4: From reward functions to environments The reward stops being a function and becomes a place the agent acts in. We walked the reset()/step() contract from Gym to LLM agents, built an OpenEnv environment and pushed it to the Hub, plugged it into TRL's GRPOTrainer, then trained a real coding agent (OpenCode) through Harbor with AsyncGRPOTrainer on Hugging Face sandboxes. The series has passed 300k views. Thank you to every speaker, to the TRL team, and to everyone who showed up live with questions. Playlist:
Show more
0
46
1.6K
233
Forward to community
What the hell happened in AI today? > frontier models got cheaper > restricted models got more accessible > Chinese models got better > secret models showed up for free Full recap: MYTHOS 5 → Anthropic opened its security scans to every Claude Enterprise customer. SOL → OpenAI just cut API pricing for the next 3 months by over 20% DEEPSEEK → V4 Flash got vision while keeping basically the same text intelligence, apparently approaching Opus 4.8 on some multimodal agent evals OX ALPHA → mysterious 1M-context multimodal model appears FREE for a week with insane capacity. nobody knows who the hell made it NVIDIA → AVO hits 100% on the ARC-AGI-3 PUBLIC SET (I repeat, PUBLIC SET) using Opus 5
Show more