Register and share your invite link to earn from video plays and referrals.

Yuhang Yao
@yuhang_yao
Senior Research Scientist @ZOOM | PhD @CarnegieMellon | Graph Agent Research
117 Following    1K Followers
Grok 4.6 just topped the VISTA Leaderboard 👑 🥇 Grok 4.6 — 0.552 🥈 GPT-5.6-sol — 0.538 🥉 fable-5 — 0.533 VISTA tests how well coding agents turn Figma designs into functional web apps—not traditional coding problems. Grok 4.6 improves significantly over 4.5’s 0.517. One caveat: 4.6 ran at High effort, while 4.5 ran at Medium, so the longer runtime is expected. AI coding is moving beyond “does it run?” to “does it look right and actually work?” Models are shipping faster than ever. Looking forward to Grok 4.7—and more challengers. 🚀 #Grok46# #AICoding# #VISTA#
Show more
🚀 MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale How can we make small models stronger for on-device agent deployment? MERA uses stronger models to guide an iterative loop of RL/GRPO, skill learning, and router optimization. Student failures become verified demonstrations, reusable SkillBook procedures, and LoRA updates, helping the small model take on more work over time. 🔥 Results: Qwen2.5-Coder-1.5B: 28.7% → 49.7% coding pass Qwen3.5-2B on TAU-2: 14/35 → 18/35 Fine-tuned 2B matches an unadapted 4B model Don’t just route around small models. Evolve them. 📄 💻
Show more
🚀 MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale What if small models could improve themselves, not just be routed around? MERA evolves both small-model skills and weights, guided by stronger models. Student failures become verified teacher demonstrations, reusable SkillBook procedures, and LoRA updates. 🔥 Results: Qwen2.5-Coder-1.5B: 28.7% → 49.7% direct pass on coding tasks Qwen3.5-2B on TAU-2: 14/35 (40.0%) → 18/35 (51.4%) The adapted 2B model matches an unadapted 4B model Small models don’t just get routed around. They get better. 📄 Paper: 💻 Code:
Show more
📣 KDD 2026 is next week! Check out our Blue Sky Ideas work on why stronger models and bolted-on guardrails are not enough for trustworthy multi-agent systems. 🤖 Our GPT and Claude demos reveal failures caused by agent composition, semantic misalignment, data re-identification, and operational loops. Trust must be baked in—not bolted on. 🔐 Demo: Paper: GitHub:
Show more
Grok 4.5 (Med) vs fable-5 (High) on VISTA 🏆 Score: 0.517 vs 0.533 ⚡ Tokens: 161K vs 757K 💰 Cost: $1.83 vs $12.04 Nearly the same performance, far lower cost.
Finally received the X money card.
🔥 Grok 4.5 ranks No. 1 on the VISTA Leaderboard! 🚀With Cursor Agent, Grok 4.5 achieves a 0.367 combined score, outperforming GPT-5.6-sol and Fable-5. ⚡It only uses half the tokens of GPT-5.6-sol and one-fifth of Fable-5. Leaderboard:
Show more
GLM-5.2 is now just 0.005 away from Fable-5 in visual spec-to-app development. On VISTA’s latest C4 leaderboard: 🥇 Claude Code + Fable-5 — 0.274 🥈 Claude Code + GLM-5.2 — 0.269 🥉 Claude Code + Opus 4.8 — 0.263   Claude Code + Sonnet 4.6 — 0.248   Cursor + Composer 2.5 — 0.212   Codex + GPT-5.5 — 0.205 The story is no longer just “which model writes better code.” For real app-building agents, performance is increasingly shaped by the full stack: 🧠 model 🛠️ harness 🔁 tool loop 👁️ visual understanding ⚙️ reasoning setup The harness is becoming part of the model. 🔗
Show more
Excited to share that TwinRouterBench has been accepted to the #RLEval# Workshop at #CAIS2026# 🎉 As LLM apps become long-horizon agents, one request can trigger many model calls across planning, tool use, retrieval, coding, and verification. That makes per-step LLM routing a core infrastructure problem: sending each call to the cheapest sufficient model without breaking downstream success. TwinRouterBench introduces: ⚡ Static track: 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench 🚀 Dynamic track: live SWE-bench Verified evaluation with official task resolution + realized API spend Key result: a router trained on static labels achieves comparable SWE-bench resolve rate while cutting API cost by ~53% vs. an unrouted Opus 4.6 baseline. Paper: Code: Dataset: Website: #LLM# #AgenticAI# #LLMRouting# #Benchmark# #SWEBench#
Show more
Thrilled to share that our vision paper “Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On” has been accepted to the KDD Blue Sky Idea Track! 🎉✨ Trust in A2A networks should be baked in from the start, not patched on afterward. 🤖🔗🔐 📄 #KDD# #TrustworthyAI# #AIAgents# #LLM#
Show more
Agentic RL.
Introducing RL Commons. An open research initiative for the reinforcement learning era. Shared compute. Open benchmarks. Collaborative work. Founding phase Project Aster is now open.
Show more
Interesting discussion. We explore the trust problem in agent ecosystems here:
OpenClaw 2026.3.8 🦞 🔒 ACP provenance — your agent finally knows who's talking to it 💾 openclaw backup — because YOLO deploys need a safety net 📱 Telegram dupes killed 🛡️ 12+ security fixes We fixed more things than we broke. Progress.
Show more