TwiScan
Hot
Communities
Account collections
Login
Register
English
日本語
한국의
简体中文
繁体中文
Register and share your invite link to earn from video plays and referrals.
Register now
Xudong Han
@Xudong07452910
🎓PhD ing @ University of Sussex | LLM & AI Agents 📖Building Personal AI Agents | Vibe Coding 🛠️分享AI 技术解析 | AI 应用心得 📩DM for collab
Joined November 2020
563
Following
9.6K
Followers
Xudong Han
@Xudong07452910
2026.08.08 04:00
Agent 出错以后,到底该修模型,还是修 harness? Scale AI 这篇《Model or Harness?》专门研究 Agent 的故障定位。 今天的 Agent 已经是一个复杂系统:模型之外,还有 Context、Memory、Tool、Grader、用户和运行环境。最后看到的失败结果相同,背后的原因可能完全不同。 比如 Agent 忽略了一条早期指令。 可能是上下文压缩把它删掉了,需要改 harness;也可能信息一直都在,只是模型没有正确使用,这时才该训练模型。 论文整理了 41 类常见失败,并把每个问题定位到具体的「组件交互 + 责任方」。 这样一次失败就能进一步回答: 该做模型 post-training,改 context / memory / tool,还是重新设计环境和评测。 作者还让不同前沿模型独立给失败案例分类,最好的结果与人工标注达到 Cohen’s κ = 0.76,说明这套分类有一定一致性。 我觉得这篇很适合现在的 Agent 发展阶段。 随着 harness 越做越复杂,「Agent 失败了」这个结论已经太粗了。 真正有用的 debugging,要继续追到第一处无法恢复的错误,再决定到底该修哪一层。 📎 arxiv:
Show more
0
0
83
446
76
Forward to community
Most Popular Users
New York Post
@nypost
4.1M Followers
Serenity
@aleabitoreddit
1M Followers
billboard
@billboard
16.5M Followers
オリコンニュース
@oricon
1.9M Followers
Reuters
@Reuters
26.3M Followers
First Squawk
@FirstSquawk
556.5K Followers
BTS_official
@bts_bighit
45.3M Followers
Hello! Project Link
@UpFrontLink
15.5K Followers
一劍浣春秋
@chee828
0 Followers
Elon Musk
@elonmusk
0 Followers
Pirat_Nation 🔴
@Pirat_Nation
342K Followers
BLACKPINKOFFICIAL
@BLACKPINK
11.6M Followers
i-dle (아이들)
@official_i_dle
2.4M Followers
吴说区块链
@wublockchain12
0 Followers
INI
@official__INI
510.1K Followers