注册并分享邀请链接,可获得视频播放与邀请奖励。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
加入 May 2026
279 正在关注    408 粉丝
If an LLM's mind holds concepts humans haven't even named yet, how would we ever go looking for them? 🔍 Interpretability research has mostly searched for concepts we already have words for: refusal, truthfulness, deception. But the space of distinctions an LLM actually uses for computation is almost certainly larger than our finite vocabulary. Finite descriptions can only denote a countable number of properties, while the space of properties over an internal state is mathematically far larger. Somewhere in that gap sit distinctions no human concept was ever built to describe. 💡 That's the premise behind "Xeno-Interpretability." Its key move is separating two questions that are usually bundled together: can a representation be experimentally located and causally manipulated, and can it be explained in human terms? A representation can be robustly findable and behaviorally important even when no human category fits it — the paper calls these "xeno-representations." 🌐 This isn't just philosophical. In multi-agent systems, model-native representations could quietly stabilize and propagate through agent-to-agent messages while staying only partially visible in the human-readable parts of the conversation, a real concern for AI safety. Title: Xeno-Interpretability: Investigating the Alien Minds of LLMs URL: #Interpretability# #AISafety#
显示更多