注册并分享邀请链接,可获得视频播放与邀请奖励。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
加入 February 2026
1.7K 正在关注    12.4K 粉丝
Maybe we're seeing a closer academic approach toward the Meta-RL endpoint.
(1/n) Introduce LLM-as-a-Coach for non-verifiable tasks. Scalar rewards discard rich feedback. LLM-as-a-Coach guides policy with context, leveraging high bandwidth to provide rich experiential knowledge that preserves fine-grained preference on non-verifiable tasks.
显示更多
0
2
99
10
转发到社区