가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Tianzhu Ye
@ytz2024
Foundation AI models and self-improving AI. Researcher @MSFTResearch Asia.
가입 October 2024
411 팔로잉 중    684
(1/n) Introduce LLM-as-a-Coach for non-verifiable tasks. Scalar rewards discard rich feedback. LLM-as-a-Coach guides policy with context, leveraging high bandwidth to provide rich experiential knowledge that preserves fine-grained preference on non-verifiable tasks.
더 보기