注册并分享邀请链接,可获得视频播放与邀请奖励。

Rei
@rei_labs
Applied artificial intelligence research • Led by @0xreisearch
加入 November 2024
17 正在关注    16.1K 粉丝
A reward can tell an RL/online learner that something worked without telling it which combination of internal signals made it work. Today, we’re unlocking two learning rules in Adapt-1 Preview: Counterfactual Utility Plasticity (CUP) and Temporal Context Projection (TCP).
显示更多
0
27
283
63
转发到社区