註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rei
@rei_labs
Applied artificial intelligence research • Led by @0xreisearch
加入 November 2024
17 正在關注    16.1K 粉絲
A reward can tell an RL/online learner that something worked without telling it which combination of internal signals made it work. Today, we’re unlocking two learning rules in Adapt-1 Preview: Counterfactual Utility Plasticity (CUP) and Temporal Context Projection (TCP).
顯示更多
0
27
283
63
轉發到社區