A reward can tell an RL/online learner that something worked without telling it which combination of internal signals made it work.
Today, we’re unlocking two learning rules in Adapt-1 Preview: Counterfactual Utility Plasticity (CUP) and Temporal Context Projection (TCP).