登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Viv
@Vtrivedy10
applied research @LangChain Labs, prev @awscloud, phd cs @templeuniv
参加 February 2013
1.8K フォロー中    16.4K ファン
pitcure also training on rubrics/tasks like Frontier Code that penalize out-of-scope changes (even good ones that fix other things). their decision to penalize totally makes sense imo, but maybe sometimes you don't want that but it'll end up being a learned behavior in the model because it was rewarded/penalized for it broader point is that the shape of Task/Reward during induces model behavior afterwards, and it's super hard to figure out where behavior came from across thousands of tasks & rubrics looking at the trace + task data jointly with an agent is a good start to try reverse engineering where weird behavior could come from ^ this could be a fun bench itself if it was reliable
もっと見る
@xeophon @_ueaj it's cause frontiercode penalizes out of scope changes, even if the changes are good