登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Jan Dubiński
@jan_dubinski_
TruthfulAI | PhD in AI | Warsaw University of Technology and NASK, working on AI safety and generative models
参加 July 2022
1.4K フォロー中    686 ファン
"Inoculation Midtraining with Learned Neologisms". They introduce a token that tells the model where misalignment should occur. Also very cool that the authors test vulnerability to conditional misalignment (some triggers still increase misalignment a bit). A nice idea! Authors: Kyle O'Brien, Edward Young, @RadmardPuria, @n_lie_k, @cam_tice, @tomekkorbak, @DavidDAfrica
もっと見る