注册并分享邀请链接,可获得视频播放与邀请奖励。

Jan Dubiński
@jan_dubinski_
TruthfulAI | PhD in AI | Warsaw University of Technology and NASK, working on AI safety and generative models
加入 July 2022
1.4K 正在关注    686 粉丝
"Inoculation Midtraining with Learned Neologisms". They introduce a token that tells the model where misalignment should occur. Also very cool that the authors test vulnerability to conditional misalignment (some triggers still increase misalignment a bit). A nice idea! Authors: Kyle O'Brien, Edward Young, @RadmardPuria, @n_lie_k, @cam_tice, @tomekkorbak, @DavidDAfrica
显示更多