가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Jan Dubiński
@jan_dubinski_
TruthfulAI | PhD in AI | Warsaw University of Technology and NASK, working on AI safety and generative models
가입 July 2022
1.4K 팔로잉 중    686 팬
"Inoculation Midtraining with Learned Neologisms". They introduce a token that tells the model where misalignment should occur. Also very cool that the authors test vulnerability to conditional misalignment (some triggers still increase misalignment a bit). A nice idea! Authors: Kyle O'Brien, Edward Young, @RadmardPuria, @n_lie_k, @cam_tice, @tomekkorbak, @DavidDAfrica
더 보기