"Inoculation Midtraining with Learned Neologisms". They introduce a token that tells the model where misalignment should occur. Also very cool that the authors test vulnerability to conditional misalignment (some triggers still increase misalignment a bit). A nice idea!
Authors: Kyle O'Brien, Edward Young,
@RadmardPuria,
@n_lie_k,
@cam_tice,
@tomekkorbak,
@DavidDAfrica