working on AGI alignment. prev: GPT-Neo, the Pile, LM evals, RL overoptimization, scaling SAEs to GPT-4, interp via circuit sparsity. EleutherAI cofounder.
neolab onboarding be like:
openai, which was founded to be the good guys, ended up just racing to the bottom on safety, by its own hand.
anthropic, which tried to be better, also failed and ended up similarly racing.
now it is our turn to be the good guys.
mr capabees, I'm afraid to inform you that your creation, "number go up machine 3000 megacreative turbogoodharting unmonitorable edition" has made number go up in an...unexpected manner