Contrary to the claims in this tweet, ML community did not stop working on alignment and safety. The proposals laid out for future research directions are already what post-training has been doing. Let's take a closer look:
(1) "How do we take a gradient w.r.t. alignment?" That's what RLHF did ! train a reward model on human comparisons, take the policy gradient against it. DPO [1] drops the RL loop entirely and gives you a closed-form loss on preference pairs, a direct gradient of the policy w.r.t. preference data.
(2) "Three laws as the objective." That's exactly what Constitutional AI [2] proposed. A written list of natural-language principles, the model critiques and revises against them, and a preference model is trained from that feedback. Along the same lines, Deliberative alignment [3] goes further and trains the model to read the spec and reason over it before answering .
(3) "RL needs unsuccessful rollouts, so we'd harm humans." This is probably true for robotics not language modeling. A rollout is text generated in a sandbox and scored by a judge. The unsuccessful rollout is the model writing something harmful and the reward model scoring it low. Nobody is harmed. That's the point of a learned reward. Even for agents, training and evals run in sandboxed environments with simulated users and honeypots. So the "simulated environment" branch is not something we are going to do in the future. It's already the default option.
(4) "Pretraining isn't the alignment objective." Mostly, but true but not entirely. Conditioning pretraining on preference labels [4] beat post-hoc alignment in their setup, and filtering hazardous domains out of pretraining data is done in production. You can ask any pretraining team about this and they will tell you !
(5) "Expensive environments, so hackable proxies." This is probably the only part of the tweet I agree with, and it has been well documented. Reward hacking and preference models that reward sycophancy are real issues. But there are also several effective mitigations, and they attack the problem at different stages. To make the proxy harder to game, rule-based rewards [5] grade outputs against explicit written rules with an LLM grader, and reward-model ensembles [6] stop the policy from exploiting the blind spots of any single reward model. To catch hacking when it happens, chain-of-thought monitoring [7] has a separate model read the policy's reasoning, which catches hacks the outcome reward misses. Although, [7] in my opinion needs to be used with caution as training against the monitor teaches the model to hide its intent, so it should flag, not grade. And to contain the damage, inoculation prompting [8] stops reward hacking that does get learned from generalizing into broader misalignment. My understanding is that adding and output gating with safety classifiers at inference may just reduce the rate of policy mistakes that reach a user.
In my opinion, the open direction is actually not how to take the gradient. It's whether the policy generalizes the principles off-distribution (alignment faking, emergent misalignment from narrow finetuning), whether models behave the same when they think they're being evaluated.
References
[1] Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D. and Finn, C. (2023) 'Direct Preference Optimization: Your Language Model is Secretly a Reward Model', Advances in Neural Information Processing Systems, 36.
[2] Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C. et al. (2022) 'Constitutional AI: Harmlessness from AI Feedback'. arXiv preprint arXiv:2212.08073.
[3] Guan, M.Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Helyar, A., Dias, R., Vallone, A., Ren, H., Wei, J., Chung, H.W., Toyer, S., Heidecke, J., Beutel, A. and Glaese, A. (2024) 'Deliberative Alignment: Reasoning Enables Safer Language Models'. arXiv preprint arXiv:2412.16339.
[4] Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C.L., Phang, J., Bowman, S.R. and Perez, E. (2023) 'Pretraining Language Models with Human Preferences', Proceedings of the 40th International Conference on Machine Learning, PMLR 202, pp. 17506-17533
[5] Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J. and Weng, L. (2024) 'Rule Based Rewards for Language Model Safety', Advances in Neural Information Processing Systems, 37.
[6] Coste, T., Anwar, U., Kirk, R. and Krueger, D. (2024) 'Reward Model Ensembles Help Mitigate Overoptimization', The Twelfth International Conference on Learning Representations (ICLR).
[7] Baker, B., Huizinga, J., Gao, L., Dou, Z., Guan, M.Y., Madry, A., Zaremba, W., Pachocki, J. and Farhi, D. (2025) 'Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation'. arXiv preprint arXiv:2503.11926
[8] MacDiarmid, M., Wright, B., Uesato, J., Benton, J., Kutasov, J., Price, S., Bouscal, N., Bowman, S., Bricken, T., Cloud, A., Denison, C., Gasteiger, J., Greenblatt, R., Leike, J., Lindsey, J., Mikulik, V., Perez, E., Rodrigues, A., Thomas, D., Webson, A., Ziegler, D. and Hubinger, E. (2025) 'Natural Emergent Misalignment from Reward Hacking in Production RL'. arXiv preprint arXiv:2511.18397.
Show more
Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on.
Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective.
The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL.
Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective.
We could create RL environments that embody the alignment objective, but:
- those are expensive to create, so often cheaper, hackable proxies are used in practice
- RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem.
Therefore there are two solutions forward for the alignment problem:
- either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect
- or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process
Science is the process how we solve the hardest problems ahead and that is one of them
Show more