Register and share your invite link to earn from video plays and referrals.

Amanda Huang
@amandaH_333
LLM MLE @MiniMax_AI(ex-quant @RBC Capital Market, ex-Tiktoker) Role-play&Character training &Agent Climber-free solo someday not your typical algo guy
135 Following    254 Followers
Once collect human preferences, train a Reward Model to imitate them, then run RL against that RM. It's weak because the RM is only a proxy for the real objective — a giant net imitating "vibes" — so it has adversarial examples, and if you optimize too long the model games it (spamming "The the the" for near-perfect scores). Yet it still helps: it cashes in the generator–discriminator gap (judging answers is far easier for humans than writing ideal ones) and curbs hallucinations.
Show more
# RLHF is just barely RL Reinforcement Learning from Human Feedback (RLHF) is the third (and last) major stage of training an LLM, after pretraining and supervised finetuning (SFT). My rant on RLHF is that it is just barely RL, in a way that I think is not too widely appreciated. RL is powerful. RLHF is not. Let's take a look at the example of AlphaGo. AlphaGo was trained with actual RL. The computer played games of Go and trained on rollouts that maximized the reward function (winning the game), eventually surpassing the best human players at Go. AlphaGo was not trained with RLHF. If it were, it would not have worked nearly as well. What would it look like to train AlphaGo with RLHF? Well first, you'd give human labelers two board states from Go, and ask them which one they like better: Then you'd collect say 100,000 comparisons like this, and you'd train a "Reward Model" (RM) neural network to imitate this human "vibe check" of the board state. You'd train it to agree with the human judgement on average. Once we have a Reward Model vibe check, you run RL with respect to it, learning to play the moves that lead to good vibes. Clearly, this would not have led anywhere too interesting in Go. There are two fundamental, separate reasons for this: 1. The vibes could be misleading - this is not the actual reward (winning the game). This is a crappy proxy objective. But much worse, 2. You'd find that your RL optimization goes off rails as it quickly discovers board states that are adversarial examples to the Reward Model. Remember the RM is a massive neural net with billions of parameters imitating the vibe. There are board states are "out of distribution" to its training data, which are not actually good states, yet by chance they get a very high reward from the RM. For the exact same reasons, sometimes I'm a bit surprised RLHF works for LLMs at all. The RM we train for LLMs is just a vibe check in the exact same way. It gives high scores to the kinds of assistant responses that human raters statistically seem to like. It's not the "actual" objective of correctly solving problems, it's a proxy objective of what looks good to humans. Second, you can't even run RLHF for too long because your model quickly learns to respond in ways that game the reward model. These predictions can look really weird, e.g. you'll see that your LLM Assistant starts to respond with something non-sensical like "The the the the the the" to many prompts. Which looks ridiculous to you but then you look at the RM vibe check and see that for some reason the RM thinks these look excellent. Your LLM found an adversarial example. It's out of domain w.r.t. the RM's training data, in an undefined territory. Yes you can mitigate this by repeatedly adding these specific examples into the training set, but you'll find other adversarial examples next time around. For this reason, you can't even run RLHF for too many steps of optimization. You do a few hundred/thousand steps and then you have to call it because your optimization will start to game the RM. This is not RL like AlphaGo was. And yet, RLHF is a net helpful step of building an LLM Assistant. I think there's a few subtle reasons but my favorite one to point to is that through it, the LLM Assistant benefits from the generator-discriminator gap. That is, for many problem types, it is a significantly easier task for a human labeler to select the best of few candidate answers, instead of writing the ideal answer from scratch. A good example is a prompt like "Generate a poem about paperclips" or something like that. An average human labeler will struggle to write a good poem from scratch as an SFT example, but they could select a good looking poem given a few candidates. So RLHF is a kind of way to benefit from this gap of "easiness" of human supervision. There's a few other reasons, e.g. RLHF is also helpful in mitigating hallucinations because if the RM is a strong enough model to catch the LLM making stuff up during training, it can learn to penalize this with a low reward, teaching the model an aversion to risking factual knowledge when it's not sure. But a satisfying treatment of hallucinations and their mitigations is a whole different post so I digress. All to say that RLHF *is* net useful, but it's not RL. No production-grade *actual* RL on an LLM has so far been convincingly achieved and demonstrated in an open domain, at scale. And intuitively, this is because getting actual rewards (i.e. the equivalent of win the game) is really difficult in the open-ended problem solving tasks. It's all fun and games in a closed, game-like environment like Go where the dynamics are constrained and the reward function is cheap to evaluate and impossible to game. But how do you give an objective reward for summarizing an article? Or answering a slightly ambiguous question about some pip install issue? Or telling a joke? Or re-writing some Java code to Python? Going towards this is not in principle impossible but it's also not trivial and it requires some creative thinking. But whoever convincingly cracks this problem will be able to run actual RL. The kind of RL that led to AlphaGo beating humans in Go. Except this LLM would have a real shot of beating humans in open-domain problem solving.
Show more
Interesting ~ The value of a harness isn’t to preserve more context for the model, but to normalize a large out-of-distribution task into a sequence of smaller, in-distribution ones. This makes a lot of sense from a training perspective: harnesses for specialized domains often converge on similar patterns!
Show more
Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize through composition. We observe a powerful property when training RLMs: for tasks with shared structure that look different, the root model naturally learns the same trajectory, meaning it views the two task trajectories as the same! In other words, the Transformer does not need additional generalization capabilities to transfer capabilities from one task to the other, the harness induces it. We find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token! Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model. For example, RLMs can see problems of different lengths as the same: we show that RLMs can train exclusively on short tasks, and fully generalize to similar but unseen tasks 8-32x longer because it produces near identical trajectories for both. Taking this further, we show that tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect. RLMs can train on the problem of finding which essays belong to the same author and improve performance on finding math problems that share similar solutions. The full blogpost, experiments, and discussion are in the thread below.
Show more
Opened Cursor again today and was struck by how similar coding-agent interfaces are becoming. It made me wonder where lasting differentiation will come from. My guess: not just better models or harnesses, but deeper adaptation to how each person works. General features spread quickly; products that truly learn their users may build more durable loyalty—and perhaps pricing power.
Show more
Real self-correction isn’t asking a model to “think again.” It’s building a generate–verify–repair loop grounded in independent evidence, structured feedback, and hard stop conditions. Its effectiveness depends far more on the quality of the verifier and external feedback than on swapping models or adding more agent roles
Show more
I’ve recently been exploring continual learning for agents. The key question is how to evaluate agent optimizers—are they learning generalizable coding strategies, or merely overfitting to the current benchmark
Show more
There’s something exciting about kicking off an evening project with a whole team of agents by your side!😉
Welcome back! ✨🐣
hello world✨👋 finally sending my first post. (my previous X got hacked be careful of scammers)
Beyond measuring autonomous capability, evals should capture the diverse ways people interact and achieve outcomes with AI.
Today we share the worldview behind our mission. Human values don't average out. Local knowledge can't be centralized. The good future has many AIs, raised in different places, shaped by the people they serve, disagreeing with each other the way we do.
Show more
A strong training and evaluation harness raises the ceiling for what models can achieve.
new post on harness engineering for AI self-improvement: It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple. Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
Show more
🧵 Deli AutoResearch SKILL is now officially open source! 🎉 Alongside it, we’re dropping our 4th survey paper — this time on Self-play. Inspired by AlphaZero, we got a powerful insight: prior knowledge doesn’t always lift the ceiling. Models can discover more globally optimal solutions just by playing against themselves. The biggest change in this paper? For the first time, the AutoResearch Agent autonomously planned GPU experiments — and submitted actual RL runs on the DeepSeek 285B model. The entire RL pipeline — experiment design, code writing, running, debugging, and conclusion summarization — was 100% automated, with zero human intervention from me. This was incredibly difficult, but an incredibly important step. GRPO is the tool being called by the AutoResearch Agent here. We see this as the beginning of our Continual Learning research journey. 🚀 As always, this is my personal research project, unaffiliated with any organization. All views are my own. #AI# #ReinforcementLearning# #SelfPlay# #OpenSource# #AutoML# #ContinualLearning# #DeepSeek#
Show more
0
28
1.3K
211
Forward to community
Discovering the right question is more important than finding answer
Yesss! We packed this much capability into a 428B model.🫡
MiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters Weights: MiniMax Sparse Attention:
Kudos to the team🫡
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: Token Plan: 🚀New! MiniMax Code: Weights & Tech Report in ~10 Days
Show more
What's fascinating about building a companion harness: watching how an AI learns to perceive😉😌.
Ushering in the M3 multimodal agentic era😊😊
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: Token Plan: 🚀New! MiniMax Code: Weights & Tech Report in ~10 Days
Show more
Really really happy to have worked on this project! The key takeaway is not model capability in isolation, but how to close the loop between model iteration and real consumer usage in a domain that is both non-verifiable and inherently preference-driven. In the work we reframe three core questions: 1.What is Role-Play? We define Role-play as an agent’s capacity to navigate specific coordinates: {World} × {Stories}, conditioned on {User Preferences}. do we evaluate it when there is no ground truth answer? If correctness is subjective, then optimize for not being wrong. 3. How do we iterate model performance in production? Online preference learning on denoised user signals, A/B testing for validation and iteration. If you’re thinking about AI entertainment, or online learning in production usage— would love to discuss more!
Show more