Really excited to see
@humansand training models from their long-horizon interactions with people to make AI more human-aligned.
We strongly share this vision, which is why we built SWE-Together, a benchmark reconstructed from 11K+ real user–agent coding sessions that evaluates agents not just on whether they finish the task, but on how well they stay aligned with user intent.
If interaction with people is the training signal, interaction quality should be the eval.
Would love to see models trained with this recipe tested on SWE-Together 🚀
📄
💻