Really really happy to have worked on this project!
The key takeaway is not model capability in isolation, but how to close the loop between model iteration and real consumer usage in a domain that is both non-verifiable and inherently preference-driven.
In the work we reframe three core questions:
1.What is Role-Play?
We define Role-play as an agent’s capacity to navigate specific coordinates: {World} × {Stories}, conditioned on {User Preferences}.
do we evaluate it when there is no ground truth answer?
If correctness is subjective, then optimize for not being wrong.
3. How do we iterate model performance in production?
Online preference learning on denoised user signals, A/B testing for validation and iteration.
If you’re thinking about AI entertainment, or online learning in production usage— would love to discuss more!