🚀 Introducing our latest research: AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
Continual learning for language agents has not yet been clearly defined. How should we evaluate their ability to continually learn from experience and improve themselves on complex, long-horizon tasks?
- Traditional continual learning provides a useful perspective on the plasticity–stability trade-off, but its formulation does not naturally extend to non-parametric learning paradigms today.
- Many recent works on agent memory still focus on retrieval and reasoning over long (static) contexts, rather than on how agents can reuse experience from complex agentic tasks.
Across coding, deep research, and language understanding and reasoning tasks, we show that carefully designed task streams and metrics are essential for understanding continual learning in language agents.
📄 Paper:
🤗 Dataset: