Register and share your invite link to earn from video plays and referrals.

Huan Sun
@hhsun1
Prof. @OhioState, endowed CoE Innovation Scholar, advancing the capability and safety/security of LLM-based agents, understanding transformers' limitations
646 Following    6.7K Followers
how do we know an agent is “continually learning”? I think most benchmarks for evaluating continual learning in agents fall short in at least one of the desiderata: 1. Task order. There should be a sequential learning process, and thus tasks should be arranged in a sequential order. 2. Task relationship. What is the relationship between earlier and later tasks? Why are experiences from earlier tasks supposed to help later tasks? What do they share in common beyond the underlying environment? 3. Metrics. If we want to measure “continual learning”, what matters is the performance difference on certain tasks before and after the task stream (or, the sequential learning process). For example, plasticity (and stability) in continual learning should be directly measured as performance difference on newer (and older) tasks before and after the task stream. In our new work (AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents), we propose a more rigorous evaluation setup for continual learning in language agents that considers all these aspects. Check it out!
Show more