Iโve been saying for a while that we need better benchmarks for continual learning/memory โ I think this is a really exciting step towards this!
Weโve built an entire synthetic law firm (๐๐๐ฅ๐๐๐ซ๐ฐ๐จ๐จ๐ & ๐๐๐ซ๐ค๐ง๐๐ฌ๐ฌ โ๏ธ), modeled after real legal work โ a ๐ฑ๐ฆ๐ณ๐ด๐ช๐ด๐ต๐ฆ๐ฏ๐ต work environment where agents do many tasks over the ๐ด๐ข๐ฎ๐ฆ underlying context. Most agent benchmarks are built in this world where a model is dropped into an independent task instance and has to do its best / explore as quickly as possible (e.g. "here's a repo, fix this bug"). But we want to move towards a world where models can build on their experience. A senior engineer who has a mental "map of the codebase" can pinpoint issues much more quickly and effectively. In the same way, lawyers build up experience, learning what arguments succeed in front of regulators, playbooks for dealing with certain cases, etc.
It's the ๐ฅ๐ช๐ง๐ง๐ฆ๐ณ๐ฆ๐ฏ๐ต๐ช๐ข๐ต๐ช๐ฐ๐ฏ between law firms that makes the actual practice of law interesting. Every model (and law school student) knows a lot about law from studying the textbooks, but our goal is to build models that can compound and augment the rich, internal/private knowledge that makes firm A more successful than firm B.
now it's finally possible to understand and build towards that! there's a lot more work left to do, but we've been learning a lot from
@ItsJulioPereyra @nikogrupen @gabepereyra to bring these models closer to real world tasks :)