We just open-sourced 10 RL environments for M&A due diligence at
@harvey.
These are some of the longest-context knowledge work evals out there, with 80M tokens of context and 100-1,000 rubric criteria per task.
Deep dive by
@ItsJulioPereyra and
@nikogrupen: