Terminal-Bench 2.1 pass@1 run is scored 76.40% with Luna High using Pydantic AI adapter of my Agent SDK.
Next one will have same experiment with LangChain adapter and we'll see the difference.
@samuelcolvin@pydantic Lately I'm powering my filesystem simulator VSH with Monty ( to track the side effects of filesystem operations and review changes before running them (using Monty's filesystem mounts), and currently working on its Pydantic AI capability.