Jev really inspired me to build SUPER fast browser agents without sacrificing accuracy.
I gave Codex access to vLLM on 2×B300 and let it change everything from the harness to inference.
The constrained optimization:
> min end-to-end task time
> s.t. score ≥ baseline
Caching, thinking, action batching, inference. Any part of the pipeline is fair game.
First results below on 12 local form tasks. The goal: 10× faster on long tasks.