We did a comparison study on where a robot's intelligence should live by using the same model (
@OpenAI GPT6-Astra), same bimanual YAM station, same LEGO pick-and-place task, all run on
@se3labs' real-world remote eval stack:
1. Frozen code-as-policy (built on
@nvidia ENPIRE components): 6/20 placements
2. Agentic control + Astra in the loop: 18/20 placements