Register and share your invite link to earn from video plays and referrals.

justin
@justinsunyt
cofounder @capydotai • prev @pennmandt
1.7K Following    8.8K Followers
FDE is dead introducing BDE (Backward Deployed Engineering) instead of deploying an engineer to ship fixes for your customers, let your customers ship the fixes themselves with your coding agent
Show more
actually, all three beliefs are right the less harness, the better: true for frontier models working on isolated single shot tasks. mini-swe-agent outperforms codex and claude on coding benchmarks with a single bash tool, while being significantly cheaper. for tasks that don’t interact with external systems, models benefit more from less context bloat than more tools and instructions. post training wins: true for specific tools and workflows. the RL-ification of models has led them to overfit on tool schemas. hence why opus hallucinates parameters on the pi edit tool, or gpt will prefer native apply_patch. and as the labs invent new approaches to multi-agent orchestration, model behaviors around spawning subagents will branch out too. harnesses provide more value: true if you care about agent behavior around interfaces and systems. as agents evolve and move towards the cloud, we will see infinitely more permutations of agent interfaces and tools. your agent in slack shouldn’t behave the same way as it does in your terminal. some of these behaviors aren’t as easy to tack on as an AGENTS.md or a skill. even more, the labs’ incentives aren’t always aligned with the end user. they want to sell more of their tokens. but the pareto frontier of models today is spread across multiple model families. the best harnesses will take advantage of model capabilities at each price point and route automatically. these three beliefs don’t have to conflict with each other at all. the ideal harness is minimal when it needs to be (deferred context), leans into model post training (model-specific guidance/tools), and brings more value by routing the best models and aligning to their interfaces. (running evals on capy v2 right now and seeing harness optimizations pay off in real time)
Show more
Capy accidentally cracked a graph theory conjecture entirely in Slack REFUTED: Graffiti Conjecture 284 (open for ~30 years) We shared the original post in Slack and Capy (running on Grok 4.5 Medium!!!) decided it wanted to take a shot, finding a novel refutation in 8 minutes
Show more