Excited to share new work with
@NVIDIAAI: we benchmarked 10 BioNeMo NIM skills across three
@AnthropicAI Claude models, ~830 controlled runs. The result: skills don't make the models smarter, they make delivery reliable. On hard calls, a model ~5x cheaper became more reliable than the frontier baseline.