🤝
@arcee_ai, an American AI lab, built its frontier open model family on NVIDIA Blackwell Ultra.
Its flagship MoE model, Trinity-Large-Thinking, was RL post-trained using NVIDIA NeMo open libraries.
Arcee optimized agentic model inference with NVIDIA Dynamo and vLLM, along with NVIDIA accelerated networking, to deliver low token cost, achieving:
✅ 3T+ tokens served on OpenRouter in its first two months
✅ $0.90 per 1M output tokens
✅ 2nd on PinchBench for open agentic models
Explore Arcee's open models, powered by NVIDIA ➡️