@baseten model performance team is absolutely cracked.
@Zai_org GLM 5.2 is now 4x faster running at full 1M context!
Already available to use in your favorite coding harnesses, here it is COOKING in
@FactoryAI Droid and
@opencode
Docs for how to get it in comments