We built a Braintrust-native eval in collaboration with Baseten to test whether GLM-5.2 can preserve exact long-context retrieval under production serving constraints.
GLM-5.2's retrieval score is effectively flat as context grows from 25K to 50K, which is the result users most want to see from a sparse-attention long-context model.
Read the full GLM-5.2 eval →