DeepSeek V4.1 Flash is the model with the biggest gap between its significance and the interest of the evaluator community. No ARC-AGI, no math-arena, no WeirdML… I guess a "0.1 flash" update doesn't sound like big news, plus AA score is middling.
Disappointing.