Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision
very high benchmarks (beating K3), insane efficiency, and as always amazing tech report
this is probably the most novel arch i've seen in a while, pretty insane