Most agent memory benches score recall.
StateMemBench scores whether the answer uses the current fact or a superseded one.
234 multi-session scenarios. 322 graded probes. Closed-pool grading for state drift.
StateMem lifts current-state accuracy about 1.8x over the strongest same-backbone memory baseline on DeepSeek-V4-Flash (0.199 to 0.363).