We calculated okayish, but not the greatest results for
@deepseek_ai's V4F-0731 in
@sam_paech's new EQ-Bench v4 benchmark. It is an interesting question whether this is a systematic effect of the new post training of 0731: How does heavy RL affect a model's personality, could it be detrimental?