Really impressive results — self-verification scaling this well with DeepSeek V4 Flash is a big deal for open models.
I recently open-sourced a DeepSeek Harness plugin that implements Best-of-3/5 LLM-as-a-Verifier for coding tasks: isolated candidates (worktrees), auto-test filtering, then the verifier ranks the best patch (with strict double-confirmation before applying anything).
Would love feedback: