I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh...
gpt-5.6-sol: 52%
fable: 65%
whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%)
I am very confused