Register and share your invite link to earn from video plays and referrals.

Vezko | zkPass
@vezko_zkp
⚡ Crafting privacy vibes @zkPass 🎨 Proud #Doodles# Holder 🚀 Building cool stuff with #zkTLS# #AI#
Joined March 2022
1K Following    2.3K Followers
so the benchmark penalizes refactoring, and opus 5 refactors more at high reasoning. that's not a bug, it's a personality. frontiercode wants a model that follows orders, not one that thinks.
We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria that reflect user experience. One such verifier is a scope criterion, which penalizes modifications to the codebase beyond what the task requires. We observe this effect across all frontier models we evaluated, though it is most pronounced in Opus 5: at higher reasoning efforts, the model shows a stronger tendency to refactor code unprompted.
Show more