Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate.
When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average.
For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% is baseline, a model winning and losing equally often.