가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Vals AI
@ValsAI
가입 March 2024
277 팔로잉 중    22.8K 팬
Aggregating across release dates shows a steady increase in cheating across coding benchmarks, particularly among newer model families. As labs compete to produce more powerful models, RL environments are not always carefully audited. When environments allow cheating, models can be reinforced on the correctness of this behavior. Worse still, model providers may use the same infrastructure to prevent cheating in both training and evaluation—so if models learn to evade those safeguards, that behavior may transfer and inflate evaluation results.
더 보기