가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Sebastian Raschka
@rasbt
ML/AI research engineer. Ex stats professor. Author of "Build a Large Language Model From Scratch" ( & reasoning (
가입 October 2012
1.2K 팔로잉 중    511K 팬
Some food for thought when designing benchmarks... So, here's a little computer-use (visual) comparison between GPT-5.6 Astra and Qwen3.8 Max. The task here was to recreate the image in the center using the Paint UI. Super interesting how the two different LLMs+Harnesses approached this totally differently by default. I.e., Astra tried to approach this by drawing and layering geometric shapes. Qwen approached this pixel by pixel. (Of course, the pixel-by-pixel result looks closer to the original, it's essentially a low-res version of that by nature.) So, the Qwen-generated image would surely score higher in the sense that it's closer to the original. But I wouldn’t conclude from this example that one LLM generalizes better than the other on other tasks. Also, I wouldn't say Qwen has better compute-use capabilities or better visual understanding than Astra. But it highlights an interesting point about how slippery benchmarks are when they only compare final results.
더 보기