Training VLMs to use vision-only inputs to play games is not just limited to Anthropic.
We showed this was possible using Qwen3-VL-Instruct-8b prior to Fable 5 beating pokemon firered. It is great to see a scaled up version in the latest Claude release
Wondering how VLMs can be trained to play games using only visual inputs, like Anthropic’s newly released Claude Fable 5?
Check out our recent work, Odysseus:
In Odysseus, we train VLMs to play games directly from visual inputs, using Super Mario Land as a testbed, and scale RL to improve their long-horizon decision-making capabilities. Excited to see more exploration in this direction!