I shipped a fine-tuning tutorial for Muse Glimmer 30B on AI2's MolmoWeb dataset with TRL
> but merve, this model is sota on ScreenSpot-Pro?
yes, but it doesn't work well with ambiguous prompts of MolmoWeb that are close to how you interact with computer, e.g. "jump to nutrition facts"
I compared model fine-tuned on MolmoWeb format against base model zero-shot outputs converted to MolmoWeb format (coordinates on 100)
I found that base model has 13% click accuracy + 35.0% within 5% diagonal while fine-tuned model has 41% + 68% within 5% diagonal