註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Liquid AI
@liquidai
Build efficient general-purpose AI at every scale.
加入 March 2023
53 正在關注    34.4K 粉絲
Today, we release LFM2.5-VL-3B, a lightweight vision-language model that reads screens, documents, and the physical world. It handles digital screens across mobile, web, and desktop, grounds objects to coordinates, reads text and charts, and calls tools from either text or image input. Built on LFM2.5-2.6B base, with a SigLIP2 400M NaFlex vision encoder > Pre-trained on ~34T tokens > Vocab size: 128K Comparable or better scores compared to models up to 2.6x its size: > ScreenSpot-v2 80.7, ahead of Gemma-4-E4B at 51.2 > RealWorldQA 73.1, ahead of InternVL-3.5-4B at 67.7 > TextVQA 84.3, ahead of Qwen3.5-4B at 81.2 > RefCOCO-avg 87.9, up from 57.1 on LFM2-VL-3B > ToolSandbox 59.5, up from 26.4 on LFM2-VL-3B 🧵
顯示更多
0
50
1.3K
165
轉發到社區