๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Liquid AI
@liquidai
Build efficient general-purpose AI at every scale.
๊ฐ€์ž… March 2023
53 ํŒ”๋กœ์ž‰ ์ค‘    34.4K ํŒฌ
Today, we release LFM2.5-VL-3B, a lightweight vision-language model that reads screens, documents, and the physical world. It handles digital screens across mobile, web, and desktop, grounds objects to coordinates, reads text and charts, and calls tools from either text or image input. Built on LFM2.5-2.6B base, with a SigLIP2 400M NaFlex vision encoder > Pre-trained on ~34T tokens > Vocab size: 128K Comparable or better scores compared to models up to 2.6x its size: > ScreenSpot-v2 80.7, ahead of Gemma-4-E4B at 51.2 > RealWorldQA 73.1, ahead of InternVL-3.5-4B at 67.7 > TextVQA 84.3, ahead of Qwen3.5-4B at 81.2 > RefCOCO-avg 87.9, up from 57.1 on LFM2-VL-3B > ToolSandbox 59.5, up from 26.4 on LFM2-VL-3B ๐Ÿงต
๋” ๋ณด๊ธฐ