Register and share your invite link to earn from video plays and referrals.

Liquid AI
@liquidai
Build efficient general-purpose AI at every scale.
59 Following    36.5K Followers
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
Show more
We're partnering with @Qualcomm to bring personal AI context to devices powered by Snapdragon. Liquid Context, our on-device context layer, is now optimized for Snapdragon processors and runs on the Qualcomm Hexagon NPU. Our goal: Give the agents people choose an understanding of what matters to them and when they need help. With the user's permission, Liquid Context learns from device signals and builds an understanding of their routines, preferences, and needs. That understanding is built and maintained locally, and Liquid Context shares relevant context with the user's chosen agents, whether they run on the device, in the cloud, or across both. That includes third-party agents and Liquid Agent, our efficient embedded agent powered by LFM2.5-2.6B. Running on the Hexagon NPU, Liquid Context works in the background and keeps that understanding current without requiring a cloud model to process every update. For device manufacturers, this is a path to add personal context to their devices while supporting their own choice of agents and services. OEMs building embedded or hybrid agents can also work with us to evaluate Liquid Agent. As our CEO @ramin_m_h said: "Personal AI starts with understanding how you live and what you need, when you need it. Liquid Context builds that understanding on your device so the agents you choose can offer more relevant help and anticipate your needs." > Read more about our partnership: > Check out the livestream of @cristianoamon and Ramin's keynote here:
Show more
In a new article published today on the cover of @CellCellPress, we obtained Liquid Foundation Model instances that establish state-of-the-art performance on biological longevity tasks, outperforming the best frontier models such as Gemini-3.1-Pro, GPT-5, and Claude Opus. In partnership with @InSilicoMeds, we built and released: > A comprehensive eval suite of 17 biological longevity tasks (i.e., LongevityBench), to assess whether a general-purpose language model can interpret aging data spanning clinical records, DNA methylation, transcriptomics, plasma proteomics, and genetic evidence. > LFM2-1.2B-Longevity and LFM2-2.6B-Longevity: two compact models specialized for interpreting structured aging data across these tasks. These results are important! 🧵
Show more
Can a language model master chemical synthesis through deep customization? Our partners @InSilicoMeds tested this out with our LFM2.5-2.6B. See what happened next 👇
Incredible release! About time we had something easily accessible like this. Choose the best model for your needs and iPhone (data only for the 17 Pro for now) then you can try it out in @LocallyAIApp.
Show more
In case you've wondered what model to run on the device you probably use the most
This will be a game changer for local AI
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
Show more
Great to see on-device evaluations! Very proud of LFM2.5-2.6B. The way it generalizes beyond agentic tasks is crazy.
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
Show more
Liquid AI's models are the best for the scaled production deployment - pareto-optimal, beating much larger rivals. That's why we use them at Shopify.
Looking to implement intelligence on edge? Check out @ArtificialAnlys 's new testing results.
Which is the best model to run on an iPhone 17 Pro? Now you'll know: Our team at Liquid AI has partnered with Artificial Analysis to bring you an open-source benchmark to measure on-device quality, speed, latency, and memory.
Show more
This is actually a pretty slick way to approach on device benchmarking. Being able to just pick the models you want and benchmark the whole stack on your hardware is exactly what local AI has been missing. And yeah, the UI is clean af.
Show more
I’m very grateful to have been part of this release. Stay tuned for more 👀! If there’s a model, quantization, runtime, or device you’d like to see in Pipette, let us know:
Show more
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
Show more
It is amazing having such high-caliber folks partnering with you to make sure you are the shaping the frontier of SLMs @ArtificialAnlys @liquidai
🔥 How fast does my MacBook Pro M5 Max run Qwen3.8-27B? 🆕 There's a new benchmarking tool and leaderboard for that. 🎯 Liquid AI + Artificial Analysis released Pipette, an open-source benchmark built for AI running on local hardware. 👇 Pipette benchmarks the inference stack, including ... 🧠 Model 🗜️ Quantization ⚙️ Runtime 💻 Device 📚 Context length So instead of wondering how fast Qwen3.8 is, now you can find out how fast Qwen Q4 is in llama.cpp on my MacBook Pro M5 Max (or other local AI machines). According to AA and Liquid, the dataset currently has ... 📊 10,000+ benchmark results 🧪 1,000+ tested configurations ✅ model × quantization × runtime × device × context 🤖 ~35 model classes 🗜️ 7 quantization levels ⚙️ llama.cpp 📱 iPhone + Android 💻 Mac + Windows PCs 🧠 Strix Halo results coming soon And it measures what matters to Local AI users... ⚡ Decode tps 🚀 Prefill speed ⏱️ Latency 💾 Peak memory 🧠 Model quality ⚠️ Caveat, current benchmark only covers ... ✅ MacBook Pro — M5 Max ✅ iPhone 17 Pro ✅ Samsung Galaxy S26 Ultra 🔗 pipette dot liquid dot ai
Show more
Introducing Pipette ! An open source benchmarking platform for on-device intelligence.
benchmarking on real hardware is infamously challenging; amazing release by my colleagues at @liquidai and @ArtificialAnlys 🫡 *yes upper and left is better😅
Pipette is our new LLM leaderboard that shows not only eval performance but also how fast the models run and how much memory they consume
Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a common, reproducible way to measure the quality, speed, latency, and memory use of AI models on devices such as phones, laptops, PCs, AI boxes, and embedded hardware. > Pipette is open source > In Pipette, models get compated as model + quantization + runtime + device from one interface. > It comes with a warehouse of verified benchmark results, currently with 10k+ results across 35 model classes, 7 quants, llama.cpp runtimes, and 4 devices. > Pipette is a dynamic platform, allowing new contributions from day one, adding new devices, runtimes, model families, and quantization levels. 🧵
Show more