Register and share your invite link to earn from video plays and referrals.

Liquid AI
@liquidai
Build efficient general-purpose AI at every scale.
56 Following    35.7K Followers
Incredible release! About time we had something easily accessible like this. Choose the best model for your needs and iPhone (data only for the 17 Pro for now) then you can try it out in @LocallyAIApp.
Show more
In case you've wondered what model to run on the device you probably use the most
Great to see on-device evaluations! Very proud of LFM2.5-2.6B. The way it generalizes beyond agentic tasks is crazy.
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
Show more
Liquid AI's models are the best for the scaled production deployment - pareto-optimal, beating much larger rivals. That's why we use them at Shopify.
Looking to implement intelligence on edge? Check out @ArtificialAnlys 's new testing results.
Which is the best model to run on an iPhone 17 Pro? Now you'll know: Our team at Liquid AI has partnered with Artificial Analysis to bring you an open-source benchmark to measure on-device quality, speed, latency, and memory.
Show more
This is actually a pretty slick way to approach on device benchmarking. Being able to just pick the models you want and benchmark the whole stack on your hardware is exactly what local AI has been missing. And yeah, the UI is clean af.
Show more
I’m very grateful to have been part of this release. Stay tuned for more 👀! If there’s a model, quantization, runtime, or device you’d like to see in Pipette, let us know:
Show more
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
Show more
It is amazing having such high-caliber folks partnering with you to make sure you are the shaping the frontier of SLMs @ArtificialAnlys @liquidai
🔥 How fast does my MacBook Pro M5 Max run Qwen3.8-27B? 🆕 There's a new benchmarking tool and leaderboard for that. 🎯 Liquid AI + Artificial Analysis released Pipette, an open-source benchmark built for AI running on local hardware. 👇 Pipette benchmarks the inference stack, including ... 🧠 Model 🗜️ Quantization ⚙️ Runtime 💻 Device 📚 Context length So instead of wondering how fast Qwen3.8 is, now you can find out how fast Qwen Q4 is in llama.cpp on my MacBook Pro M5 Max (or other local AI machines). According to AA and Liquid, the dataset currently has ... 📊 10,000+ benchmark results 🧪 1,000+ tested configurations ✅ model × quantization × runtime × device × context 🤖 ~35 model classes 🗜️ 7 quantization levels ⚙️ llama.cpp 📱 iPhone + Android 💻 Mac + Windows PCs 🧠 Strix Halo results coming soon And it measures what matters to Local AI users... ⚡ Decode tps 🚀 Prefill speed ⏱️ Latency 💾 Peak memory 🧠 Model quality ⚠️ Caveat, current benchmark only covers ... ✅ MacBook Pro — M5 Max ✅ iPhone 17 Pro ✅ Samsung Galaxy S26 Ultra 🔗 pipette dot liquid dot ai
Show more
Introducing Pipette ! An open source benchmarking platform for on-device intelligence.
benchmarking on real hardware is infamously challenging; amazing release by my colleagues at @liquidai and @ArtificialAnlys 🫡 *yes upper and left is better😅
Pipette is our new LLM leaderboard that shows not only eval performance but also how fast the models run and how much memory they consume
Say hello to Pipette! Our benchmarking platform for on-device AI models
This one was long overdue in the local AI community! A benchmarking platform for on-device intelligence that is 1) independently validated, 2) community driven evaluation components, and 3) taking into account all aspects of local intelligence that do not show up on the cloud! It was a pleasure working with the amazing @ArtificialAnlys team on this project. ✨ enjoy
Show more
Today we release Pipette, a model evaluation suite for on-device intelligence, in partnership with @ArtificialAnlys Most benchmarking platforms are optimized to measure core capabilities and speed profile of foundation models served in the cloud. Pipette gives the field a common, reproducible way to measure the quality, speed, latency, and memory use of AI models on devices such as phones, laptops, PCs, AI boxes, and embedded hardware. > Pipette is open source > In Pipette, models get compated as model + quantization + runtime + device from one interface. > It comes with a warehouse of verified benchmark results, currently with 10k+ results across 35 model classes, 7 quants, llama.cpp runtimes, and 4 devices. > Pipette is a dynamic platform, allowing new contributions from day one, adding new devices, runtimes, model families, and quantization levels. 🧵
Show more
0
29
721
111
Forward to community
Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality. A lightweight draft model proposes a block of candidate tokens and the target model verifies them in a single forward pass. Across MATH500, GSM8K, HumanEval, MBPP, and MT-Bench at batch size 1: > Up to 3.18x throughput on an H100: LFM2.5-8B-A1B on MATH500, 428 → 1362 tok/s > Up to 2.87x on an M4 Max MacBook Pro: LFM2.5-1.2B-Instruct on HumanEval, 136 → 389 tok/s > LFM2.5-2.6B means: 2.67x on the H100 (323 → 864 tok/s), 2.27x on device (61 → 139 tok/s) > Under greedy decoding, the emitted sequence is identical to baseline by construction, so benchmark accuracy is unchanged. Each draft model is around 300M parameters, with embedding and LM head tied to its target model. The gain shows most in agentic workloads, where the model reasons before every tool call and the user waits through it all: on BFCL multi-tool scenarios, DSpark cuts LFM2.5-2.6B latency by nearly 50% on average. 🧵
Show more