LFM2.5-VL-3B: Liquid's Faster Edge Vision-Language Model
Liquid AI LFM2.5-VL-3B (~Aug 12, 2026): ~3.1B open-weight edge vision-language model, non-reasoning VL, LFM 1.0 license, self-reported screen/UI and grounding gains. Weights on Hugging Face.

At a Glance
- ID / HF:
lfm-2.5-vl-3b/LiquidAI/LFM2.5-VL-3B - Org: Liquid AI
- Params: ~3.1B multimodal VL
- Style: Non-reasoning (direct answers)
- Base: LFM2.5-2.6B text base + SigLIP2 400M NaFlex vision encoder
- Pretrain: ~34T tokens (family); 4x vision pretrain vs prior VL-3B (vendor)
- License: LFM 1.0 /
lfm1_0 - Runtimes: llama.cpp (GGUF), MLX, vLLM, SGLang, ONNX
- Vendor on-device: ~228 tok/s M5 Max; ~116 tok/s Ryzen AI Max+ 395; ~3 GB; ~20 tok/s Galaxy S26 Ultra class
What's New vs LFM2-VL-3B
- Screen/UI: ScreenSpot-v2 avg 80.7 (vendor; vs weak prior VL-3B screen scores).
- Function calling (new to VL line): ToolSandbox 59.5 (from 26.4); BFCL v4 32.5 (from 20.5).
- Grounding: RefCOCO-avg precision@1 87.9 (from 57.1).
- Multi-image: BLINK 61.5 (from 50.2); MuirBench 58.3 (from 34.9).
Selected vendor vision scores (self-reported)
| Benchmark | LFM2.5-VL-3B | Prior LFM2-VL-3B | Note |
|---|---|---|---|
| ScreenSpot-v2 (avg of splits) | 80.7 | (near-zero prior) | Vendor |
| RefCOCO-avg | 87.9 | 57.1 | P@1 avg |
| RealWorldQA | 73.1 | 71.1 | Vendor |
| DocVQA (val) | 91.1 | 89.8 | Vendor |
| ChartQA (test) | 81.3 | 80.4 | Vendor |
| BLINK | 61.5 | 50.2 | Multi-image |
| ToolSandbox | 59.5 | 26.4 | Text tool use |
| BFCL V4 | 32.5 | 20.5 | Function calling |
Compared in-blog to Gemma 4 E2B/E4B, InternVL 3.5 2B/4B, Qwen3.5 2B/4B under Liquid's harness notes. Not LLM Stats verified.
When to Use It
Good fit: On-device UI agents, document/OCR loops, grounding + tool use at the edge, WebGPU / phone demos where non-reasoning latency matters.
Not automatic: Heavy chain-of-thought VL; knowledge-heavy or coding-first workloads (prefer larger / reasoning models). Text-only agents may still prefer lfm-2.5-2.6b.
Caveats
- License is LFM 1.0, not Apache/MIT. Read terms before commercial redistribution.
- All headline numbers above are vendor-reported (non-reasoning eval mode).
- "Open-weight without restrictions" marketing language still sits under LFM license terms.
Sources
Questions
Frequently Asked Questions
- Liquid AI's ~3.1B open-weight vision-language model (~Aug 12, 2026) for edge / on-device use. Non-reasoning VL (direct answers for lower latency) on the LFM2.5-2.6B text base + SigLIP2 encoder. Hugging Face:
LiquidAI/LFM2.5-VL-3B. - Liquid's LFM 1.0 license (catalog
lfm1_0; custom; not Apache/MIT). Read terms before commercial redistribution. No. Headline vision scores (ScreenSpot-v2, RefCOCO-avg, DocVQA, ChartQA, BLINK, ToolSandbox, BFCL v4, and others) are vendor-reported in non-reasoning eval mode. Not LLM Stats verified.
- Good fit: on-device UI agents, document/OCR loops, grounding + tool use at the edge, WebGPU / phone demos where non-reasoning latency matters. Not automatic: heavy chain-of-thought VL or knowledge/coding-first workloads; text-only agents may still prefer
lfm-2.5-2.6b. Vendor on-device figures: ~228 tok/s on M5 Max; ~116 tok/s on Ryzen AI Max+ 395; ~20 tok/s Galaxy S26 Ultra class; about ~3 GB.
Continue Reading
