LFM2.5 VL 450M is a compact vision-language model designed for efficient multimodal understanding and on-device deployment, combining the LFM2.5-350M language model with an 86M SigLIP2 NaFlex vision encoder. It uses a 16-layer hybrid architecture with 1,024 hidden size, 16 attention heads, and 8 key-value heads, alternating short convolution and full-attention layers for efficient sequence processing. Supporting a 32K-token context window, it processes native-resolution images through 512×512 tiling and thumbnail encoding. The model is optimized for multilingual vision understanding, object detection, bounding box prediction, instruction following, and function calling.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (LiquidAI/LFM2.5-VL-450M). Released under LFM 1.0. Further reading: Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure
1,440 GB
Enter your email to get access to this content
Benchmarks measured by Vultr on B200 · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.