
Liquid AI
LFM2.5 8B A1B is a compact hybrid Mixture-of-Experts model designed for on-device personal assistants, tool use, instruction following, and agentic workflows. It features 8.3B total parameters with 1.5B active, using 24 layers that combine 18 double-gated convolution blocks with 6 GQA attention blocks. The architecture employs a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and 32 experts with 4 activated per token. Supporting a 128K-token context window, it is trained on 38T tokens for efficient, high-throughput multilingual inference across ten languages and diverse edge devices.

Liquid AI
LFM2.5 2.6B is a compact hybrid model designed for tool use, instruction following, multi-step agentic tasks, and efficient on-device deployment. It features 2.69B parameters across 30 layers, combining 22 double-gated short convolution blocks with 8 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 10,752-dimensional SwiGLU feed-forward network. Trained on 34T tokens with agentic reinforcement learning, it supports a 131K-token context window and multilingual text generation across 15 languages for efficient long-context agentic workloads.

Liquid AI
LFM2.5 1.2B Thinking is a compact hybrid model designed for general-purpose reasoning and efficient on-device deployment. It features 1.2B parameters across 16 layers, combining 10 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 12,288-dimensional SwiGLU feed-forward network. Trained on 28T tokens with extended pre-training and multi-stage reinforcement learning, it supports a 32K-token context window and multilingual text generation across English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.

Liquid AI
LFM2.5 1.2B Instruct is a compact hybrid model designed for general-purpose language tasks and efficient on-device deployment. It features 1.2B parameters across 16 layers, combining 10 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 2,048-dimensional hidden size, 32 attention heads, 8 KV heads, and a 12,288-dimensional SwiGLU feed-forward network. Supporting extended context through positional scaling, it is optimized for fast edge inference and operates under 1GB of memory. It supports a 32K-token context window and multilingual text generation across eight languages.

Liquid AI
LFM2.5 350M is a compact hybrid model designed for general-purpose language tasks and efficient on-device deployment. It features 350M parameters across 16 layers, combining 10 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 1,024-dimensional hidden size, 16 attention heads, 8 KV heads, and a 6,656-dimensional SwiGLU feed-forward network. Trained on 28T tokens through extended pre-training and multi-stage reinforcement learning, it supports a 32K-token context window and multilingual text generation across English, Arabic, Chinese, French, German, Japanese, Korean, Portuguese, and Spanish.

Liquid AI
LFM2.5 230M is an ultra-compact hybrid model designed for general-purpose language tasks, tool use, data extraction, and efficient on-device deployment. It features 230M parameters across 14 layers, combining 8 double-gated convolution blocks with 6 GQA attention blocks. The architecture uses a 1,024-dimensional hidden size, 16 attention heads, 8 KV heads, and a 2,560-dimensional SwiGLU feed-forward network. Distilled from LFM2.5-350M and refined through multi-stage reinforcement learning, it supports a 32K-token context window and multilingual text generation across English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, and Spanish.