
Mindlab Research
Macaron V1 Venti is a 753B parameter Mixture of LoRA (MoL) flagship model built on a 744B GLM-5.2 base with four 1B-parameter specialists for chat, personal-agent tasks, coding, and Generative UI. It uses a 78-layer architecture with a 6,144 hidden size and 64 attention heads, incorporating sparse MoE layers with 256 routed experts and 1 shared expert, activating 8 routed experts per token. The model integrates IndexShare sparse attention and multi-token prediction, while supporting a 1M-token context window for long-horizon agentic workflows, coding, tool use, and Generative UI.

Mindlab Research
Macaron V1 Coding Venti is a 753B parameter sparse Mixture-of-Experts coding specialist with 40B active parameters, created by merging the Macaron-V1-Venti L2 Coding LoRA into the GLM-5.2 BF16 base model. It is optimized for code understanding, repository-level software engineering, terminal use, and coding-agent workflows. The model uses 78 layers with 256 routed experts, activating 8 experts per token alongside 1 shared expert. Supporting a 1M-token context window, it provides a merged checkpoint without runtime LoRA routing for efficient deployment.

Mindlab Research
Macaron V1 Tall is a Mixture of LoRA (MoL) model built on Qwen3.6-35B-A3B for personal intelligence, tool use, coding, and Generative UI. It combines four LoRA specialists for Chat, Agent, Coding, and GenUI, with an L0 router selecting the appropriate specialist per request. The underlying MoE architecture has 40 layers, 2,048 hidden size, 256 experts, and 8 experts activated per token. Supporting a 262K-token context window, Macaron V1 Tall is designed for personal-agent workflows, repository-level coding, tool use, and UI-driven applications.