VultronRetrieverCore Qwen3.5 4.5B is a multimodal late-interaction retrieval model designed for visual document search and multilingual RAG across PDFs, scans, slides, and reports. It features 4.5B parameters, using a 32-layer hybrid GatedDeltaNet and full-attention architecture with 2,560 hidden size, 16 attention heads, and 320-dimensional multi-vector embeddings with MaxSim scoring. Supporting up to 262K context and 1,792 visual tokens, it delivers state-of-the-art 4–5B retrieval performance while balancing retrieval accuracy, latency, and memory efficiency.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (vultr/VultronRetrieverCore-Qwen3.5-4.5B). Released under Apache 2.0.
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure


2,048 GB
1,440 GB
Enter your email to get access to this content
No benchmark data is available for this configuration yet.