VultronRetrieverPrime Qwen3.5 8B is a multimodal late-interaction retrieval model designed for visual document search and multilingual RAG across PDFs, scans, slides, and reports. It features 8B parameters, using a 32-layer hybrid GatedDeltaNet and full-attention architecture with 4,096 hidden size, 16 attention heads, and 320-dimensional multi-vector embeddings with MaxSim scoring. Supporting up to 262K context and 1,792 visual tokens, it delivers state-of-the-art retrieval accuracy while maintaining a compact index and efficient large-scale serving.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (vultr/VultronRetrieverPrime-Qwen3.5-8B). Released under Apache 2.0.
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure


2,048 GB
1,440 GB
Enter your email to get access to this content
No benchmark data is available for this configuration yet.