Ling 3.0 Flash Fin is a finance-enhanced Mixture-of-Experts model designed for financial research, valuation, spreadsheet workflows, source-grounded analysis, and long-horizon agentic tasks. It features 124B total parameters with 5.1B activated per token, using a 42-layer architecture with 2,560 hidden size, 32 attention heads, and 32 KV heads. The architecture employs a hybrid KDA and attention design, 512 routed experts with 8 activated per token, and 1 shared expert, alongside a native Multi-Token Prediction layer. Supporting a 256K-token context window, it is optimized for multi-document financial reasoning, tool-intensive workflows, and professional research outputs.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (inclusionAI/Ling-3.0-flash-Fin). Released under MIT.
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure
1,440 GB
Enter your email to get access to this content
Benchmarks measured by Vultr on B200 · SGLang. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.