Fine-grained concurrency sweep from 500 to 1,000 concurrent requests (step 50) to identify the exact saturation knee for each model. Each concurrency level was tested across 3 independent runs with 200 requests per level.
All models are fully saturated by 500 concurrent requests. Operating beyond 750 concurrent provides no throughput benefit and only increases tail latency. For production deployments, target the 200-500 range for optimal throughput-to-latency tradeoff.
0 Comments
Be the first to comment and share your perspective with the community.