Latest ContentInference Cookbook Model Library

Tiny Aya Water

Tiny Aya Water is a 3.35B-parameter multilingual language model optimized for European and Asia Pacific languages. It uses a 36-layer transformer with 2,048 hidden size, 16 attention heads, and 4 KV heads. The architecture interleaves sliding window attention (4,096 window) with periodic global attention layers for efficient long-range interaction. It supports around 8K context length and uses RoPE for positional encoding. Trained across 70+ languages, it is designed for strong regional performance while remaining efficient for local deployment and downstream adaptation.

Type	Dense LLM
Capabilities	Text Generation, Instruction Following, Text Classification, Multilingual
Release Date	17 February, 2026
Links	Paper\|Blog\|HF Model Card
License	CC-BY-NC-4.0

Inference Instructions

Deploy and run this model on NVIDIA B200 GPUs using the command below. Copy the command to get started with inference.

CONSOLE

docker run -it --rm 
 --runtime=nvidia 
 --gpus all 
 --ipc=host 
 --shm-size=128g 
 -p 8000:8000 
 -v ~/.cache/huggingface:/root/.cache/huggingface 
 -e HF_TOKEN='YOUR_HF_TOKEN' 
 vllm/vllm-openai:v0.18.0 
 CohereLabs/tiny-aya-water 
  --tensor-parallel-size 8 
 --max-model-len auto 
  --max-num-batched-tokens 65536 
 --gpu-memory-utilization 0.95 
 --max-num-seqs 1024 
 --trust-remote-code

Note

Serving CohereLabs/tiny-aya-water requires gated model access via Hugging Face.

Model Benchmarks

Each model was tested with a fixed input size and total token volume while increasing concurrency to measure serving performance under load.

Tiny Aya Water

Inference Instructions

Model Benchmarks

ITL vs Concurrency

Time to First Token

Throughput Scaling

Total Tokens/sec vs Avg TTFT

NVIDIA HGX B200

Products

Features

Solutions

Marketplace

Resources

Company

Tech Talks

Vultr Blogs