--- title: NVIDIA LiteLLM Router emoji: 🚀 colorFrom: green colorTo: blue sdk: docker pinned: false license: apache-2.0 --- # NVIDIA LiteLLM Router Auto-routes across 31 free NVIDIA NIM models with latency-based routing, tier selection, and automatic failover. ## Usage ```bash curl https://tomoritemitopex-nvidia-litellm-router.hf.space/v1/chat/completions \ -H "Authorization: Bearer sk-litellm-master" \ -H "Content-Type: application/json" \ -d '{"model": "nvidia-auto", "messages": [{"role": "user", "content": "hello"}]}' ``` ### Model Groups | Model | Description | |-------|-------------| | `nvidia-auto` | Fastest across ALL models | | `nvidia-coding` | Fastest coding model | | `nvidia-reasoning` | Fastest reasoning model | | `nvidia-general` | Fastest general model | | `nvidia-fast` | Fastest small/efficient model | | `` | Direct access (e.g. `deepseek-v3.2`, `kimi-k2-instruct`) | ### Features - Latency-based routing picks the fastest model automatically - 429 rate limit? Retries 3x with backoff, then failover - Model slow? Deprioritized automatically - Model down? 60s cooldown, auto-recovers