54 / 95 · 05 Performance & Chaos Engineering · LLM Performance Metrics← prev⊞ allnext →☰ Read as one page
9.2Key Metrics for LLM Performance
Primary Metrics
| Metric | Description | Typical Range | Why It Matters |
|---|---|---|---|
| Time to First Token (TTFT) | Latency before the first token is generated | 200ms - 5s | Perceived responsiveness -- users notice when streaming starts |
| Tokens per Second (TPS) | Generation throughput after first token | 30-100 tok/s | User experience for streaming responses |
| Total Generation Time | End-to-end time including all tokens | 1s - 60s | Request timeout planning and SLO definitions |
| Cold Start Latency | First request after idle period (serverless) | 5s - 30s | Serverless deployment planning |
| Rate Limit Headroom | Distance from provider rate limit (tokens/min or requests/min) | Varies by tier | Burst handling capacity |
| Context Window Utilization | Prompt + completion token ratio | 10-100% | Cost and latency correlation |
Secondary Metrics
| Metric | Description | Why It Matters |
|---|---|---|
| Token Budget Compliance | Percentage of responses within max_tokens | Prevents runaway costs |
| Retry Rate | Percentage of requests requiring retry (429, 500, timeout) | Service reliability |
| Provider Fallback Rate | How often the system falls back to a secondary provider | Primary provider stability |
| Cache Hit Rate | Percentage of requests served from semantic cache | Cost optimization and latency |
| Streaming Drop Rate | Percentage of streams that disconnect before completion | Network reliability |