Inference Benchmark · SGLang
Two load profiles measured against GLM-5.3 (NVFP4) with DFlash speculative decoding — a short low-latency chat profile and a long cached-context agentic profile. 64 prompts each, concurrency 8.
| Metric | Avg | p50 | p95 | p99 |
|---|---|---|---|---|
| TTFT | 977 | 761 | 2196 | 2573 |
| ITL | 14.5 | 3.7 | 63.2 | 213 |
| E2E | 2413 | 2041 | — | 5968 |
| TPOT | 14.5 | 11.5 | — | 38.4 |
| Metric | Avg | p50 | p95 | p99 |
|---|---|---|---|---|
| TTFT | 2773 | 2425 | 5842 | 6302 |
| ITL | 10.1 | 2.9 | 19.1 | 147 |
| E2E | 4889 | 4732 | — | 11661 |
| TPOT | 4.2 | 2.4 | — | 16.9 |
Percentile latency, lower is better. The agentic profile pays a larger prefill cost (40k-token prompts) despite 90% cache hit.
Output throughput (tokens/sec, aggregate across 8 concurrent streams) and median time-per-output-token. Higher throughput / lower TPOT is better.