GLM-5.3 Benchmarks

Inference Benchmark · SGLang

GLM-5.3 Benchmarks

Two load profiles measured against GLM-5.3 (NVFP4) with DFlash speculative decoding — a short low-latency chat profile and a long cached-context agentic profile. 64 prompts each, concurrency 8.

4×B300 · TP4/EP4 NVFP4 · fp8 KV DFlash spec-decode 64 prompts · conc 8
Scenario A · 50.89 s · 64/64 ok

Real-time generation

3,000 in100 outno cachecustom format
3.18 req/s
Requests
318 tok/s
Output
4.14
Accept len
Latency (ms)
MetricAvgp50p95p99
TTFT97776121962573
ITL14.53.763.2213
E2E24132041—5968
TPOT14.511.5—38.4
Scenario B · 1 m 19 s · 64/64 ok

Online agentic

40,000 in36,000 cached500 outOpenAI format
1.55 req/s
Requests
775 tok/s
Output
4.50
Accept len
Latency (ms)
MetricAvgp50p95p99
TTFT2773242558426302
ITL10.12.919.1147
E2E48894732—11661
TPOT4.22.4—16.9

Time to First Token

Percentile latency, lower is better. The agentic profile pays a larger prefill cost (40k-token prompts) despite 90% cache hit.

Real-time (3k in) Agentic (40k in)

Decode speed & throughput

Output throughput (tokens/sec, aggregate across 8 concurrent streams) and median time-per-output-token. Higher throughput / lower TPOT is better.

Real-time Agentic