Cheap, fast inference for models worth building with.

GLM 5.3

z-ai/glm-5.3
Latency
4.14s
ThroughputP50 · rolling 5 min
40.0 tps
Uptime
100.00%
TTFT
2.34s
Recent uptimeUpdated programmatically.
100.00%

GLM 5.3 Flash

z-ai/glm-5.3-flash
Latency
8.81s
ThroughputP50 · rolling 5 min
17.5 tps
Uptime
100.00%
TTFT
5.37s
Recent uptimeUpdated programmatically.
100.00%

DeepSeek V4 Flash

deepseek/deepseek-v4-flash
Latency
7.77s
ThroughputP50 · rolling 5 min
28.7 tps
Uptime
100.00%
TTFT
2.36s
Recent uptimeUpdated programmatically.
100.00%

Built with experience and support from