Models
We host these ourselves. Speeds are measured on our own hardware at a single stream, not quoted from a datasheet.
| Model | tokens/s | Context | Input / 1M | Output / 1M |
|---|---|---|---|---|
glm-4.7-flashFast, tool calls, logprobs, 202k context. The default: the lower-priced of the two and the one most workloads should use. | 62.8 | 202,752 | 0.25 | 1.50 |
qwen3.8-27b27.3B dense at F16, 262k context, tool calls, logprobs. Slower per token, priced 3x the flash model because it holds the GPUs 3.6x longer. | 17.5 | 262,144 | 0.75 | 4.50 |
Prices exclude 5% UAE VAT, which applies to UAE customers.