Zanii LLM
Menu

Models

We host these ourselves. Speeds are measured on our own hardware at a single stream, not quoted from a datasheet.

Modeltokens/sContextInput / 1MOutput / 1M
glm-4.7-flash
Fast, tool calls, logprobs, 202k context. The default: the lower-priced of the two and the one most workloads should use.
62.8202,7520.251.50
qwen3.8-27b
27.3B dense at F16, 262k context, tool calls, logprobs. Slower per token, priced 3x the flash model because it holds the GPUs 3.6x longer.
17.5262,1440.754.50

Prices exclude 5% UAE VAT, which applies to UAE customers.