زاني إل إل إم
القائمة

Chat completions

POST /v1/chat/completions takes the OpenAI shape, including stream, tools and max_tokens.

Request fields behave as your client expects. These response headers are ours:

HeaderMeaning
X-Request-IdThe id for this call. Send your own to make a retry idempotent: the same id is metered and charged once.
X-Zanii-Cost-MicroWhat the call cost, in micro-AED, a millionth of a dirham.
X-Zanii-Balance-MicroWhat is left after it.
X-Zanii-Receipt-HashThe receipt, when the key has receipts enabled.
X-Zanii-Verify-UrlWhere anyone can check that receipt.

Streaming

Set stream: true. Frames arrive as the model produces them, and the final frame carries the token

usage, which is what the call is billed from. You do not need to set stream_options: usage is

always requested, because a stream that cannot be counted cannot be billed honestly.

Reasoning, and what it costs you

These models reason before they answer, and reasoning tokens are billed like any other because the

GPU produced them. They also come out of the same max_tokens budget as the reply, so a budget too

small for both returns a successful call with empty content: the tokens were produced and billed,

but none reached you.

Two things follow. Give a reasoning model room, a few hundred tokens beyond the answer you expect.

And turn the reasoning off when the question does not need it:

{"model": "glm-4.7-flash", "messages": [...],
 "chat_template_kwargs": {"enable_thinking": false}}

Measured on glm-4.7-flash with a one-sentence question: 412 output tokens in 4.4 s with

reasoning on, 22 tokens in 0.6 s with it off, for the same answer. That is roughly nineteen

times the cost and seven times the wait, so it is worth a thought per endpoint rather than per

project. Keep reasoning for work where it earns its tokens.

In the Python SDK this is thinking=False. When answer.text comes back empty,

answer.empty_because says why and answer.reasoning_tokens says how much of the budget went on

thinking.

Models

GET /v1/models lists what your key may call, with the price per million tokens. A model that is

not in the list cannot be called: an unpriced call would be an unbillable one.

GET /v1/models/{id} is one model on its own: its price, its context length and whether it is

live.

Embeddings

POST /v1/embeddings takes the OpenAI shape and returns it. Priced on prompt tokens only, at the

same rate a completion pays for its prompt. If the model server behind your deployment has no

embedding model, this answers 501 embeddings_not_supported rather than pretending.

Counting tokens before you spend any

POST /v1/tokenize takes {"model": "...", "prompt": "..."} (or messages) and returns the

token count and what the prompt would cost, without calling the model or charging anything.

The count comes from the model server's own tokenizer. If it cannot tokenize, you get

503 tokenizer_unavailable instead of an estimate: a number you plan a budget against should be

the real one.

Usage, balance, receipts and invoices

GET /v1/usage returns what this key's account has spent, and GET /v1/balance what is left.

GET /v1/usage/daily is the same by day, including today.

GET /v1/receipts lists this account's receipts and what the ledger did with each, and

GET /v1/invoices lists its invoices. Both are paged with ?limit= and ?offset=, and both are

authenticated by the API key itself, so you can reconcile our invoice against your own records,

or check that every call you were charged for has a proof, from a script rather than a browser.

Retrying safely

Send Idempotency-Key (or X-Request-Id) and a retry of the same request will not be charged

twice.

A request id that was already served is refused with 409 idempotent_replay, and the error says

when the original ran. It is refused rather than answered again because we do not store

completions: there is nothing to replay, and calling the model a second time for free would be a

gap in the billing rather than a kindness. A first attempt that failed before it was billed leaves

no record, so a genuine retry after a failure goes through normally.

Spend caps

An account can carry a monthly ceiling, set in the console. A call that would take the month past

it is refused with 402 spend_cap_reached before the model is asked, so the limit costs nothing to

enforce and a runaway agent cannot spend the balance.

markdown · التوثيق