Sign in
Docs
0%

Concepts

Usage and units

How a call is counted, what each category costs, the monthly allowance, and what happens when it runs out.

Units

Every operation in the Reference belongs to a category, and a category costs a fixed number of units per call, per 1000 tokens the answer used, or both. An account's usage is one number, units, the same across products, so an expensive call and a cheap one are counted on one scale, and a long model answer costs more than a short one.

CategoryPriceFor
read1 per callanswers from what is stored
compute4 per callanswers that run something small
write20 per callanything that changes state
run40 per calla program in the sandbox
generate4 per 1000 tokensa model generation, by its tokens

Tokens are the prompt and the completion together, as the answer's usage reports them, and are rounded up per count: a generation of 2 000 tokens costs 8 units.

GET /v1/accounts/units/categories answers the same list (scope usage:read, or signed in). A price change applies to calls made after it; what was already counted keeps the price it was counted at, and the changelog says when a price moves.

What is counted

Every answered call made for your account: with any of its API keys, and by you or your team on the inorbit.hr sites. A stream counts once, when it opens, and its tokens with its last message; a call over the WebSocket or MCP counts like the same call over REST.

Not counted:

  • a call the gateway refused (no token, a missing scope, an unknown route, rate_limited, quota_exceeded);
  • a call that failed on our side (internal, unavailable, timeout, unimplemented);
  • GET /v1/me and the OpenAPI documents.

The monthly allowance

An account's plan gives it an allowance of units per calendar month, UTC. Every account starts on free: 100 000 units a month. The allowance renews at 00:00 UTC on the 1st.

The allowance is the account's one budget. When it runs out on a plan that blocks, which free does, the account's calls are refused with 429 quota_exceeded until it renews or the allowance is raised: those made with its API keys, and its people's own use of the sites alike. A plan with overage keeps serving and records the excess. The products also bound one person's use per day (tokens, sandbox runs, pages), well above any plan, against abuse; anyone without an account has a smaller daily limit of their own.

{
  "code": "quota_exceeded",
  "error": "this account has used its units for the month; they renew on the 1st (UTC), or sooner when its plan or allowance is raised",
  "details": []
}

Do not retry quota_exceeded in a loop: nothing changes until the month turns or the allowance does. Counting happens a few seconds after each call, so an account can run a few calls past zero before its calls are refused.

Seeing it

The console's Usage page shows the month: units used, the allowance, what remains, where the month ends at the pace so far, a bar per day, the last 7, 30 or 90 days split by key, category or operation, and every operation by key with its category. With a key, GET /v1/accounts/orgs/{org_id}/units answers the month (scope usage:read), and GET /v1/accounts/orgs/{org_id}/units/series?by=who the days (by is who, category, operation or empty; from and to are UTC days, 92 at most). GET /v1/accounts/orgs/{org_id}/usage keeps answering what each product reported by day, units among them.

Both also say how the calls went: client_errors (calls that ended with your error, a bad argument, not found or refused; answered and billed like any call), platform_errors (calls we failed: unavailable, internal or a timeout; never billed), and p50_ms and p95_ms, the time at the gateway in which half and 95 in 100 of the calls were answered. The Usage page shows them as the error rate and answer time, and your keys and tokens each show their error rate.