Billing & usage
Every model is paid for through one of two lanes, and the gateway adds no markup on either. Platform-funded calls draw down your credits; bring-your-own-key calls are billed by the provider directly.
cost_micro_usd.estimated_cost_micro_usd for attribution only.Which lane a model rides is decided per-provider by its waterfall: a deployment backed by one of your provider connections is pass-through; a platform-seeded deployment is platform-funded. Either way, zero markup.
A credit is the platform’s spendable unit, pegged at a flat one cent. It is what everything is priced in, so display, checkout, and spend never drift between dollars and tokens. Today a credit is simply a cent under a friendlier name: the only thing that draws credits is routed token usage, which stays zero margin (platform-funded calls draw credits at the provider’s catalog price, with nothing added on top).
You get credits two ways, both at one cent each. The Free plan refreshes a monthly allotment. A Pro plan grants a larger monthly allotment and unlocks the Pro features. One-off top-ups buy credits without a plan, at the same flat rate. There is no markup on routed tokens; any margin comes from plans, not from a credit spread.
A new organization starts with a welcome credit grant. Your balance is the credit granted minus your billable (platform-funded) spend; pass-through usage does not count against it. Balance, spend, adding credits, and auto-recharge live in the dashboard at Credits.
GET /api/gateway/usage/daily for spend by day, model, or member. See Telemetry.You can connect more than one account for the same provider — two Anthropic keys, two OpenAI organizations — each under its own handle. They form a pool: the gateway serves your traffic on the first account in your order, and rotates to the next one when an account runs out of quota, or is rate-limited in a sustained way (a burst of throttles over fifteen minutes; a single throttle never rotates, because switching accounts busts the prompt cache you have built on the current one). Rotation is a verdict written from your own traffic every five minutes; a later successful key check re-admits the account.
Manage the pool on the Credits page: drag accounts to set the order, switch each account’s “rotate when out of quota” and “rotate on sustained rate limit” off to fail on it instead of spending on a sibling, and read every account’s usage on its own key. The same controls are one call for an agent holding your org key: GET /api/orgs/{org_id}/provider-connections/accounts/usage, POST /api/orgs/{org_id}/provider-connections/{provider}/accounts/{setup_alias}/routing, and POST /api/orgs/{org_id}/provider-connections/reorder.
Spend is bounded at three levels, all configured in the dashboard:
A key can read its own effective limits over the API. GET /api/gateway/keys/<api_key_id>/limits returns the three ceilings with platform defaults folded in; a null value means uncapped, and source is explicit when set on the key or default otherwise. Setting limits is an admin dashboard action.
curl "https://api-pr-1391.preview.experientiallabs.ai/api/gateway/keys/$API_KEY_ID/limits" \-H "Authorization: Bearer $EXPLABS_API_KEY"
| field | Meaning |
|---|---|
| daily_spend_cap_micro_usd | Max platform-funded spend per day for this key (micro-USD). |
| requests_per_minute | Request-rate ceiling for this key. |
| tokens_per_minute | Token-rate (TPM) ceiling for this key. |
When your credit balance or a spend limit is exhausted, calls fail with 429 insufficient_quota; the message says which: a daily org or per-model cap, a budget, or your credits. It is not transient: retrying does not clear it.
The full error contract is in Errors.