Credits & billing
Free models need no credit. Paid models are charged against a balance, in rupees.
The 1,000,000 free tokens you start with
Every new account starts with 1,000,000 free tokens — real, spendable credit (₹100 under the hood) with one restriction: it works on our open-model tier, priced up to ₹100/Mtok output (the DeepSeek V3 / Llama / Qwen / Sarvam / Gemma class, most of the catalog). At that price a million output tokens is the floor; input-heavy prompts stretch it further. Frontier and reasoning models (DeepSeek R1, GLM 4.5, Kimi K3, …) need topped-up credit.
Promotional credit is spent first: as long as any remains, eligible
requests draw it down before touching money you've added. The
promo_balance_inr field below shows how much is left.
Checking your balance
curl https://api.thali.ai/api/v1/credits -H "Authorization: Bearer $THALI_API_KEY"
{
"data": {
"total_credits_inr": 500.0,
"total_usage_inr": 0.0007,
"balance_inr": 499.9993,
"promo_balance_inr": 100.0,
"balance_micropaise": 49999930000
}
}
Both units are returned deliberately. balance_inr is for display;
balance_micropaise is the exact integer, for anyone reconciling.
promo_balance_inr is the portion of your balance that came from the signup
grant and is restricted to eligible models.
Why micropaise
Pricing is quoted per million tokens, so a small request costs a fraction of a paise. A real example from this gateway: 2 prompt tokens and 11 completion tokens against a ₹20/₹60-per-Mtok model costs 70,000 micropaise — ₹0.0007, which is 7 hundredths of a paise.
Round that to paise and you either charge a full paise for it (a 1,400% markup) or charge nothing (and serve a million of them free). So balances are held as integers in micropaise, 10⁻⁶ of a paise:
cost_micropaise = tokens × price_paise_per_mtok # exact, no division
No floating point touches money at any point. Rupees appear only at the edge, for humans.
Adding credit
Top-up
curl -X POST https://api.thali.ai/api/v1/credits/top-up \
-H "Authorization: Bearer $THALI_API_KEY" \
-H 'content-type: application/json' \
-d '{"amount_inr": 500}'
{
"trade_no": "TH01KYMP9MKS51MMQ1RNE4SFQHKC",
"amount_inr": 500.0,
"credit_inr": 500.0,
"status": "created",
"payment_url": null
}
The trade_no is ours, not the payment provider's, and the record exists before
you are sent anywhere to pay. Keep it: it is the reference for any query about a
payment that did not land. Minimum top-up is ₹10.
Razorpay is not wired yet, so payment_url is null during the MVP and top-ups
are settled by hand.
Redemption codes
Prepaid credit that needs no payment method at redemption time.
curl -X POST https://api.thali.ai/api/v1/credits/redeem \
-H "Authorization: Bearer $THALI_API_KEY" \
-H 'content-type: application/json' \
-d '{"code": "thali-gift-XXXX-XXXX-XXXX-XXXX"}'
Each code works exactly once.
What a request cost
Every response carries three headers: X-Thali-Served-By (the backend that ran),
X-Thali-Model (the model that actually served, which differs from the one you
asked for when a fallback fired), and X-Thali-Generation-Id — the id of this
request's usage record. That id is how you look the cost up afterwards:
ID=$(curl -sD- https://api.thali.ai/api/v1/chat/completions ... \
| grep -i x-thali-generation-id | cut -d' ' -f2 | tr -d '\r')
curl https://api.thali.ai/api/v1/generation/$ID -H "Authorization: Bearer $THALI_API_KEY"
On a stream the header arrives with the response headers, before the first token.
{
"data": {
"model": "thali/dev-paid",
"served_by": "01KYKK0X8RD0KFZ8CBB0GS6K3S",
"tokens_prompt": 2,
"tokens_cached": 0,
"tokens_completion": 11,
"cost_inr": 0.0007,
"cost_micropaise": 70000,
"usage_estimated": false
}
}
tokens_cached is the part of tokens_prompt the provider served from its own
cache — a subset, not an addition. Where a model has a cached-input rate, those
tokens are billed at it instead of the normal input rate, which is why two prompts
of the same length can cost different amounts.
usage_estimated: true means the client disconnected mid-stream, so the token
counts are a floor rather than a report from the model. Flagged rather than
hidden, so nobody reconciles against a number we already know is approximate.
Full movement history is at GET /api/v1/credits/history.
Per-key spend limits
An account has one balance. A key may be capped below it — useful for handing a key to a contractor or embedding one in a client:
{"key_limit_inr": 500.0, "key_spent_inr": 12.4}
Two rules, and they are not the same mechanism:
- Rate limits are always per account. Five keys never mean five times the throughput — see rate-limits.md.
- A key's spend cap only ever binds tighter than the account pool. Setting a ₹500 key limit on a ₹100 account gives you ₹100, not ₹500.
Inspect the calling key at GET /api/v1/key.
When you run out
Paid models return 402:
{
"error": {
"message": "Insufficient credit for the paid model 'x'. Balance is ₹0.00. …",
"type": "insufficient_credit",
"code": "insufficient_credit"
}
}
402 and not 429 on purpose: a rate limit means wait, and a credit failure
means pay. A client that retries the second one on a backoff would loop until
it gave up.
Free models keep working at a zero balance. That is the point of them.
Timing
Charges are applied by a background worker, not on the request path — a
completion never waits on a ledger write. Your balance can therefore lag by a
second or two, and a burst can push it slightly negative. The overshoot is
bounded by max_tokens_per_request × max_concurrent, which is one of the reasons
both limits exist.
Capping what one key can spend
A key may carry a spend limit that binds tighter than your balance. Useful for handing a key to a contractor, a CI job, or an agent you would rather not trust with the whole wallet.
# at creation
curl -X POST https://api.thali.ai/api/v1/keys \
-H "Authorization: Bearer $THALI_API_KEY" \
-d '{"name": "ci", "limit_inr": 500}'
# raise, lower, or remove it later
curl -X PATCH https://api.thali.ai/api/v1/keys/$KEY_ID \
-H "Authorization: Bearer $THALI_API_KEY" \
-d '{"limit_inr": 750}'
"limit_inr": null removes the cap. Omitting the field leaves it unchanged —
the two are different requests.
A key limit can only ever bind tighter than the account pool, never looser: a ₹500 key on a ₹100 balance can still spend ₹100. When a key exhausts its own cap while the account has funds left, the 402 says so explicitly, so you can tell "this key is done" from "you are out of credit".