Thali

Credits & billing

Free models need no credit. Paid models are charged against a balance, in rupees.

The 1,000,000 free tokens you start with

Every new account starts with 1,000,000 free tokens — real, spendable credit (₹100 under the hood) with one restriction: it works on our open-model tier, priced up to ₹100/Mtok output (the DeepSeek V3 / Llama / Qwen / Sarvam / Gemma class, most of the catalog). At that price a million output tokens is the floor; input-heavy prompts stretch it further. Frontier and reasoning models (DeepSeek R1, GLM 4.5, Kimi K3, …) need topped-up credit.

Promotional credit is spent first: as long as any remains, eligible requests draw it down before touching money you've added. The promo_balance_inr field below shows how much is left.

Checking your balance

curl https://api.thali.ai/api/v1/credits -H "Authorization: Bearer $THALI_API_KEY"
{
  "data": {
    "total_credits_inr": 500.0,
    "total_usage_inr": 0.0007,
    "balance_inr": 499.9993,
    "promo_balance_inr": 100.0,
    "balance_micropaise": 49999930000
  }
}

Both units are returned deliberately. balance_inr is for display; balance_micropaise is the exact integer, for anyone reconciling. promo_balance_inr is the portion of your balance that came from the signup grant and is restricted to eligible models.

Why micropaise

Pricing is quoted per million tokens, so a small request costs a fraction of a paise. A real example from this gateway: 2 prompt tokens and 11 completion tokens against a ₹20/₹60-per-Mtok model costs 70,000 micropaise — ₹0.0007, which is 7 hundredths of a paise.

Round that to paise and you either charge a full paise for it (a 1,400% markup) or charge nothing (and serve a million of them free). So balances are held as integers in micropaise, 10⁻⁶ of a paise:

cost_micropaise = tokens × price_paise_per_mtok      # exact, no division

No floating point touches money at any point. Rupees appear only at the edge, for humans.

Adding credit

Top-up

curl -X POST https://api.thali.ai/api/v1/credits/top-up \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -H 'content-type: application/json' \
  -d '{"amount_inr": 500}'
{
  "trade_no": "TH01KYMP9MKS51MMQ1RNE4SFQHKC",
  "amount_inr": 500.0,
  "credit_inr": 500.0,
  "status": "created",
  "payment_url": null
}

The trade_no is ours, not the payment provider's, and the record exists before you are sent anywhere to pay. Keep it: it is the reference for any query about a payment that did not land. Minimum top-up is ₹10.

Razorpay is not wired yet, so payment_url is null during the MVP and top-ups are settled by hand.

Redemption codes

Prepaid credit that needs no payment method at redemption time.

curl -X POST https://api.thali.ai/api/v1/credits/redeem \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -H 'content-type: application/json' \
  -d '{"code": "thali-gift-XXXX-XXXX-XXXX-XXXX"}'

Each code works exactly once.

What a request cost

Every response carries three headers: X-Thali-Served-By (the backend that ran), X-Thali-Model (the model that actually served, which differs from the one you asked for when a fallback fired), and X-Thali-Generation-Id — the id of this request's usage record. That id is how you look the cost up afterwards:

ID=$(curl -sD- https://api.thali.ai/api/v1/chat/completions ... \
       | grep -i x-thali-generation-id | cut -d' ' -f2 | tr -d '\r')

curl https://api.thali.ai/api/v1/generation/$ID -H "Authorization: Bearer $THALI_API_KEY"

On a stream the header arrives with the response headers, before the first token.

{
  "data": {
    "model": "thali/dev-paid",
    "served_by": "01KYKK0X8RD0KFZ8CBB0GS6K3S",
    "tokens_prompt": 2,
    "tokens_cached": 0,
    "tokens_completion": 11,
    "cost_inr": 0.0007,
    "cost_micropaise": 70000,
    "usage_estimated": false
  }
}

tokens_cached is the part of tokens_prompt the provider served from its own cache — a subset, not an addition. Where a model has a cached-input rate, those tokens are billed at it instead of the normal input rate, which is why two prompts of the same length can cost different amounts.

usage_estimated: true means the client disconnected mid-stream, so the token counts are a floor rather than a report from the model. Flagged rather than hidden, so nobody reconciles against a number we already know is approximate.

Full movement history is at GET /api/v1/credits/history.

Per-key spend limits

An account has one balance. A key may be capped below it — useful for handing a key to a contractor or embedding one in a client:

{"key_limit_inr": 500.0, "key_spent_inr": 12.4}

Two rules, and they are not the same mechanism:

  • Rate limits are always per account. Five keys never mean five times the throughput — see rate-limits.md.
  • A key's spend cap only ever binds tighter than the account pool. Setting a ₹500 key limit on a ₹100 account gives you ₹100, not ₹500.

Inspect the calling key at GET /api/v1/key.

When you run out

Paid models return 402:

{
  "error": {
    "message": "Insufficient credit for the paid model 'x'. Balance is ₹0.00. …",
    "type": "insufficient_credit",
    "code": "insufficient_credit"
  }
}

402 and not 429 on purpose: a rate limit means wait, and a credit failure means pay. A client that retries the second one on a backoff would loop until it gave up.

Free models keep working at a zero balance. That is the point of them.

Timing

Charges are applied by a background worker, not on the request path — a completion never waits on a ledger write. Your balance can therefore lag by a second or two, and a burst can push it slightly negative. The overshoot is bounded by max_tokens_per_request × max_concurrent, which is one of the reasons both limits exist.

Capping what one key can spend

A key may carry a spend limit that binds tighter than your balance. Useful for handing a key to a contractor, a CI job, or an agent you would rather not trust with the whole wallet.

# at creation
curl -X POST https://api.thali.ai/api/v1/keys \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -d '{"name": "ci", "limit_inr": 500}'

# raise, lower, or remove it later
curl -X PATCH https://api.thali.ai/api/v1/keys/$KEY_ID \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -d '{"limit_inr": 750}'

"limit_inr": null removes the cap. Omitting the field leaves it unchanged — the two are different requests.

A key limit can only ever bind tighter than the account pool, never looser: a ₹500 key on a ₹100 balance can still spend ₹100. When a key exhausts its own cap while the account has funds left, the 402 says so explicitly, so you can tell "this key is done" from "you are out of credit".