Rate limits
The short version
Your quota is tokens per day, shared across every API key on your account.
curl https://api.thali.ai/api/v1/usage -H "Authorization: Bearer $THALI_API_KEY"
The limits
| Free | Unlocked (₹99 one-time) | |
|---|---|---|
| Tokens / day (primary) | see /api/v1/usage |
see /api/v1/usage |
| Requests / day | 50 | 1,000 |
| Requests / minute | 20 | 20 |
| Concurrent requests | 2 | 4 |
max_tokens per request |
4,096 | 8,192 |
All limits apply across every :free model combined, and reset at 00:00 UTC.
Why tokens and not just requests
Requests are a bad proxy for cost. One request with a 131k-token context and an 8k-token completion occupies a GPU for minutes; one request saying "hi" occupies it for milliseconds. A limit that counts them the same either throttles the second user pointlessly or lets the first user take the whole cluster.
So the quota that actually governs you is tokens. Requests/minute and requests/day still exist, but only as guards against pathological traffic.
Keys don't multiply your quota
You may hold up to 5 API keys. They are for separating environments, not for multiplying throughput — every limit above is enforced per account, and all your keys draw on the same pool.
What a 429 tells you
{
"error": {
"message": "Daily token quota exhausted: 10,000 of 10,000 tokens used on tier 'free'. Resets in 41203s at 00:00 UTC.",
"type": "rate_limit_error",
"param": null,
"code": "tokens_per_day"
}
}
code is always one of tokens_per_day, requests_per_day,
requests_per_minute, or concurrency, and there is always a Retry-After
header. Back off on it rather than retrying immediately.
Exceeding max_tokens
We return 400, not a quietly truncated response:
{
"error": {
"message": "max_tokens=32000 exceeds the 4096 limit for tier 'free'. Lower max_tokens or upgrade your tier.",
"type": "invalid_request_error",
"param": "max_tokens",
"code": "max_tokens_exceeded"
}
}
Silently capping a request you asked for is worse than refusing it — you'd spend an afternoon wondering why your outputs are short.
Requests that omit max_tokens don't escape the ceiling: the gateway
injects your tier's limit before forwarding, so the cap holds whether or not
you declare a value. Set max_tokens explicitly if you want less than the
ceiling.
Streams still count
If you disconnect halfway through a stream, the tokens generated up to that point are still billed against your quota. They were produced on a GPU whether or not anything read them.
Upgrading
A one-time ₹99 top-up moves your account to the unlocked tier permanently.
During the MVP this is applied manually; Razorpay self-serve arrives in Phase 2.
When the service says 503
Two cases, distinguished by code:
no_healthy_backend— every backend for that model is down. Try another model;/api/v1/modelsmarks which areavailable.free_budget_exhausted— the free tier's global monthly budget is spent. Free models resume next month. This is how we keep a free tier that exists at all rather than one that quietly bankrupts the service.