For inference providers
List your models on Thali
Thali routes Indian developers, SMEs and enterprises to your models through one OpenAI-compatible API. You set the price; we handle distribution, billing in rupees, and settlement. Traffic follows performance — providers with better latency and uptime earn more of it.
You set the price
Per-token pricing is yours to publish and change, and you receive your published rate in full on every request. Thali's margin is added on top of it for the customer, never deducted from what you earn.
We handle the rupees
Indian customers pay us in INR. You receive monthly settlement backed by per-request metering you can audit.
Performance earns traffic
Routing weights latency, throughput and uptime. Serve well and your share grows — no listing fees, no pay-for-placement.
Residency is a feature
If you serve from Indian datacentres, your models carry
hosted_in: "in" and surface in our residency tier — demand no
global aggregator sends you.
What we ask of providers
| API | An OpenAI-compatible /chat/completions endpoint with SSE
streaming and token usage reported on both streamed and non-streamed
responses. |
| Catalog | A /models endpoint declaring pricing, context length,
feature support and — critically for us — where inference physically
runs. We verify residency claims before they reach our catalog. |
| Reliability | Sustained uptime. Our health checks and circuit breakers route around failures automatically; consistently unhealthy backends lose traffic and, eventually, their listing. |
| Data policy | A published statement of retention, logging and training use. Thali stores no prompts or completions, and we hold providers to the policy they publish. |
| Settlement | An Indian entity able to invoice with GST, or willingness to settle through one. Monthly, reconciled against our metering. |