Thali

Quickstart

Thali is an OpenAI-compatible API. If your code already talks to OpenAI, you change two lines: the base URL and the key.

Target: zero to your first completion in under five minutes.


1. Get a key

# Send yourself a verification code
curl -X POST https://api.thali.ai/api/v1/auth/otp/start \
  -H 'content-type: application/json' \
  -d '{"phone": "+919876543210"}'

# Exchange the code for an account and an API key
curl -X POST https://api.thali.ai/api/v1/auth/otp/verify \
  -H 'content-type: application/json' \
  -d '{"phone": "+919876543210", "code": "123456"}'

The response contains your key:

{
  "account_id": "01J9…",
  "tier": "free",
  "new_account": true,
  "api_key": "thali-sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
  "key_id": "01J9…"
}

Copy it now. We store only a SHA-256 hash, so this is the only time the full key exists anywhere outside your terminal. Lost it? Create another (POST /api/v1/keys, up to 5 per account).


2. Your first completion

curl

curl https://api.thali.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -H 'content-type: application/json' \
  -d '{
    "model": "openai/gpt-oss-20b:free",
    "messages": [{"role": "user", "content": "Explain UPI in one sentence."}]
  }'

Python (official openai SDK)

from openai import OpenAI

client = OpenAI(
    base_url="https://api.thali.ai/api/v1",   # <- the only two lines that change
    api_key="thali-sk-...",
)

completion = client.chat.completions.create(
    model="openai/gpt-oss-20b:free",
    messages=[{"role": "user", "content": "Explain UPI in one sentence."}],
)
print(completion.choices[0].message.content)

TypeScript (official openai SDK)

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.thali.ai/api/v1",
  apiKey: process.env.THALI_API_KEY,
});

const completion = await client.chat.completions.create({
  model: "openai/gpt-oss-20b:free",
  messages: [{ role: "user", content: "Explain UPI in one sentence." }],
});
console.log(completion.choices[0].message.content);

3. Streaming

Standard SSE, terminated by data: [DONE].

stream = client.chat.completions.create(
    model="openai/gpt-oss-20b:free",
    messages=[{"role": "user", "content": "Write a haiku about Chennai rain."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Want token counts on a stream? Ask for them:

stream = client.chat.completions.create(
    ...,
    stream=True,
    stream_options={"include_usage": True},
)

The final chunk then carries usage and an empty choices list. If you don't ask, you won't get it — the stream stays byte-identical to what OpenAI sends.


4. Embeddings

result = client.embeddings.create(
    model="nvidia/llama-nemotron-embed-vl-1b-v2:free",
    input=["chennai", "madurai"],
)
print(len(result.data[0].embedding))

5. What's available

curl https://api.thali.ai/api/v1/models

No key required. Each entry carries the standard OpenAI fields plus:

Field Meaning
pricing.prompt_inr_per_mtok ₹ per million input tokens. 0 for :free.
pricing.completion_inr_per_mtok ₹ per million output tokens.
license The model's actual licence. We only serve what we've cleared.
data_policy zero_retention — we never store your prompts or completions.
hosted_in in — served from Indian infrastructure.
available Whether a healthy backend is serving it right now.
context_length Maximum context window.

Note on pricing: OpenRouter returns strings of USD-per-token. We return numbers in ₹ per million tokens, because we bill in rupees. It's the one field where a port from OpenRouter needs a change.


6. Limits

See rate-limits.md. The short version: your quota is tokens per day, shared across every key on your account, and you can check it any time:

curl https://api.thali.ai/api/v1/usage -H "Authorization: Bearer $THALI_API_KEY"

7. Errors

Identical in shape to OpenAI's, so your SDK raises the exceptions it already knows:

{
  "error": {
    "message": "Daily token quota exhausted: 10,000 of 10,000 tokens used on tier 'free'. Resets in 41,203s at 00:00 UTC.",
    "type": "rate_limit_error",
    "param": null,
    "code": "tokens_per_day"
  }
}
Status When
400 Bad request — including max_tokens above your tier's ceiling. We reject rather than silently truncate.
401 Missing, malformed, unknown, or revoked key.
404 Unknown model. The message names /api/v1/models.
429 A limit was hit. code names which one; Retry-After says when.
502 The upstream model rejected the request.
503 No healthy backend, or free-tier budget exhausted for the month.

Running it locally

cp .env.example .env
# set FREE_MONTHLY_TOKEN_BUDGET and *_TOKENS_PER_DAY -- they default to 0,
# which disables free inference on purpose
docker compose -f infra/docker-compose.yml --profile dev up --build

That brings up the gateway, Postgres, Redis, the metering worker, and a mock backend that speaks the OpenAI protocol without needing a GPU. Then point any OpenAI SDK at http://localhost:8000/api/v1.