Thali

Routing

By default Thali picks the healthiest backend for the model you asked for. When that isn't what you want, say so.

Falling back to another model

{
  "model": "openai/gpt-oss-20b:free",
  "models": ["nvidia/nemotron-nano-9b-v2:free"],
  "messages": [...]
}

Tried in order. The model that actually served is reported in the response body's model field and in the X-Thali-Model header — never the one you asked for. If a fallback fired, you will know.

A model in the list that doesn't exist is a 404, not a silent skip. A typo in a fallback list would otherwise degrade quietly and permanently.

Fallback also respects your balance per model, at the moment each is reached. So {"model": "paid-one", "models": ["free-one:free"]} on an empty balance serves the free model rather than refusing the request.

Choosing a backend

{
  "model": "openai/gpt-oss-20b:free",
  "provider": {
    "order": ["e2e-blr-1"],
    "ignore": ["old-node"],
    "sort": "latency",
    "allow_fallbacks": true
  },
  "messages": [...]
}
Field Effect
order Named backends first, in this order. Others follow unless allow_fallbacks is false.
ignore Never use these.
sort price, throughput, latency, or priority.
allow_fallbacks false restricts the request to order only.

Backends are named by label, which you'll find in GET /api/v1/status.

The default order is cost-aware

When you express no preference, healthy backends for a model are tried cheapest first — cheapest by what serving your request actually costs, which is how the same model gets served from whichever provider is most efficient at that moment without you doing anything. A backend whose cost is unknown sorts last; a provider that throttles or fails simply yields to the next one in line.

Writing "sort": "priority" explicitly is the opt-out: it restores the operator's static ordering for that request.

Preferences narrow. They never widen.

This is the important rule. provider.order cannot reach a backend that is unhealthy, licence-blocked, or sitting behind a tripped circuit breaker. Preferences filter the set we were already willing to serve from — they are not a way to request something we decided not to serve.

If your preferences exclude everything, you get a clean 503 no_healthy_backend rather than a silent fallback to something you asked us not to use.

Variants

Suffix Meaning
:free Part of the model's name, not a routing mode.
:nitro Sort backends by measured throughput.
:floor Sort candidate models by price, cheapest first.
{"model": "openai/gpt-oss-20b:free:nitro"}
{"model": "expensive-model:floor", "models": ["cheaper-model"]}

An explicit provider.sort beats a variant suffix — you wrote it out in full, so you meant it.

How latency and throughput are measured

Each gateway keeps a rolling average per backend, fed from completed requests. It is per instance and not shared, deliberately: latency is a property of the path between that gateway and that backend, and two instances in different racks legitimately disagree about which node is fastest. Averaging them across the fleet would produce a number true for neither.

A freshly started instance routes on priority until it has observed a few requests, and a backend nobody has measured yet always sorts behind one that has. Current measurements are visible at GET /api/v1/status.

What isn't here

No automatic model selection — no equivalent of an "auto" router that picks a model for you. That is the ROI router, and it is deliberately not in this release. The seam it will plug into is select_backend(); everything on this page is built on the same one.


Your own MCP servers

Thali brokers tools the way it routes models. Register a server and its tools appear to any MCP client you connect, billed and rate-limited exactly like the built-in ones.

curl -X POST https://thaliai.in/api/v1/mcp/servers \
  -H "Authorization: Bearer $THALI_API_KEY" \
  -d '{
        "label": "github",
        "url": "https://mcp.example.com/mcp",
        "auth_scheme": "bearer",
        "auth_value": "ghp_..."
      }'

The label namespaces that server's tools: its search becomes github__search, so two servers can both expose a search and neither collides with a built-in. Labels are lowercased, must be 3–32 characters of letters, digits and hyphens, and are unique within your account — two customers may both use github.

GET /api/v1/mcp/servers list yours
GET /api/v1/mcp/servers/{id} one, with its tools
PATCH /api/v1/mcp/servers/{id} change url, enabled, or auth
POST /api/v1/mcp/servers/{id}/refresh re-probe and re-read tools/list
DELETE /api/v1/mcp/servers/{id} remove it

Things worth knowing

Tool lists are cached. Agents call tools/list constantly, so we serve each server's tools from the last successful probe rather than fanning out to every one of your servers on every call — otherwise your slowest server becomes everyone's latency. Added a tool? Call refresh.

auth_value is write-only. It is sent once and never returned by any endpoint; responses carry has_auth_value: true instead. Omitting it on a PATCH keeps the stored value, so you can change a URL without re-sending a secret you cannot read back.

Your server must be on the public internet. URLs resolving to loopback, private, link-local or reserved addresses are refused — at registration and again immediately before every call, so a hostname that changes where it points after registration does not get a free pass. Redirects are not followed.

A broken server is your server's outage, not ours. A failed call marks that server unhealthy with the reason on last_error, returns an MCP tool error, and leaves the rest of your session working.

Brokered calls count. They pass through the same quota and metering as the built-in tools — registering your own server is not a way around the rate limit.