/docs
Baton docs
How to point a client at Baton, how it grades and routes a request, and every header, limit and error of its API.
Introduction
Baton is an LLM router with an OpenAI-compatible API and an Anthropic-compatible API: you point a client at it and send one key. It grades how hard each request is and sends it to one of the models you have put in order on your ladder, using your own provider keys. Hard requests stay on the model at the top, medium requests leave a model once half of its daily budget is spent, and easy requests go to the model at the bottom. When no model on your ladder can answer, the house model answers and is paid from your prepaid credit.
Quickstart
1. Get a key
Open the sign-in page and sign in with a Solana wallet, or choose to continue without one. The first sign-in creates a workspace and its first API key. The key is shown once, so copy it before you go on. If you lose it, create another on the Keys page.
2. Send a request
Point your client at one of the two base URLs and put your key where the samples say YOUR_BATON_KEY. The model name baton means "let the router decide".
- OpenAI-style base URL
- https://www.batonapi.com/v1
- Anthropic-style base URL
- https://www.batonapi.com
- Model name
- baton
curl https://www.batonapi.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_BATON_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "baton",
"messages": [{"role": "user", "content": "Explain what a mutex is in one sentence."}]
}'# pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://www.batonapi.com/v1",
api_key="YOUR_BATON_KEY",
)
completion = client.chat.completions.create(
model="baton",
messages=[{"role": "user", "content": "Explain what a mutex is in one sentence."}],
)
print(completion.choices[0].message.content)// npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://www.batonapi.com/v1",
apiKey: "YOUR_BATON_KEY",
});
const completion = await client.chat.completions.create({
model: "baton",
messages: [{ role: "user", content: "Explain what a mutex is in one sentence." }],
});
console.log(completion.choices[0].message.content);curl https://www.batonapi.com/v1/messages \
-H "x-api-key: YOUR_BATON_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "baton",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Explain what a mutex is in one sentence."}]
}'# pip install anthropic
import anthropic
client = anthropic.Anthropic(
base_url="https://www.batonapi.com",
api_key="YOUR_BATON_KEY",
)
message = client.messages.create(
model="baton",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain what a mutex is in one sentence."}],
)
for block in message.content:
if block.type == "text":
print(block.text)// npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://www.batonapi.com",
apiKey: "YOUR_BATON_KEY",
});
const message = await client.messages.create({
model: "baton",
max_tokens: 1024,
messages: [{ role: "user", content: "Explain what a mutex is in one sentence." }],
});
for (const block of message.content) {
if (block.type === "text") console.log(block.text);
}A new workspace has no models on its ladder, so the house model answers and the request is paid from credit. When the workspace has no credit the request fails with status 402, and on a deployment without a house model it fails with 503. To route between models of your own, add a provider key and a model.
To try a request without a client, use the Playground in the app. It sends a real request through the router and shows which model answered.
3. Read the response headers
Every routed response says where it went. With curl, add -i to see the headers:
x-baton-model: claude-opus-4-8
x-baton-provider: anthropic
x-baton-rung: 0
x-baton-tier: mediumx-baton-modelis the model that answered, by the id its provider knows it by.x-baton-provideris the provider that served it, orhousefor the house model.x-baton-rungis the model's place on your ladder.0is the top and-1is the house model.x-baton-tieris what the request was routed as:hard,mediumoreasy. The sample request is graded medium; the grader section shows why.
The Requests page in the app lists the same facts for every request Baton has routed for you.
Keys and sign-in
Sending a key
Every request to the API and to the MCP server carries an API key, in either of two headers. OpenAI clients send the first, Anthropic clients the second:
Authorization: Bearer YOUR_BATON_KEY
x-api-key: YOUR_BATON_KEYWhen a request carries both, the bearer token is the one that is checked. An Authorization header that is not a bearer token is ignored in favour of x-api-key. A missing, unknown or revoked key gets status 401.
What is stored
A key is bt_ followed by 40 letters and digits. Only its SHA-256 hash is stored, next to a masked form for the Keys page: the prefix with the four characters after it, and the last four. The key itself is shown once, when it is made, and cannot be shown again.
Workspaces and sign-in
Keys belong to a workspace. There are two kinds of workspace, and they differ only in how you get back in.
- Wallet-linked. You sign in with a Solana wallet by signing a message. Nothing is sent on chain and it costs nothing. The first sign-in with a wallet creates its workspace and a first key named "Default"; later sign-ins with the same wallet open the same workspace.
- Key-only. "Continue without a wallet" creates a workspace and a first key. That key, or any other active key of the workspace, is then the only way to sign back in: paste it on the sign-in page. Revoking the last active key of a key-only workspace leaves no way to sign back in.
You can link a wallet to a key-only workspace later, on the Settings page. A wallet belongs to one workspace, and once linked it cannot be changed. The message a wallet signs is valid for 5 minutes. A dashboard session lasts 30 days.
Making and revoking keys
Keys are made and revoked on the Keys page. A workspace can have any number of keys, and all of them have the same access. Revoking takes effect at once and cannot be undone; the key stays in the list, marked as revoked. The page also shows when each key was last used; that time is updated at most once every 60 seconds.
Connect your tools
Claude Code
Set three environment variables in the shell you start Claude Code from:
export ANTHROPIC_BASE_URL="https://www.batonapi.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_BATON_KEY"
export ANTHROPIC_API_KEY=""
claudeANTHROPIC_BASE_URLis where Claude Code sends its requests. It adds/v1/messagesitself, so the URL has no/v1at the end.ANTHROPIC_AUTH_TOKENis your Baton key. Claude Code sends it as a bearer token.ANTHROPIC_API_KEYis left empty on purpose. An Anthropic key that is already in your environment would otherwise be sent in place of the token. The empty value keeps that key out of the request.
Claude Code asks for Anthropic model names. Baton grades and routes a request whatever model it names, so nothing in Claude Code's model settings has to change. Its token counting calls are answered by the token counting endpoint. When a request is answered by a model that is not Anthropic's, it is translated on the way; Formats lists what does not cross.
OpenAI-compatible clients
Any client that lets you set an OpenAI base URL works. It needs three settings:
- Base URL
- https://www.batonapi.com/v1
- API key
- YOUR_BATON_KEY
- Model
- baton
| Client | Where the settings go |
|---|---|
| OpenAI SDKs | base_url and api_key in Python, baseURL and apiKey in JavaScript, as in the quickstart. The SDKs also read OPENAI_BASE_URL and OPENAI_API_KEY from the environment. |
| Cursor | In the model settings, switch on the OpenAI base URL override, enter the base URL and the key, and add baton as a custom model. |
| Continue | A model entry with provider: openai, apiBase, apiKey and model: baton. |
| Aider | --openai-api-base, --openai-api-key and --model openai/baton. |
| LangChain | ChatOpenAI with base_url, api_key and model="baton". |
The client has to use the chat completions API. The Responses API, embeddings and the older completions API are not implemented, and a request to them gets status 404.
Anthropic SDKs
Set base_url (baseURL in JavaScript) to https://www.batonapi.com and api_key (apiKey) to your Baton key, as in the quickstart samples. The SDK adds /v1/messages to the base URL and sends the key in x-api-key. Use baton as the model.
API reference
All endpoints are under https://www.batonapi.com. OpenAI clients take https://www.batonapi.com/v1 as their base URL and Anthropic clients take https://www.batonapi.com.
| Endpoint | What it does |
|---|---|
| POST /v1/chat/completions | Routes an OpenAI chat completions request to one of your models. The reply is in the chat completions format. |
| POST /v1/messages | Routes an Anthropic Messages request to one of your models. The reply is in the Messages format. |
| POST /v1/messages/count_tokens | Counts the input tokens of a Messages request without running it. |
| GET /v1/models | Lists the model names a key can ask for: the names Baton defines and the models on your ladder. |
| POST /mcp | The MCP server, for agents that call tools. |
These are the only API endpoints. Before a request is routed, Baton checks that the body is a JSON object with a messages array of at least one message, each with a role its format knows. Everything else in the body is checked by the provider that answers.
The model field
The model name in a request never picks a model from your ladder. It only says how to choose the tier:
batonhas the request graded.baton-hard,baton-mediumandbaton-easyroute it as that tier without grading.- Any other name, including the id of a model on your ladder, is treated like
baton. That is why clients with fixed model names work unchanged.
Names are compared without regard to case. The response headers tell you which model answered.
Token counting
POST /v1/messages/count_tokens takes the body of a Messages request and answers { "input_tokens": 1234 }. When the first switched-on model of your ladder is an Anthropic model with a working key, the count comes from Anthropic for that model and is exact. Otherwise it is an estimate: one token for every four characters of the system prompt, the messages and the tool definitions. The x-baton-token-count response header says which of the two you got.
Model list
GET /v1/models lists baton, baton-hard, baton-medium and baton-easy, then every switched-on model of your ladder once. The list is in OpenAI's format unless the request has an anthropic-version header, in which case it is in Anthropic's.
Request headers
| Header | What it does |
|---|---|
| authorization | Bearer followed by your key. |
| x-api-key | Your key, when there is no bearer token. |
| x-baton-tier | hard, medium or easy: routes the request as that tier and skips the grader. Any other value is ignored. |
| x-baton-conversation | An id of your choosing that names the conversation, so that stickiness does not depend on the text of the request. |
| anthropic-version | Passed on when a Messages request is answered by an Anthropic model; without it, 2023-06-01 is sent. On GET /v1/models its presence selects Anthropic's list format. |
| anthropic-beta | Passed on when a Messages request is answered by an Anthropic model, and dropped otherwise. |
No other header of yours reaches a provider. Your Baton key in particular is never passed on: each provider is called with the key you stored for it.
Response headers
| Header | Value |
|---|---|
| x-baton-model | The id of the model that answered. none when no model did. |
| x-baton-provider | anthropic, openai, google, openrouter or custom; house for the house model; none when no model answered. |
| x-baton-rung | The answering model's place on the ladder, from 0 at the top. -1 for the house model and when no model answered. |
| x-baton-tier | hard, medium or easy: the tier the request was routed as. |
| x-baton-token-count | On the token counting endpoint only: exact or estimate. |
The four routing headers are on every response that got as far as routing: an answer, a provider's refusal of the request that is passed back to you, and the 402, 502 and 503 that mean no model could answer. Errors raised before routing, such as a bad key or a body that is not JSON, carry none of them.
Streaming
Set stream to true in the body and the reply is a stream of server-sent events in the format of your request, with the routing headers on the response as usual.
- When the answering model speaks your format, its events are passed through byte for byte. The one exception: Baton always asks OpenAI-style providers for the usage chunk, and removes it again unless you set
stream_options.include_usageyourself. - When it speaks the other format, the events are translated one by one. A translated chat completions stream ends with
data: [DONE], and its chunks havecreatedset to0. - Another model is only tried before the first byte. If the provider's stream breaks off part of the way through, your stream ends with one error event in your format, and the HTTP status stays 200.
- If you close the connection, the call to the provider is cancelled. What was generated up to then is still counted.
CORS
The API can be called from a browser on any origin. Every response has access-control-allow-origin: * and exposes the five response headers above to scripts. A preflight OPTIONS request is answered with status 204 and may be cached for 24 hours. It allows these request headers, plus any others the browser announces:
authorization, x-api-key, content-type, anthropic-version, anthropic-beta, x-baton-tier, x-baton-conversation
The API never reads cookies, so a page can only act with a key it already holds. Keep keys out of pages you serve to other people.
Limits
| Limit | Value | What happens |
|---|---|---|
| Request body | 25 MB | A larger body gets status 413. The MCP endpoint reads at most 4 MB. |
| Start of the answer | 60 seconds | A provider that has not started answering by then is treated as unavailable and the next model is tried. |
| Whole answer, not streaming | 10 minutes | Counted from the request to the last byte of the provider's reply. A stream has no such limit. |
| Run time of a request | 300 seconds | The longest the inference and MCP routes ask the hosting platform to keep one request running. |
| Exact token count | 15 seconds | When Anthropic takes longer to count, the endpoint answers with an estimate. |
How routing works
A request goes through three steps. It is graded hard, medium or easy. The tier and the state of your ladder decide which models to try, in which order. The first one that answers is the one you hear from.
The grader
Grading follows fixed rules and calls no model, so the same request always gets the same grade. What is graded is the last thing a person typed: the latest user message that contains text. Turns that only carry tool results are skipped.
Every request starts at 45 points. Signals add or take away points, and the total, kept between 0 and 100, decides the tier: 65 or more is hard, 35 to 64 is medium, and anything lower is easy.
| Group | Points | Signal |
|---|---|---|
| Hard work | +30 | The message asks for a design, a refactor, a migration, debugging, root-cause analysis, a proof, a plan, a security review or a performance investigation. |
| Hard work | +18 | An architecture question, a comparison of approaches, a move from one technology to another, a failure that is hard to reproduce, concurrency, or work across many files or steps. |
| Hard work | +12 | A hard subject that is only mentioned: migrations, security, optimisation, distributed systems, production risk, algorithms. |
| Easy work | -30 | An easy verb is in charge of the message: rename, fix a typo or the wording, format, summarise, translate, convert between simple formats, or write a commit message or a title. Also a greeting or thanks that asks for nothing. |
| Easy work | -20 | A lookup or a definition ("what is", "who", "how many") of at most 240 characters, with no code attached and no other work asked for. |
| Easy work | -20 | A bare follow-up such as "ok", "yes" or "go ahead", outside an agent loop. |
| Easy work | -15 | A yes-or-no question of at most 120 characters, or a request for a list. |
| Attachments | +5, +9, +13 | The message is at least 400, 1,200 or 4,000 characters long. |
| Attachments | +6 | It includes code. |
| Attachments | +12, +4 | It includes a stack trace, or other error output. |
| Attachments | +10, +20 | It asks for three or four separate things, or for five or more. |
| Brevity | -5, -15 | The message is at most 59 characters long, or at most 24. |
| Session | +12 | The request asks for reasoning. |
| Session | +5, +9 | The whole request holds at least 24,000 characters, or at least 100,000. |
| Session | +8 | Eight or more tools are offered, and the message is neither easy work nor a bare follow-up. |
| Session | +6 | The conversation includes an image. |
- Within hard work and within easy work, the strongest signal counts in full and each further one counts half. Hard-work signals about the same thing count once. Hard work adds at most 45 points and easy work takes away at most 40.
- When an easy verb is in charge and nothing else is asked, the rest of the message is material to work on: hard words in it, its length and its code do not count.
- A "how do I" question of at most 200 characters is a question about hard work, not the work itself: its hard-work signals count 18 points at most.
- Brevity is not counted when a 30-point signal fired, or in an agent loop, which is a conversation that already holds tool calls or tool results. There a short message steers work in progress.
- Session signals add at most 19 points together. They can tip a request that already carries some weight, but they never make a plain one hard.
- A request asks for reasoning when a Messages request has
thinkingset to anything butdisabledorbetween_tools, or whenoutput_config.effortorreasoning_effortismediumor higher. - With eight or more tools offered, an easy job such as a summary or a translation is not graded easy when its subject is something in the workspace, such as a repository or a pull request, and the material is not in the message. The agent has to go and read it first.
- The word lists are English. A request in another language is graded on its length, its attachments and the session. Of a very long message, the first 12,000 and the last 4,000 characters are read for wording.
Three requests, each sent on its own with no tools and no reasoning settings:
Refactor the auth module to use dependency injection.
- Every request starts at
- 45
- asks for a refactor
- +30
- Score and tier
- 75, hard
The message is short, but nothing is taken off for that: a verb worth 30 points outweighs brevity.
Explain what a mutex is in one sentence.
- Every request starts at
- 45
- concurrency
- +18
- short message
- -5
- Score and tier
- 58, medium
This is the request from the quickstart. It opens with a verb that asks for work, so it is not treated as a lookup.
Rename getUser to fetchUser.
- Every request starts at
- 45
- asks for a rename
- -30
- short message
- -5
- Score and tier
- 10, easy
An easy verb is in charge of the whole message. Reasoning settings, tools and a large context could add at most 19 points, which still leaves it easy.
Only the tier is reported, in the x-baton-tier response header. The grader on the home page runs the same rules on whatever you type.
The ladder policy
The ladder is your models in order, with rung 0 at the top. A rung is usable when it is switched on, its provider has a working key, it is not cooling down, and its spend today is below its daily budget. A rung with no budget is never over it.
| Tier | Starts at | Stays on a rung while |
|---|---|---|
| hard | Rung 0 | Less than 100% of the rung's daily budget is spent. |
| medium | Rung 0 | Less than 50% is spent. On the last rung: less than 100%. |
| easy | The last rung | Less than 100% is spent. |
- If the conversation is pinned to a rung and that rung is usable, it is the first choice, whatever the tier and the percentages say.
- Otherwise the router walks down from the rung the tier starts at and takes the first usable rung that is under the tier's percentage.
- If no rung is under it, the router walks down again and takes the first usable rung. The 50% mark is a preference; a spent budget is a limit.
- The fallbacks are every usable rung below the first choice, in order, and after them the house model. Each is tried only when the one before it could not answer.
Some things follow from these rules that are easy to miss:
- Apart from a pin, nothing moves a request up the ladder past the rung its tier starts at. An easy request without a pin therefore uses the last rung, then the house model, and never a rung above, even when those are free.
- The last rung is the last row of the ladder, whether or not it is switched on. When that row is switched off, has no working key, is cooling down or is over budget, easy requests go straight to the house model.
- With a single rung, all three tiers use it until its budget is spent.
- Budgets are checked before a request is sent, so the requests already in flight can take a rung past its budget.
An example. Rung 0 is claude-opus-4-8 with a budget of $10.00 a day, rung 1 is claude-sonnet-4-6 with $5.00 a day and less than half of it spent, and rung 2 is claude-haiku-4-5 with no budget. This is where a request without a pin goes first:
| Spent today on rung 0 | hard | medium | easy |
|---|---|---|---|
| $0.00 to $4.99 | Rung 0 | Rung 0 | Rung 2 |
| $5.00 to $9.99 | Rung 0 | Rung 1 | Rung 2 |
| $10.00 or more | Rung 1 | Rung 1 | Rung 2 |
Budgets and the daily reset
A rung's spend is the cost of the requests it answered today: input tokens times the input price plus output tokens times the output price, at the prices you set on the rung in US dollars per million tokens. Token counts are the ones the provider reports. When a provider reports none, they are estimated at four characters per token.
- The day is the UTC day. Every rung's spend starts again from zero at 00:00 UTC.
- Cached input tokens are counted at the full input price.
- Only answers count. A request the provider refused adds nothing, and a stream that was cut short adds what had been generated.
- A rung whose prices are both 0 never accumulates spend, so its budget is never reached. Set prices on every rung that has a budget.
This spend is an estimate that exists for routing. What you owe a provider is on that provider's invoice.
Rate limits, cooldowns and provider errors
What a provider answers decides whether the next candidate is tried. A cooldown belongs to one rung, and every request skips that rung until the cooldown has passed.
| The provider answers | What happens |
|---|---|
| 429, 529 or 402 | The model is out for now: a rate limit, an overload, or a provider account without credit. The rung cools down and the next candidate is tried. The cooldown is the time the provider asks for in retry-after-ms or retry-after, kept between 5 seconds and 60 minutes, or 60 seconds when it names none. |
| 401 or 403 | The provider key is marked as rejected, whatever the provider's reason, and the next candidate is tried. Every rung of that provider is then skipped until you save the key again or test it with success on the Ladder page. |
| 500 and above, 408, a redirect, no connection, or no start of an answer in time | The rung cools down for 30 seconds and the next candidate is tried. |
| 200 with a body that is not an answer | The same as an outage. This covers a body that is not JSON, JSON that holds only an error, and a reply to a streaming request that is not an event stream. |
| Any other 4xx, such as 400, 404 or 422 | The request itself is at fault, so no other model is tried. The provider's error is passed back with the same status, in the format of your request. |
The house model is the last candidate, so there is nothing to fall back to: when it fails in any of these ways, the request ends with status 502. When no candidate was left at all, it ends with 402 or 503.
Conversation stickiness
Prompt caches and thinking blocks belong to one model, so by default a conversation stays on the rung that answered it last. The API keeps no conversations, so one is recognised by a fingerprint: a hash of your workspace, the system prompt and the text of the first user message, which are the same on every turn. Only the hash is stored.
- After each answer from a rung, the conversation is pinned to that rung for 6 hours, counted from its latest answer.
- The next turn goes to the pinned rung first, as long as it is usable. When it is not usable, the turn is routed as if there were no pin. When it fails, the rungs below it are tried. Either way the pin moves to the rung that answers.
- The pin does not look at the tier. A conversation that opens with an easy message is pinned to the last rung, and a later hard message in it is answered there too.
- Answers from the house model set no pin.
- In a chat completions request, the system prompt is the system and developer messages that come before the first other message. When the first user message has no text and no id is sent, there is nothing to key on and the request is routed without a pin.
To name a conversation yourself, send any id in the x-baton-conversation header. The fingerprint is then made from your workspace and that id, and the text of the request no longer matters. Use it when your client rewrites the system prompt or the first message between turns, when separate conversations open with the same text, or to start a conversation afresh under a new id.
curl https://www.batonapi.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_BATON_KEY" \
-H "Content-Type: application/json" \
-H "x-baton-conversation: ticket-4821" \
-d '{
"model": "baton",
"messages": [{"role": "user", "content": "Why does the nightly export time out?"}]
}'To turn stickiness off for the whole workspace, switch off "Keep conversations on one model" on the Settings page. Pins are then neither read nor written.
Forcing a tier
There are two ways to skip the grader. Send hard, medium or easy in the x-baton-tier header, or ask for the model baton-hard, baton-medium or baton-easy. When a request has both, the header wins.
curl https://www.batonapi.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_BATON_KEY" \
-H "Content-Type: application/json" \
-H "x-baton-tier: hard" \
-d '{
"model": "baton",
"messages": [{"role": "user", "content": "Is this lock ordering safe?"}]
}'A forced tier changes where the request starts on the ladder and nothing else. Budgets and cooldowns apply as usual, and a conversation pin still comes first. Baton reports the tier it used in the response header either way.
Formats
You can send either format, whatever is on your ladder. Anthropic models are called in the Messages format. OpenAI, Google, OpenRouter, custom endpoints and the house model are called in the chat completions format. When your request and the answering model differ in format, the request is translated on the way in and the reply on the way back, streams included.
A chat completions request answered by an Anthropic model
- System and developer messages, wherever they stand in the conversation, are joined into the
systemfield. - Consecutive messages of one role become one turn.
toolmessages becometool_resultblocks at the start of a user turn, and an assistant'stool_callsbecometool_useblocks. - Function tools become Anthropic tools with the same name, description and schema. In
tool_choice,requiredbecomesanyand a named function becomes that tool.parallel_tool_calls: falsebecomesdisable_parallel_tool_use. - Tool call ids are reduced to letters, digits,
_and-, the characters Anthropic accepts, and made unique within the request. - An
image_urlpart becomes an image block: base64 for adata:URL, a URL source for any other. Audio and file parts are left out. max_completion_tokens, or elsemax_tokens, becomesmax_tokens. Anthropic requires it, so a request with neither gets 16,000, or 64,000 when it streams.stopbecomesstop_sequences, andtemperatureis kept within 0 to 1.reasoning_effortbecomesoutput_config.effort, withminimalandnonesent aslow. Aresponse_formatof typejson_schemabecomesoutput_config.format.userbecomesmetadata.user_id.- Left out, because the Messages format has no place for them:
n,logprobs,seed,logit_bias, the two penalties,response_formatof typejson_object, a tool'sstrictflag, and the olderfunctionsparameter andfunctionmessages.
The translated request is then fitted to the model like any other Messages request, as described under models that share a format.
A Messages request answered by an OpenAI-style model
systembecomes a system message at the start.- In a user turn, each
tool_resultblock becomes atoolmessage, followed by one user message with the rest. A failed result hasError:put before its text. Images inside tool results move to the user message, because a tool message holds only text. - Image blocks become
image_urlparts; base64 data is sent as adata:URL. - In an assistant turn, the text blocks are joined and the
tool_useblocks becometool_calls. - Your own tools become function tools. In
tool_choice,anybecomesrequiredand a named tool becomes that function.disable_parallel_tool_usebecomesparallel_tool_calls: false. max_tokensis sent to OpenAI asmax_completion_tokensand to other providers asmax_tokens.stop_sequencesbecomesstop; OpenAI gets the first 4, the most it accepts.temperatureandtop_ppass through.top_kis left out.output_config.effortbecomesreasoning_effortfor OpenAI models that accept it, withxhighandmaxsent ashigh. Other providers do not get it. Anoutput_config.formatof typejson_schemabecomes aresponse_format.- Left out: the
thinkingsetting,cache_controlmarkers, document blocks andmetadata.
Replies
Text stays text, and tool calls become tool calls in the other format, with their arguments turned from a JSON string into an object or back. Stop reasons are mapped like this when an Anthropic model answers a chat completions request:
| Messages stop_reason | Chat completions finish_reason |
|---|---|
| end_turn, stop_sequence, pause_turn | stop |
| max_tokens, model_context_window_exceeded | length |
| tool_use | tool_calls |
| refusal | content_filter |
In the other direction, length becomes max_tokens and content_filter becomes refusal. Any other reply that calls tools is tool_use, whatever reason the provider gave, and everything else is end_turn.
- Usage keeps its meaning. Anthropic counts cached input apart from
input_tokens; in a chat completions replyprompt_tokensis the sum of all input, with the cached part inprompt_tokens_details.cached_tokens. In a Messages reply, cached tokens an OpenAI-style provider reports are taken out ofinput_tokensand put incache_read_input_tokens. - Ids are kept as the provider gave them, so a chat completions reply can carry an id that starts with
msg_. Themodelfield of a translated reply is the id of the model that answered. - A refusal from an OpenAI-style model arrives as text. Reasoning text is not part of a translated reply, in either direction.
- A provider's error keeps its message and gets the error shape of your format.
Between models that share a format
When your request and the answering model share a format, the body is sent as you wrote it, with model replaced. Messages, the system prompt and tools are never touched. Only parameters the target model would reject are changed.
For a Messages request sent to an Anthropic model:
temperature,top_pandtop_kare removed for models that refuse sampling parameters. For models that taketemperatureortop_pbut not both,top_pis removed when both are set.max_tokensis lowered to the model's limit, where it is known.output_config.effortis lowered to the nearest level the model accepts, or removed.thinkingis fitted: a token budget becomes adaptive thinking on models that have no budgets, and adaptive thinking becomes a budget of half ofmax_tokens, between 1,024 and 32,000 tokens, on models that only have budgets.disabledis removed for models that cannot switch thinking off.- A
tool_choicethat forces a tool becomesautoon models that do not accept forcing. - A model id that is not recognised as a Claude model gets the strictest of these rules.
For a chat completions request sent to an OpenAI-style model:
- A stream always asks the provider for usage with
stream_options.include_usage. The extra chunk is taken out of your stream again unless you asked for it. - For OpenAI itself,
max_tokensis renamed tomax_completion_tokens, andreasoning_effortis removed for models whose names start withgpt-3.5,gpt-4orchatgpt-4, which reject it.
What cannot cross
- Anthropic server tools. Tools that Anthropic runs, such as web search and code execution, exist only there. When a Messages request goes to an OpenAI-style model they are left out, together with a
tool_choicethat forces one and any of their calls and results in the history. - Thinking blocks. Only the model that wrote them can use them. They are left out when a Messages request goes to an OpenAI-style model, and thinking in a reply is not translated. Between two Anthropic models they are forwarded as they are.
- Prompt caches. A cache belongs to one model at one provider. A chat completions request sent to an Anthropic model sets no cache markers, so nothing is cached there.
- Beta features. The
anthropic-betaheader reaches Anthropic models only.
This is why conversations are sticky by default: every move to another model gives up whatever of the above the conversation was using.
The house model
The house model is a model this deployment provides to every workspace. It is not on your ladder and needs no provider key of yours. Which model it is and what it costs is up to the operator; the Ladder page shows both under your last rung.
When it answers
The house model is the last candidate of every request. It is asked when:
- your ladder is empty,
- no rung is usable for the request, or
- every rung that was tried failed.
It then answers only if all of this holds:
- The operator has configured a house model.
- "Fall back to the house model" is on for your workspace. It is on by default; the switch is on the Settings page.
- Your workspace has credit, unless the operator has set the house model's price to zero.
An answer from the house model has house in x-baton-provider and -1 in x-baton-rung. It sets no conversation pin and counts toward no rung's budget.
How it is priced
Each answer is charged to your credit: input tokens times the input price plus output tokens times the output price, both per million tokens. The counts are the ones the house model reports, or an estimate at four characters per token when it reports none. The charge is made after the answer, and the Requests page shows it for each request.
At zero credit
With no credit, the house model is not asked and the request ends with status 402 and the code insufficient_credit. The check looks at the balance your workspace had when the request arrived, not at the size of the request: a request that starts with any credit is answered, and its charge takes the balance to zero at most.
When the house model is not configured or is switched off, the status is 503 instead. When it was asked and failed, the status is 502. The house model is not retried.
For operators
The house model is any endpoint that speaks the chat completions format. Set HOUSE_BASE_URL to the part of its address before /chat/completions. Requests are sent there with model set to HOUSE_MODEL and, when HOUSE_API_KEY is set, that key as a bearer token. The two prices are HOUSE_PRICE_IN_USD_PER_MTOK and HOUSE_PRICE_OUT_USD_PER_MTOK. All of them are in the table of variables.
- Without
HOUSE_BASE_URLthere is no house model. In development a built-in mock stands in for it: it calls nothing, costs nothing, and answers with a fixed text that names the tier of the request.HOUSE_MOCKswitches it on or off. - With both prices at 0 the house model is free, and it answers workspaces that have no credit.
REQUIRE_CREDITmakes credit a condition for the whole API. A workspace without credit then gets 402 for every request, including those its own keys could have answered.
Credits
Credit is a prepaid balance in US dollars that belongs to a workspace. It pays for answers from the house model and for nothing else. Requests that your own provider keys answer are billed by those providers and leave your credit alone.
Adding credit
On this deployment, credit is granted by the operator. There is nothing to buy in the app: the operator adds credit to a workspace by its id, which is on the Settings page, or by the wallet linked to it. A new workspace starts with credit when the operator has set a starter amount.
Where to see it
- The Credits page shows the balance, what was used today (UTC) and the additions so far.
- The Requests page shows the credit charged for each request.
- An agent can read the balance with the
balancetool of the MCP server.
For operators
POST /api/credits/grant adds credit to a workspace. It exists only when ADMIN_SECRET is set, and takes that secret as a bearer token. The body names the workspace by workspaceId or by wallet, one of the two, with usd above zero and an optional note of up to 200 characters.
curl https://www.batonapi.com/api/credits/grant \
-H "Authorization: Bearer $ADMIN_SECRET" \
-H "Content-Type: application/json" \
-d '{"workspaceId": "ws_3fKx8mQ2LpZr7TnVbY4c", "usd": 5, "note": "Trial credit"}'{
"ok": true,
"workspaceId": "ws_3fKx8mQ2LpZr7TnVbY4c",
"grantedMicros": 5000000,
"balanceMicros": 5000000
}Amounts in the reply are in millionths of a dollar. A wrong secret gets status 401, and a workspace that does not exist gets 404; a wallet counts only once its owner has signed in with it. STARTER_CREDIT_USD gives every new workspace a first amount without a call.
Your own models
Baton routes between models you bring. On the Ladder page you store a key for each provider you use, then put that provider's models on the ladder in the order they should be tried.
Providers
| Provider | Called in | Base URL |
|---|---|---|
| Anthropic | Messages | https://api.anthropic.com |
| OpenAI | Chat completions | https://api.openai.com/v1 |
| Chat completions | https://generativelanguage.googleapis.com/v1beta/openai | |
| OpenRouter | Chat completions | https://openrouter.ai/api/v1 |
| Custom endpoint | Chat completions | The base URL you enter |
- A workspace holds one key per provider. Saving another key for the same provider replaces the first.
- A key is encrypted with AES-256-GCM before it is stored. After that it is shown only in masked form and used only to call that provider for you.
- A key is tried against the provider as soon as you save it, and again whenever you press Test. A key the provider rejects is marked as such, and the models that depend on it are skipped.
- Removing a key leaves its models on the ladder. They are skipped until the provider has a key again.
Custom endpoints
The custom provider is any endpoint that speaks the chat completions format: vLLM, Ollama, LM Studio, a hosted gateway, your own server. A workspace has one. Enter the part of its address before /chat/completions; the model list is read from /models under the same address, and the key is sent as a bearer token. The key field cannot be empty, so for an endpoint that needs no key, enter any word.
The base URL has to pass these rules:
- It is an
httporhttpsURL with no user name or password in it. A query string, a fragment and trailing slashes are removed. - Unless the operator allows private addresses, which is the default only in development, it has to use
httpsand name a public host. Refused arelocalhost, private, link-local and carrier-grade NAT addresses, names without a dot, and names that end in.localhost,.localor.internal. - Under the same condition, the host name is looked up again before every request, and the request is not sent when any of its addresses is private.
- Redirects are never followed, for any provider. A redirect counts as the endpoint being unavailable.
Models, prices and budgets
A ladder holds up to 8 models. The picker lists the models your key can call, read live from the provider, and you can type a model id instead. A new model goes to the bottom; move it to where it belongs. Each model on the ladder has:
- Prices for input and output, in US dollars per million tokens. When you pick a model they are filled in where a list price is known: from a built-in table for current Claude models, and from OpenRouter's public price list for OpenAI, Google and OpenRouter models. Check them. Prices are what a rung's spend is worked out from, and a rung priced at 0 never reaches its budget.
- A daily budget in US dollars, or none. Without a budget a rung is left only when its provider rate-limits or fails. A budget of $0 is refused: switch the model off instead.
- A switch. A model that is switched off keeps its place and is skipped. Keep in mind that easy requests start at the last row, switched on or not.
- A label, if you want one. It is shown in the dashboard in place of the model id.
Changes apply from the next request. The page also shows what each model has spent today and whether it is ready, cooling down, over budget or without a working key.
The Savings page shows what the ladder saved over the last 7 or 30 days. It prices the tokens of every answered request at the current prices of your top model (the first one that is switched on) and sets that against what was paid: the cost at the prices of the model that answered, plus credit for the house model. It is an estimate, since another model would not have used exactly the same number of tokens, and it needs prices on the top model.
Alerts
On the Settings page you can give a webhook address, and Baton posts a message to it when a model passes half of its daily budget, when a model's daily budget is spent, and when a provider rate-limits a model. Each budget message is sent once a day per model, and a rate-limited model is announced at most once every 15 minutes. You choose which of the three you want, and can send a test message.
A Discord webhook gets the message as content and a Slack one as text. Any other address gets a JSON body with text, event (budget_half, budget_spent or rate_limited), workspaceId, model and the amounts or the time that go with the event. The address must be public and use https, it is stored encrypted, and a webhook that redirects is not followed.
MCP server
Baton is also an MCP server, so an agent can ask a routed question and read your ladder, usage and balance as tools.
- URL
- https://www.batonapi.com/mcp
- Transport
- Streamable HTTP
- Authentication
- Authorization: Bearer YOUR_BATON_KEY
The key is an ordinary API key from the Keys page. A client configuration looks like this; it is the form Claude Code reads from .mcp.json, and other clients take the same three facts under their own names.
{
"mcpServers": {
"baton": {
"type": "http",
"url": "https://www.batonapi.com/mcp",
"headers": {
"Authorization": "Bearer YOUR_BATON_KEY"
}
}
}
}On the wire
- The server is stateless. Every message is a
POSTthat carries one JSON-RPC message and is answered as JSON. There are no sessions and no event streams;GETandDELETEget status 405. - The key is checked on every message,
initializeincluded. Without a valid one the answer is status 401 with aWWW-Authenticate: Bearerheader. There is no OAuth flow: set the key as a header in the client.x-api-keyis read as well. - A client that speaks HTTP itself has to send
Content-Type: application/jsonandAccept: application/json, text/event-stream. - Batches are refused with status 400, and a body over 4 MB with status 413.
Tools
| Tool | Input | What it returns |
|---|---|---|
| ask | prompt (text, required), system (text), tier (hard, medium or easy), max_tokens (a whole number of 1 or more) | Sends one prompt through the router and returns the reply, followed by a line that names the model, provider, rung and tier that answered. As data: text, model, provider, rung, tier and finish_reason. |
| ladder | None | Your models in ladder order. For each: its state (ok, disabled, no_key, cooling_down or out_of_budget), what it has spent today and its daily budget. Then the house model, and whether it is switched on for the workspace. |
| usage | None | Today's traffic, in UTC: requests, tokens in and out, the estimated cost on your own provider keys and the credit used by the house model, by tier and by answering model. |
| balance | None | Your credit balance in US dollars. |
- Every result is readable text, with the same facts as structured content.
askis an ordinary routed request. It is graded unlesstieris set, it counts toward budgets and credit, and it appears on the Requests page under the key that made it. It does not stream: the tool returns when the whole reply is there.- Each
askstands alone, with no memory of earlier calls. Stickiness still applies: the same prompt with the same system prompt goes to the rung that answered it before. - When the router cannot answer, the tool result is marked as an error and carries the message of the error an API call would have got.
Errors
An error comes in the shape of the format you sent the request in, with an HTTP status that says what kind of error it is.
The two shapes
{
"error": {
"message": "No API key was sent. Create a key in the Baton dashboard at https://www.batonapi.com and send it as \"Authorization: Bearer <key>\" or in the x-api-key header.",
"type": "authentication_error",
"param": null,
"code": "invalid_api_key"
}
}{
"type": "error",
"error": {
"type": "authentication_error",
"message": "No API key was sent. Create a key in the Baton dashboard at https://www.batonapi.com and send it as \"Authorization: Bearer <key>\" or in the x-api-key header."
}
}The chat completions shape carries the router's code in error.code. The Messages shape has no field for it, so there you tell errors apart by the status and by error.type.
| Status | Chat completions error.type | Messages error.type |
|---|---|---|
| 400 | invalid_request_error | invalid_request_error |
| 401 | authentication_error | authentication_error |
| 402 | insufficient_quota | permission_error |
| 413 | invalid_request_error | request_too_large |
| 429 | rate_limit_error | rate_limit_error |
| 500 | server_error | api_error |
| 502 | server_error | api_error |
| 503 | server_error | api_error |
The router's own errors
| Status | Code | When | What to do |
|---|---|---|---|
| 400 | invalid_json | The body is not valid JSON. | Send a JSON body. |
| 400 | invalid_request | The body is not a JSON object, has no messages array or an empty one, or holds a message without a role its format knows. The message says which. | Correct the request. Sending it again unchanged gives the same error. |
| 401 | invalid_api_key | No key was sent, or the key is unknown or revoked. | Send a key from the Keys page in one of the two key headers. |
| 401 | workspace_not_found | The key is valid but its workspace no longer exists. | Use a key of a workspace that exists. |
| 402 | insufficient_credit | Only the house model was left to answer and the workspace has no credit. Also every request of a workspace without credit, where the operator requires credit. | Add credit, or wait until a model on your ladder is usable again. The message names what stopped each one. |
| 429 | rate_limited | The workspace sent more requests in one clock minute than the router allows: 120 unless the operator set another number. This is the router's own limit; a provider's rate limit never reaches you as an error, it moves the request down the ladder. | Wait the number of seconds in the Retry-After header, or send requests more slowly. |
| 413 | request_too_large | The body is larger than 25 MB. | Send less: a shorter history, or fewer and smaller images. |
| 500 | internal_error | Something failed inside the router. | Try again in a moment. If it keeps happening, tell the operator. |
| 502 | upstream_error | The house model was asked and failed. Also the last event of a stream that the provider broke off. | Send the request again. |
| 503 | no_model_available | No model on the ladder could answer, and the house model is not configured or is switched off. The message names what stopped each model. | Act on the message: wait out a cooldown or the daily reset, replace a rejected provider key, add a model, or switch the house model on. |
Keys are made on the Keys page, credit is on the Credits page, and provider keys, models and the state of each rung are on the Ladder page.
Errors from a provider
Most provider failures never reach you: the router tries the next model. The exception is a provider that refuses the request itself, with a 4xx status other than 401, 402, 403, 408 and 429. That status and the provider's message are passed back to you.
- In the chat completions shape, the provider's own
type,codeandparamare kept where it sent them. - In the Messages shape,
error.typeis Anthropic's own when Anthropic sent the error, and otherwise the one that goes with the status. - When the provider's body holds no message that can be read, the message is "The model provider returned an error" with the status.
Codes in the request log
The Requests page shows a code under the status of every request that did not end well. Errors raised before routing, such as a bad key or a body that is not JSON, are not logged.
| Code | Meaning |
|---|---|
| provider_error | A provider refused the request itself. The status is the provider's. |
| stream_interrupted | The provider's stream broke off part of the way through. The status is 200. |
| client_closed | You closed the connection. The status is 200 when a stream was under way, and 499 when nothing had answered yet. |
| upstream_error | The house model failed: status 502. |
| no_model_available | Nothing could answer: status 503. |
| insufficient_credit | Only the house model was left and there was no credit: status 402. |
Privacy
The request log
One row is written for every request that is routed. It holds:
- an id for the row
- the workspace
- which of your API keys made the request
- the time
- the format of the request, chat completions or Messages
- the tier it was routed as
- the rung that answered
- that rung's place on the ladder at the time
- the provider
- the model
- the number of input tokens
- the number of output tokens
- the estimated cost at the prices you set
- the credit charged
- how long the request took
- the HTTP status of the answer
- whether it streamed
- an error code, when there was one
- the conversation fingerprint, which is a hash
That is all of it. Rows are kept until the operator deletes them; nothing removes them automatically.
What is never stored
- Prompts, system prompts and replies.
- Tool definitions, tool calls and tool results.
- Images and other attachments.
- The headers of your requests.
- Your API keys. Only a hash of each is kept, with a masked form for display.
The grader and the conversation fingerprint read the text of a request in memory while it is being routed. The fingerprint that is stored is a hash, and the text cannot be recovered from it.
What else is stored
- The workspace: its name, the address of a linked wallet, its credit balance and its routing settings.
- API keys: the name you gave each, its hash and masked form, and when it was made, last used and revoked.
- Provider keys, encrypted with AES-256-GCM, each with a masked form and, for a custom endpoint, its base URL.
- The ladder, and for each rung and UTC day its spend, token counts and number of requests.
- Conversation pins: a fingerprint and the rung it is pinned to. A pin expires after 6 hours and is deleted when the workspace next starts a new conversation.
- Additions of credit: the amount, where it came from and a note.
A dashboard session is a signed cookie; no list of sessions is kept on the server. Signing in with a wallet signs a message in the wallet, which never hands over a key.
Where a request goes
A request is sent to the provider whose model answers it, with the key you stored for that provider, or to the house model's endpoint. What happens to it there is up to that provider. Baton sends the body of the request and, to Anthropic, the anthropic-version and anthropic-beta headers. Your Baton key, your cookies and your other headers are not sent on. Requests to OpenRouter also carry the name and address of this site, which is how OpenRouter attributes traffic to an app.
Self-hosting
These notes are for the operator of a deployment. Baton is a Next.js application that keeps its data in one Postgres database.
GET /api/healthanswers{ "ok": true }when the database can be reached, and status 503 when it cannot.- Inference and MCP requests ask the platform for up to 300 seconds of run time. A model that writes for longer than the platform allows is cut off.
- The API reads request bodies of up to 25 MB. Vercel limits a function's request body to 4.5 MB, so there the lower limit applies.
- A truthy flag is
trueor1. A number that cannot be read falls back to its default. - The product's name, the key prefix, the header names and the model names all come from one file,
src/lib/brand.ts. A rename there changes them everywhere, these docs included.
Environment variables
"Production" below means NODE_ENV is production. How the house model variables work together is described under the house model.
| Variable | What it does | When it is not set |
|---|---|---|
| DATABASE_URL | The Postgres connection string. The tables are created on the first connection. | An embedded database in .data/pglite, for development. On Vercel, every request that needs the database fails without it. |
| SESSION_SECRET | Signs session cookies and wallet sign-in messages. Use at least 32 characters. Changing it signs everyone out. | Required in production. A fixed development value otherwise. |
| ENCRYPTION_KEY | Encrypts stored provider keys. Any string; it is hashed to a 256-bit key. After a change, the keys stored before it cannot be read and have to be saved again. | Required in production. A fixed development value otherwise. |
| ADMIN_SECRET | The bearer secret of the endpoint that grants credit. | The endpoint answers 404. |
| STARTER_CREDIT_USD | Credit, in US dollars, given to every new workspace. | 0 in production, 1 in development. |
| REQUIRE_CREDIT | When true, the API answers 402 to a workspace without credit, even for requests its own provider keys could answer. | false |
| ALLOW_PRIVATE_UPSTREAMS | When true, custom endpoints may use http and may point at localhost and private networks. | false in production, true in development. |
| HOUSE_MOCK | When true and HOUSE_BASE_URL is not set, a built-in mock answers as the house model. | false in production, true in development. |
| HOUSE_BASE_URL | The house model's endpoint, in the chat completions format, without /chat/completions. | No house model. |
| HOUSE_API_KEY | Sent to the house model's endpoint as a bearer token. | No key is sent. |
| HOUSE_MODEL | The model id sent to that endpoint and reported in the model response header. | house |
| HOUSE_LABEL | The house model's name in the dashboard. | The value of HOUSE_MODEL, or "House model". |
| HOUSE_PRICE_IN_USD_PER_MTOK | What the house model charges to credit per million input tokens, in US dollars. | 0.2 |
| HOUSE_PRICE_OUT_USD_PER_MTOK | The same per million output tokens. | 0.8 |
| RATE_LIMIT_REQUESTS_PER_MINUTE | API requests one workspace may send in a clock minute before it gets a 429. 0 switches the limit off. | 120 |
| RATE_LIMIT_SIGNUPS_PER_HOUR | Workspaces one address may create in an hour by signing in without a wallet. 0 switches the limit off. | 5 |
| RATE_LIMIT_WALLET_CHALLENGES_PER_HOUR | Wallet sign-in messages one address may ask for in an hour. 0 switches the limit off. | 30 |
Other variables
| Variable | What it does | When it is not set |
|---|---|---|
| NEXT_PUBLIC_SITE_URL | The public address of the deployment. Every URL in these docs and in the dashboard's samples, and the site named in the wallet sign-in message, comes from it. Read when the app is built. | http://localhost:3000 |
| NEXT_PUBLIC_X_URL | A link to an X account, shown in the site header. Read when the app is built. | No link. |
| NEXT_PUBLIC_GITHUB_URL | A link to the source repository, shown in the site footer. Read when the app is built. | No link. |
| PROVIDER_BASE_URL_<ID> | Replaces a provider's base URL for the whole deployment, for example to send its traffic through a gateway. <ID> is one of ANTHROPIC, OPENAI, GOOGLE, OPENROUTER, CUSTOM. | The provider's own address. |
| PGLITE_DIR | Where the embedded development database keeps its files. Used only without DATABASE_URL. | .data/pglite |