Skip to content
baton
On this page: Introduction

/docs

Baton docs

How to point a client at Baton, how it grades and routes a request, and every header, limit and error of its API.

Introduction

Baton is an LLM router with an OpenAI-compatible API and an Anthropic-compatible API: you point a client at it and send one key. It grades how hard each request is and sends it to one of the models you have put in order on your ladder, using your own provider keys. Hard requests stay on the model at the top, medium requests leave a model once half of its daily budget is spent, and easy requests go to the model at the bottom. When no model on your ladder can answer, the house model answers and is paid from your prepaid credit.

Quickstart

1. Get a key

Open the sign-in page and sign in with a Solana wallet, or choose to continue without one. The first sign-in creates a workspace and its first API key. The key is shown once, so copy it before you go on. If you lose it, create another on the Keys page.

2. Send a request

Point your client at one of the two base URLs and put your key where the samples say YOUR_BATON_KEY. The model name baton means "let the router decide".

OpenAI-style base URL
https://www.batonapi.com/v1
Anthropic-style base URL
https://www.batonapi.com
Model name
baton
Terminal
curl https://www.batonapi.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_BATON_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "baton",
    "messages": [{"role": "user", "content": "Explain what a mutex is in one sentence."}]
  }'

A new workspace has no models on its ladder, so the house model answers and the request is paid from credit. When the workspace has no credit the request fails with status 402, and on a deployment without a house model it fails with 503. To route between models of your own, add a provider key and a model.

To try a request without a client, use the Playground in the app. It sends a real request through the router and shows which model answered.

3. Read the response headers

Every routed response says where it went. With curl, add -i to see the headers:

x-baton-model: claude-opus-4-8
x-baton-provider: anthropic
x-baton-rung: 0
x-baton-tier: medium
  • x-baton-model is the model that answered, by the id its provider knows it by.
  • x-baton-provider is the provider that served it, or house for the house model.
  • x-baton-rung is the model's place on your ladder. 0 is the top and -1 is the house model.
  • x-baton-tier is what the request was routed as: hard, medium or easy. The sample request is graded medium; the grader section shows why.

The Requests page in the app lists the same facts for every request Baton has routed for you.

Keys and sign-in

Sending a key

Every request to the API and to the MCP server carries an API key, in either of two headers. OpenAI clients send the first, Anthropic clients the second:

Authorization: Bearer YOUR_BATON_KEY
x-api-key: YOUR_BATON_KEY

When a request carries both, the bearer token is the one that is checked. An Authorization header that is not a bearer token is ignored in favour of x-api-key. A missing, unknown or revoked key gets status 401.

What is stored

A key is bt_ followed by 40 letters and digits. Only its SHA-256 hash is stored, next to a masked form for the Keys page: the prefix with the four characters after it, and the last four. The key itself is shown once, when it is made, and cannot be shown again.

Workspaces and sign-in

Keys belong to a workspace. There are two kinds of workspace, and they differ only in how you get back in.

  • Wallet-linked. You sign in with a Solana wallet by signing a message. Nothing is sent on chain and it costs nothing. The first sign-in with a wallet creates its workspace and a first key named "Default"; later sign-ins with the same wallet open the same workspace.
  • Key-only. "Continue without a wallet" creates a workspace and a first key. That key, or any other active key of the workspace, is then the only way to sign back in: paste it on the sign-in page. Revoking the last active key of a key-only workspace leaves no way to sign back in.

You can link a wallet to a key-only workspace later, on the Settings page. A wallet belongs to one workspace, and once linked it cannot be changed. The message a wallet signs is valid for 5 minutes. A dashboard session lasts 30 days.

Making and revoking keys

Keys are made and revoked on the Keys page. A workspace can have any number of keys, and all of them have the same access. Revoking takes effect at once and cannot be undone; the key stays in the list, marked as revoked. The page also shows when each key was last used; that time is updated at most once every 60 seconds.

Connect your tools

Claude Code

Set three environment variables in the shell you start Claude Code from:

Terminal
export ANTHROPIC_BASE_URL="https://www.batonapi.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_BATON_KEY"
export ANTHROPIC_API_KEY=""
claude
  • ANTHROPIC_BASE_URL is where Claude Code sends its requests. It adds /v1/messages itself, so the URL has no /v1 at the end.
  • ANTHROPIC_AUTH_TOKEN is your Baton key. Claude Code sends it as a bearer token.
  • ANTHROPIC_API_KEY is left empty on purpose. An Anthropic key that is already in your environment would otherwise be sent in place of the token. The empty value keeps that key out of the request.

Claude Code asks for Anthropic model names. Baton grades and routes a request whatever model it names, so nothing in Claude Code's model settings has to change. Its token counting calls are answered by the token counting endpoint. When a request is answered by a model that is not Anthropic's, it is translated on the way; Formats lists what does not cross.

OpenAI-compatible clients

Any client that lets you set an OpenAI base URL works. It needs three settings:

Base URL
https://www.batonapi.com/v1
API key
YOUR_BATON_KEY
Model
baton
Where OpenAI-compatible clients take the settings
ClientWhere the settings go
OpenAI SDKsbase_url and api_key in Python, baseURL and apiKey in JavaScript, as in the quickstart. The SDKs also read OPENAI_BASE_URL and OPENAI_API_KEY from the environment.
CursorIn the model settings, switch on the OpenAI base URL override, enter the base URL and the key, and add baton as a custom model.
ContinueA model entry with provider: openai, apiBase, apiKey and model: baton.
Aider--openai-api-base, --openai-api-key and --model openai/baton.
LangChainChatOpenAI with base_url, api_key and model="baton".

The client has to use the chat completions API. The Responses API, embeddings and the older completions API are not implemented, and a request to them gets status 404.

Anthropic SDKs

Set base_url (baseURL in JavaScript) to https://www.batonapi.com and api_key (apiKey) to your Baton key, as in the quickstart samples. The SDK adds /v1/messages to the base URL and sends the key in x-api-key. Use baton as the model.

API reference

All endpoints are under https://www.batonapi.com. OpenAI clients take https://www.batonapi.com/v1 as their base URL and Anthropic clients take https://www.batonapi.com.

Endpoints
EndpointWhat it does
POST /v1/chat/completionsRoutes an OpenAI chat completions request to one of your models. The reply is in the chat completions format.
POST /v1/messagesRoutes an Anthropic Messages request to one of your models. The reply is in the Messages format.
POST /v1/messages/count_tokensCounts the input tokens of a Messages request without running it.
GET /v1/modelsLists the model names a key can ask for: the names Baton defines and the models on your ladder.
POST /mcpThe MCP server, for agents that call tools.

These are the only API endpoints. Before a request is routed, Baton checks that the body is a JSON object with a messages array of at least one message, each with a role its format knows. Everything else in the body is checked by the provider that answers.

The model field

The model name in a request never picks a model from your ladder. It only says how to choose the tier:

  • baton has the request graded.
  • baton-hard, baton-medium and baton-easy route it as that tier without grading.
  • Any other name, including the id of a model on your ladder, is treated like baton. That is why clients with fixed model names work unchanged.

Names are compared without regard to case. The response headers tell you which model answered.

Token counting

POST /v1/messages/count_tokens takes the body of a Messages request and answers { "input_tokens": 1234 }. When the first switched-on model of your ladder is an Anthropic model with a working key, the count comes from Anthropic for that model and is exact. Otherwise it is an estimate: one token for every four characters of the system prompt, the messages and the tool definitions. The x-baton-token-count response header says which of the two you got.

Model list

GET /v1/models lists baton, baton-hard, baton-medium and baton-easy, then every switched-on model of your ladder once. The list is in OpenAI's format unless the request has an anthropic-version header, in which case it is in Anthropic's.

Request headers

Request headers the router reads
HeaderWhat it does
authorizationBearer followed by your key.
x-api-keyYour key, when there is no bearer token.
x-baton-tierhard, medium or easy: routes the request as that tier and skips the grader. Any other value is ignored.
x-baton-conversationAn id of your choosing that names the conversation, so that stickiness does not depend on the text of the request.
anthropic-versionPassed on when a Messages request is answered by an Anthropic model; without it, 2023-06-01 is sent. On GET /v1/models its presence selects Anthropic's list format.
anthropic-betaPassed on when a Messages request is answered by an Anthropic model, and dropped otherwise.

No other header of yours reaches a provider. Your Baton key in particular is never passed on: each provider is called with the key you stored for it.

Response headers

Response headers the router sets
HeaderValue
x-baton-modelThe id of the model that answered. none when no model did.
x-baton-provideranthropic, openai, google, openrouter or custom; house for the house model; none when no model answered.
x-baton-rungThe answering model's place on the ladder, from 0 at the top. -1 for the house model and when no model answered.
x-baton-tierhard, medium or easy: the tier the request was routed as.
x-baton-token-countOn the token counting endpoint only: exact or estimate.

The four routing headers are on every response that got as far as routing: an answer, a provider's refusal of the request that is passed back to you, and the 402, 502 and 503 that mean no model could answer. Errors raised before routing, such as a bad key or a body that is not JSON, carry none of them.

Streaming

Set stream to true in the body and the reply is a stream of server-sent events in the format of your request, with the routing headers on the response as usual.

  • When the answering model speaks your format, its events are passed through byte for byte. The one exception: Baton always asks OpenAI-style providers for the usage chunk, and removes it again unless you set stream_options.include_usage yourself.
  • When it speaks the other format, the events are translated one by one. A translated chat completions stream ends with data: [DONE], and its chunks have created set to 0.
  • Another model is only tried before the first byte. If the provider's stream breaks off part of the way through, your stream ends with one error event in your format, and the HTTP status stays 200.
  • If you close the connection, the call to the provider is cancelled. What was generated up to then is still counted.

CORS

The API can be called from a browser on any origin. Every response has access-control-allow-origin: * and exposes the five response headers above to scripts. A preflight OPTIONS request is answered with status 204 and may be cached for 24 hours. It allows these request headers, plus any others the browser announces:

authorization, x-api-key, content-type, anthropic-version, anthropic-beta, x-baton-tier, x-baton-conversation

The API never reads cookies, so a page can only act with a key it already holds. Keep keys out of pages you serve to other people.

Limits

Limits and timeouts
LimitValueWhat happens
Request body25 MBA larger body gets status 413. The MCP endpoint reads at most 4 MB.
Start of the answer60 secondsA provider that has not started answering by then is treated as unavailable and the next model is tried.
Whole answer, not streaming10 minutesCounted from the request to the last byte of the provider's reply. A stream has no such limit.
Run time of a request300 secondsThe longest the inference and MCP routes ask the hosting platform to keep one request running.
Exact token count15 secondsWhen Anthropic takes longer to count, the endpoint answers with an estimate.

How routing works

A request goes through three steps. It is graded hard, medium or easy. The tier and the state of your ladder decide which models to try, in which order. The first one that answers is the one you hear from.

The grader

Grading follows fixed rules and calls no model, so the same request always gets the same grade. What is graded is the last thing a person typed: the latest user message that contains text. Turns that only carry tool results are skipped.

Every request starts at 45 points. Signals add or take away points, and the total, kept between 0 and 100, decides the tier: 65 or more is hard, 35 to 64 is medium, and anything lower is easy.

The grader's signals
GroupPointsSignal
Hard work+30The message asks for a design, a refactor, a migration, debugging, root-cause analysis, a proof, a plan, a security review or a performance investigation.
Hard work+18An architecture question, a comparison of approaches, a move from one technology to another, a failure that is hard to reproduce, concurrency, or work across many files or steps.
Hard work+12A hard subject that is only mentioned: migrations, security, optimisation, distributed systems, production risk, algorithms.
Easy work-30An easy verb is in charge of the message: rename, fix a typo or the wording, format, summarise, translate, convert between simple formats, or write a commit message or a title. Also a greeting or thanks that asks for nothing.
Easy work-20A lookup or a definition ("what is", "who", "how many") of at most 240 characters, with no code attached and no other work asked for.
Easy work-20A bare follow-up such as "ok", "yes" or "go ahead", outside an agent loop.
Easy work-15A yes-or-no question of at most 120 characters, or a request for a list.
Attachments+5, +9, +13The message is at least 400, 1,200 or 4,000 characters long.
Attachments+6It includes code.
Attachments+12, +4It includes a stack trace, or other error output.
Attachments+10, +20It asks for three or four separate things, or for five or more.
Brevity-5, -15The message is at most 59 characters long, or at most 24.
Session+12The request asks for reasoning.
Session+5, +9The whole request holds at least 24,000 characters, or at least 100,000.
Session+8Eight or more tools are offered, and the message is neither easy work nor a bare follow-up.
Session+6The conversation includes an image.
  • Within hard work and within easy work, the strongest signal counts in full and each further one counts half. Hard-work signals about the same thing count once. Hard work adds at most 45 points and easy work takes away at most 40.
  • When an easy verb is in charge and nothing else is asked, the rest of the message is material to work on: hard words in it, its length and its code do not count.
  • A "how do I" question of at most 200 characters is a question about hard work, not the work itself: its hard-work signals count 18 points at most.
  • Brevity is not counted when a 30-point signal fired, or in an agent loop, which is a conversation that already holds tool calls or tool results. There a short message steers work in progress.
  • Session signals add at most 19 points together. They can tip a request that already carries some weight, but they never make a plain one hard.
  • A request asks for reasoning when a Messages request has thinking set to anything but disabled or between_tools, or when output_config.effort or reasoning_effort is medium or higher.
  • With eight or more tools offered, an easy job such as a summary or a translation is not graded easy when its subject is something in the workspace, such as a repository or a pull request, and the material is not in the message. The agent has to go and read it first.
  • The word lists are English. A request in another language is graded on its length, its attachments and the session. Of a very long message, the first 12,000 and the last 4,000 characters are read for wording.

Three requests, each sent on its own with no tools and no reasoning settings:

  • Refactor the auth module to use dependency injection.

    Every request starts at
    45
    asks for a refactor
    +30
    Score and tier
    75, hard

    The message is short, but nothing is taken off for that: a verb worth 30 points outweighs brevity.

  • Explain what a mutex is in one sentence.

    Every request starts at
    45
    concurrency
    +18
    short message
    -5
    Score and tier
    58, medium

    This is the request from the quickstart. It opens with a verb that asks for work, so it is not treated as a lookup.

  • Rename getUser to fetchUser.

    Every request starts at
    45
    asks for a rename
    -30
    short message
    -5
    Score and tier
    10, easy

    An easy verb is in charge of the whole message. Reasoning settings, tools and a large context could add at most 19 points, which still leaves it easy.

Only the tier is reported, in the x-baton-tier response header. The grader on the home page runs the same rules on whatever you type.

The ladder policy

The ladder is your models in order, with rung 0 at the top. A rung is usable when it is switched on, its provider has a working key, it is not cooling down, and its spend today is below its daily budget. A rung with no budget is never over it.

Where each tier starts and how long it stays
TierStarts atStays on a rung while
hardRung 0Less than 100% of the rung's daily budget is spent.
mediumRung 0Less than 50% is spent. On the last rung: less than 100%.
easyThe last rungLess than 100% is spent.
  1. If the conversation is pinned to a rung and that rung is usable, it is the first choice, whatever the tier and the percentages say.
  2. Otherwise the router walks down from the rung the tier starts at and takes the first usable rung that is under the tier's percentage.
  3. If no rung is under it, the router walks down again and takes the first usable rung. The 50% mark is a preference; a spent budget is a limit.
  4. The fallbacks are every usable rung below the first choice, in order, and after them the house model. Each is tried only when the one before it could not answer.

Some things follow from these rules that are easy to miss:

  • Apart from a pin, nothing moves a request up the ladder past the rung its tier starts at. An easy request without a pin therefore uses the last rung, then the house model, and never a rung above, even when those are free.
  • The last rung is the last row of the ladder, whether or not it is switched on. When that row is switched off, has no working key, is cooling down or is over budget, easy requests go straight to the house model.
  • With a single rung, all three tiers use it until its budget is spent.
  • Budgets are checked before a request is sent, so the requests already in flight can take a rung past its budget.

An example. Rung 0 is claude-opus-4-8 with a budget of $10.00 a day, rung 1 is claude-sonnet-4-6 with $5.00 a day and less than half of it spent, and rung 2 is claude-haiku-4-5 with no budget. This is where a request without a pin goes first:

First choice per tier on the example ladder
Spent today on rung 0hardmediumeasy
$0.00 to $4.99Rung 0Rung 0Rung 2
$5.00 to $9.99Rung 0Rung 1Rung 2
$10.00 or moreRung 1Rung 1Rung 2

Budgets and the daily reset

A rung's spend is the cost of the requests it answered today: input tokens times the input price plus output tokens times the output price, at the prices you set on the rung in US dollars per million tokens. Token counts are the ones the provider reports. When a provider reports none, they are estimated at four characters per token.

  • The day is the UTC day. Every rung's spend starts again from zero at 00:00 UTC.
  • Cached input tokens are counted at the full input price.
  • Only answers count. A request the provider refused adds nothing, and a stream that was cut short adds what had been generated.
  • A rung whose prices are both 0 never accumulates spend, so its budget is never reached. Set prices on every rung that has a budget.

This spend is an estimate that exists for routing. What you owe a provider is on that provider's invoice.

Rate limits, cooldowns and provider errors

What a provider answers decides whether the next candidate is tried. A cooldown belongs to one rung, and every request skips that rung until the cooldown has passed.

What the router does with a provider's answer
The provider answersWhat happens
429, 529 or 402The model is out for now: a rate limit, an overload, or a provider account without credit. The rung cools down and the next candidate is tried. The cooldown is the time the provider asks for in retry-after-ms or retry-after, kept between 5 seconds and 60 minutes, or 60 seconds when it names none.
401 or 403The provider key is marked as rejected, whatever the provider's reason, and the next candidate is tried. Every rung of that provider is then skipped until you save the key again or test it with success on the Ladder page.
500 and above, 408, a redirect, no connection, or no start of an answer in timeThe rung cools down for 30 seconds and the next candidate is tried.
200 with a body that is not an answerThe same as an outage. This covers a body that is not JSON, JSON that holds only an error, and a reply to a streaming request that is not an event stream.
Any other 4xx, such as 400, 404 or 422The request itself is at fault, so no other model is tried. The provider's error is passed back with the same status, in the format of your request.

The house model is the last candidate, so there is nothing to fall back to: when it fails in any of these ways, the request ends with status 502. When no candidate was left at all, it ends with 402 or 503.

Conversation stickiness

Prompt caches and thinking blocks belong to one model, so by default a conversation stays on the rung that answered it last. The API keeps no conversations, so one is recognised by a fingerprint: a hash of your workspace, the system prompt and the text of the first user message, which are the same on every turn. Only the hash is stored.

  • After each answer from a rung, the conversation is pinned to that rung for 6 hours, counted from its latest answer.
  • The next turn goes to the pinned rung first, as long as it is usable. When it is not usable, the turn is routed as if there were no pin. When it fails, the rungs below it are tried. Either way the pin moves to the rung that answers.
  • The pin does not look at the tier. A conversation that opens with an easy message is pinned to the last rung, and a later hard message in it is answered there too.
  • Answers from the house model set no pin.
  • In a chat completions request, the system prompt is the system and developer messages that come before the first other message. When the first user message has no text and no id is sent, there is nothing to key on and the request is routed without a pin.

To name a conversation yourself, send any id in the x-baton-conversation header. The fingerprint is then made from your workspace and that id, and the text of the request no longer matters. Use it when your client rewrites the system prompt or the first message between turns, when separate conversations open with the same text, or to start a conversation afresh under a new id.

Terminal
curl https://www.batonapi.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_BATON_KEY" \
  -H "Content-Type: application/json" \
  -H "x-baton-conversation: ticket-4821" \
  -d '{
    "model": "baton",
    "messages": [{"role": "user", "content": "Why does the nightly export time out?"}]
  }'

To turn stickiness off for the whole workspace, switch off "Keep conversations on one model" on the Settings page. Pins are then neither read nor written.

Forcing a tier

There are two ways to skip the grader. Send hard, medium or easy in the x-baton-tier header, or ask for the model baton-hard, baton-medium or baton-easy. When a request has both, the header wins.

Terminal
curl https://www.batonapi.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_BATON_KEY" \
  -H "Content-Type: application/json" \
  -H "x-baton-tier: hard" \
  -d '{
    "model": "baton",
    "messages": [{"role": "user", "content": "Is this lock ordering safe?"}]
  }'

A forced tier changes where the request starts on the ladder and nothing else. Budgets and cooldowns apply as usual, and a conversation pin still comes first. Baton reports the tier it used in the response header either way.

Formats

You can send either format, whatever is on your ladder. Anthropic models are called in the Messages format. OpenAI, Google, OpenRouter, custom endpoints and the house model are called in the chat completions format. When your request and the answering model differ in format, the request is translated on the way in and the reply on the way back, streams included.

A chat completions request answered by an Anthropic model

  • System and developer messages, wherever they stand in the conversation, are joined into the system field.
  • Consecutive messages of one role become one turn. tool messages become tool_result blocks at the start of a user turn, and an assistant's tool_calls become tool_use blocks.
  • Function tools become Anthropic tools with the same name, description and schema. In tool_choice, required becomes any and a named function becomes that tool. parallel_tool_calls: false becomes disable_parallel_tool_use.
  • Tool call ids are reduced to letters, digits, _ and -, the characters Anthropic accepts, and made unique within the request.
  • An image_url part becomes an image block: base64 for a data: URL, a URL source for any other. Audio and file parts are left out.
  • max_completion_tokens, or else max_tokens, becomes max_tokens. Anthropic requires it, so a request with neither gets 16,000, or 64,000 when it streams.
  • stop becomes stop_sequences, and temperature is kept within 0 to 1.
  • reasoning_effort becomes output_config.effort, with minimal and none sent as low. A response_format of type json_schema becomes output_config.format. user becomes metadata.user_id.
  • Left out, because the Messages format has no place for them: n, logprobs, seed, logit_bias, the two penalties, response_format of type json_object, a tool's strict flag, and the older functions parameter and function messages.

The translated request is then fitted to the model like any other Messages request, as described under models that share a format.

A Messages request answered by an OpenAI-style model

  • system becomes a system message at the start.
  • In a user turn, each tool_result block becomes a tool message, followed by one user message with the rest. A failed result has Error: put before its text. Images inside tool results move to the user message, because a tool message holds only text.
  • Image blocks become image_url parts; base64 data is sent as a data: URL.
  • In an assistant turn, the text blocks are joined and the tool_use blocks become tool_calls.
  • Your own tools become function tools. In tool_choice, any becomes required and a named tool becomes that function. disable_parallel_tool_use becomes parallel_tool_calls: false.
  • max_tokens is sent to OpenAI as max_completion_tokens and to other providers as max_tokens. stop_sequences becomes stop; OpenAI gets the first 4, the most it accepts.
  • temperature and top_p pass through. top_k is left out.
  • output_config.effort becomes reasoning_effort for OpenAI models that accept it, with xhigh and max sent as high. Other providers do not get it. An output_config.format of type json_schema becomes a response_format.
  • Left out: the thinking setting, cache_control markers, document blocks and metadata.

Replies

Text stays text, and tool calls become tool calls in the other format, with their arguments turned from a JSON string into an object or back. Stop reasons are mapped like this when an Anthropic model answers a chat completions request:

Stop reasons in the two formats
Messages stop_reasonChat completions finish_reason
end_turn, stop_sequence, pause_turnstop
max_tokens, model_context_window_exceededlength
tool_usetool_calls
refusalcontent_filter

In the other direction, length becomes max_tokens and content_filter becomes refusal. Any other reply that calls tools is tool_use, whatever reason the provider gave, and everything else is end_turn.

  • Usage keeps its meaning. Anthropic counts cached input apart from input_tokens; in a chat completions reply prompt_tokens is the sum of all input, with the cached part in prompt_tokens_details.cached_tokens. In a Messages reply, cached tokens an OpenAI-style provider reports are taken out of input_tokens and put in cache_read_input_tokens.
  • Ids are kept as the provider gave them, so a chat completions reply can carry an id that starts with msg_. The model field of a translated reply is the id of the model that answered.
  • A refusal from an OpenAI-style model arrives as text. Reasoning text is not part of a translated reply, in either direction.
  • A provider's error keeps its message and gets the error shape of your format.

Between models that share a format

When your request and the answering model share a format, the body is sent as you wrote it, with model replaced. Messages, the system prompt and tools are never touched. Only parameters the target model would reject are changed.

For a Messages request sent to an Anthropic model:

  • temperature, top_p and top_k are removed for models that refuse sampling parameters. For models that take temperature or top_p but not both, top_p is removed when both are set.
  • max_tokens is lowered to the model's limit, where it is known.
  • output_config.effort is lowered to the nearest level the model accepts, or removed.
  • thinking is fitted: a token budget becomes adaptive thinking on models that have no budgets, and adaptive thinking becomes a budget of half of max_tokens, between 1,024 and 32,000 tokens, on models that only have budgets. disabled is removed for models that cannot switch thinking off.
  • A tool_choice that forces a tool becomes auto on models that do not accept forcing.
  • A model id that is not recognised as a Claude model gets the strictest of these rules.

For a chat completions request sent to an OpenAI-style model:

  • A stream always asks the provider for usage with stream_options.include_usage. The extra chunk is taken out of your stream again unless you asked for it.
  • For OpenAI itself, max_tokens is renamed to max_completion_tokens, and reasoning_effort is removed for models whose names start with gpt-3.5, gpt-4 or chatgpt-4, which reject it.

What cannot cross

  • Anthropic server tools. Tools that Anthropic runs, such as web search and code execution, exist only there. When a Messages request goes to an OpenAI-style model they are left out, together with a tool_choice that forces one and any of their calls and results in the history.
  • Thinking blocks. Only the model that wrote them can use them. They are left out when a Messages request goes to an OpenAI-style model, and thinking in a reply is not translated. Between two Anthropic models they are forwarded as they are.
  • Prompt caches. A cache belongs to one model at one provider. A chat completions request sent to an Anthropic model sets no cache markers, so nothing is cached there.
  • Beta features. The anthropic-beta header reaches Anthropic models only.

This is why conversations are sticky by default: every move to another model gives up whatever of the above the conversation was using.

The house model

The house model is a model this deployment provides to every workspace. It is not on your ladder and needs no provider key of yours. Which model it is and what it costs is up to the operator; the Ladder page shows both under your last rung.

When it answers

The house model is the last candidate of every request. It is asked when:

  • your ladder is empty,
  • no rung is usable for the request, or
  • every rung that was tried failed.

It then answers only if all of this holds:

  • The operator has configured a house model.
  • "Fall back to the house model" is on for your workspace. It is on by default; the switch is on the Settings page.
  • Your workspace has credit, unless the operator has set the house model's price to zero.

An answer from the house model has house in x-baton-provider and -1 in x-baton-rung. It sets no conversation pin and counts toward no rung's budget.

How it is priced

Each answer is charged to your credit: input tokens times the input price plus output tokens times the output price, both per million tokens. The counts are the ones the house model reports, or an estimate at four characters per token when it reports none. The charge is made after the answer, and the Requests page shows it for each request.

At zero credit

With no credit, the house model is not asked and the request ends with status 402 and the code insufficient_credit. The check looks at the balance your workspace had when the request arrived, not at the size of the request: a request that starts with any credit is answered, and its charge takes the balance to zero at most.

When the house model is not configured or is switched off, the status is 503 instead. When it was asked and failed, the status is 502. The house model is not retried.

For operators

The house model is any endpoint that speaks the chat completions format. Set HOUSE_BASE_URL to the part of its address before /chat/completions. Requests are sent there with model set to HOUSE_MODEL and, when HOUSE_API_KEY is set, that key as a bearer token. The two prices are HOUSE_PRICE_IN_USD_PER_MTOK and HOUSE_PRICE_OUT_USD_PER_MTOK. All of them are in the table of variables.

  • Without HOUSE_BASE_URL there is no house model. In development a built-in mock stands in for it: it calls nothing, costs nothing, and answers with a fixed text that names the tier of the request. HOUSE_MOCK switches it on or off.
  • With both prices at 0 the house model is free, and it answers workspaces that have no credit.
  • REQUIRE_CREDIT makes credit a condition for the whole API. A workspace without credit then gets 402 for every request, including those its own keys could have answered.

Credits

Credit is a prepaid balance in US dollars that belongs to a workspace. It pays for answers from the house model and for nothing else. Requests that your own provider keys answer are billed by those providers and leave your credit alone.

Adding credit

On this deployment, credit is granted by the operator. There is nothing to buy in the app: the operator adds credit to a workspace by its id, which is on the Settings page, or by the wallet linked to it. A new workspace starts with credit when the operator has set a starter amount.

Where to see it

  • The Credits page shows the balance, what was used today (UTC) and the additions so far.
  • The Requests page shows the credit charged for each request.
  • An agent can read the balance with the balance tool of the MCP server.

For operators

POST /api/credits/grant adds credit to a workspace. It exists only when ADMIN_SECRET is set, and takes that secret as a bearer token. The body names the workspace by workspaceId or by wallet, one of the two, with usd above zero and an optional note of up to 200 characters.

Terminal
curl https://www.batonapi.com/api/credits/grant \
  -H "Authorization: Bearer $ADMIN_SECRET" \
  -H "Content-Type: application/json" \
  -d '{"workspaceId": "ws_3fKx8mQ2LpZr7TnVbY4c", "usd": 5, "note": "Trial credit"}'
Reply
{
  "ok": true,
  "workspaceId": "ws_3fKx8mQ2LpZr7TnVbY4c",
  "grantedMicros": 5000000,
  "balanceMicros": 5000000
}

Amounts in the reply are in millionths of a dollar. A wrong secret gets status 401, and a workspace that does not exist gets 404; a wallet counts only once its owner has signed in with it. STARTER_CREDIT_USD gives every new workspace a first amount without a call.

Your own models

Baton routes between models you bring. On the Ladder page you store a key for each provider you use, then put that provider's models on the ladder in the order they should be tried.

Providers

Providers
ProviderCalled inBase URL
AnthropicMessageshttps://api.anthropic.com
OpenAIChat completionshttps://api.openai.com/v1
GoogleChat completionshttps://generativelanguage.googleapis.com/v1beta/openai
OpenRouterChat completionshttps://openrouter.ai/api/v1
Custom endpointChat completionsThe base URL you enter
  • A workspace holds one key per provider. Saving another key for the same provider replaces the first.
  • A key is encrypted with AES-256-GCM before it is stored. After that it is shown only in masked form and used only to call that provider for you.
  • A key is tried against the provider as soon as you save it, and again whenever you press Test. A key the provider rejects is marked as such, and the models that depend on it are skipped.
  • Removing a key leaves its models on the ladder. They are skipped until the provider has a key again.

Custom endpoints

The custom provider is any endpoint that speaks the chat completions format: vLLM, Ollama, LM Studio, a hosted gateway, your own server. A workspace has one. Enter the part of its address before /chat/completions; the model list is read from /models under the same address, and the key is sent as a bearer token. The key field cannot be empty, so for an endpoint that needs no key, enter any word.

The base URL has to pass these rules:

  • It is an http or https URL with no user name or password in it. A query string, a fragment and trailing slashes are removed.
  • Unless the operator allows private addresses, which is the default only in development, it has to use https and name a public host. Refused are localhost, private, link-local and carrier-grade NAT addresses, names without a dot, and names that end in .localhost, .local or .internal.
  • Under the same condition, the host name is looked up again before every request, and the request is not sent when any of its addresses is private.
  • Redirects are never followed, for any provider. A redirect counts as the endpoint being unavailable.

Models, prices and budgets

A ladder holds up to 8 models. The picker lists the models your key can call, read live from the provider, and you can type a model id instead. A new model goes to the bottom; move it to where it belongs. Each model on the ladder has:

  • Prices for input and output, in US dollars per million tokens. When you pick a model they are filled in where a list price is known: from a built-in table for current Claude models, and from OpenRouter's public price list for OpenAI, Google and OpenRouter models. Check them. Prices are what a rung's spend is worked out from, and a rung priced at 0 never reaches its budget.
  • A daily budget in US dollars, or none. Without a budget a rung is left only when its provider rate-limits or fails. A budget of $0 is refused: switch the model off instead.
  • A switch. A model that is switched off keeps its place and is skipped. Keep in mind that easy requests start at the last row, switched on or not.
  • A label, if you want one. It is shown in the dashboard in place of the model id.

Changes apply from the next request. The page also shows what each model has spent today and whether it is ready, cooling down, over budget or without a working key.

The Savings page shows what the ladder saved over the last 7 or 30 days. It prices the tokens of every answered request at the current prices of your top model (the first one that is switched on) and sets that against what was paid: the cost at the prices of the model that answered, plus credit for the house model. It is an estimate, since another model would not have used exactly the same number of tokens, and it needs prices on the top model.

Alerts

On the Settings page you can give a webhook address, and Baton posts a message to it when a model passes half of its daily budget, when a model's daily budget is spent, and when a provider rate-limits a model. Each budget message is sent once a day per model, and a rate-limited model is announced at most once every 15 minutes. You choose which of the three you want, and can send a test message.

A Discord webhook gets the message as content and a Slack one as text. Any other address gets a JSON body with text, event (budget_half, budget_spent or rate_limited), workspaceId, model and the amounts or the time that go with the event. The address must be public and use https, it is stored encrypted, and a webhook that redirects is not followed.

MCP server

Baton is also an MCP server, so an agent can ask a routed question and read your ladder, usage and balance as tools.

URL
https://www.batonapi.com/mcp
Transport
Streamable HTTP
Authentication
Authorization: Bearer YOUR_BATON_KEY

The key is an ordinary API key from the Keys page. A client configuration looks like this; it is the form Claude Code reads from .mcp.json, and other clients take the same three facts under their own names.

.mcp.json
{
  "mcpServers": {
    "baton": {
      "type": "http",
      "url": "https://www.batonapi.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_BATON_KEY"
      }
    }
  }
}

On the wire

  • The server is stateless. Every message is a POST that carries one JSON-RPC message and is answered as JSON. There are no sessions and no event streams; GET and DELETE get status 405.
  • The key is checked on every message, initialize included. Without a valid one the answer is status 401 with a WWW-Authenticate: Bearer header. There is no OAuth flow: set the key as a header in the client. x-api-key is read as well.
  • A client that speaks HTTP itself has to send Content-Type: application/json and Accept: application/json, text/event-stream.
  • Batches are refused with status 400, and a body over 4 MB with status 413.

Tools

The tools of the MCP server
ToolInputWhat it returns
askprompt (text, required), system (text), tier (hard, medium or easy), max_tokens (a whole number of 1 or more)Sends one prompt through the router and returns the reply, followed by a line that names the model, provider, rung and tier that answered. As data: text, model, provider, rung, tier and finish_reason.
ladderNoneYour models in ladder order. For each: its state (ok, disabled, no_key, cooling_down or out_of_budget), what it has spent today and its daily budget. Then the house model, and whether it is switched on for the workspace.
usageNoneToday's traffic, in UTC: requests, tokens in and out, the estimated cost on your own provider keys and the credit used by the house model, by tier and by answering model.
balanceNoneYour credit balance in US dollars.
  • Every result is readable text, with the same facts as structured content.
  • ask is an ordinary routed request. It is graded unless tier is set, it counts toward budgets and credit, and it appears on the Requests page under the key that made it. It does not stream: the tool returns when the whole reply is there.
  • Each ask stands alone, with no memory of earlier calls. Stickiness still applies: the same prompt with the same system prompt goes to the rung that answered it before.
  • When the router cannot answer, the tool result is marked as an error and carries the message of the error an API call would have got.

Errors

An error comes in the shape of the format you sent the request in, with an HTTP status that says what kind of error it is.

The two shapes

{
  "error": {
    "message": "No API key was sent. Create a key in the Baton dashboard at https://www.batonapi.com and send it as \"Authorization: Bearer <key>\" or in the x-api-key header.",
    "type": "authentication_error",
    "param": null,
    "code": "invalid_api_key"
  }
}

The chat completions shape carries the router's code in error.code. The Messages shape has no field for it, so there you tell errors apart by the status and by error.type.

Error types per status
StatusChat completions error.typeMessages error.type
400invalid_request_errorinvalid_request_error
401authentication_errorauthentication_error
402insufficient_quotapermission_error
413invalid_request_errorrequest_too_large
429rate_limit_errorrate_limit_error
500server_errorapi_error
502server_errorapi_error
503server_errorapi_error

The router's own errors

The router's own errors
StatusCodeWhenWhat to do
400invalid_jsonThe body is not valid JSON.Send a JSON body.
400invalid_requestThe body is not a JSON object, has no messages array or an empty one, or holds a message without a role its format knows. The message says which.Correct the request. Sending it again unchanged gives the same error.
401invalid_api_keyNo key was sent, or the key is unknown or revoked.Send a key from the Keys page in one of the two key headers.
401workspace_not_foundThe key is valid but its workspace no longer exists.Use a key of a workspace that exists.
402insufficient_creditOnly the house model was left to answer and the workspace has no credit. Also every request of a workspace without credit, where the operator requires credit.Add credit, or wait until a model on your ladder is usable again. The message names what stopped each one.
429rate_limitedThe workspace sent more requests in one clock minute than the router allows: 120 unless the operator set another number. This is the router's own limit; a provider's rate limit never reaches you as an error, it moves the request down the ladder.Wait the number of seconds in the Retry-After header, or send requests more slowly.
413request_too_largeThe body is larger than 25 MB.Send less: a shorter history, or fewer and smaller images.
500internal_errorSomething failed inside the router.Try again in a moment. If it keeps happening, tell the operator.
502upstream_errorThe house model was asked and failed. Also the last event of a stream that the provider broke off.Send the request again.
503no_model_availableNo model on the ladder could answer, and the house model is not configured or is switched off. The message names what stopped each model.Act on the message: wait out a cooldown or the daily reset, replace a rejected provider key, add a model, or switch the house model on.

Keys are made on the Keys page, credit is on the Credits page, and provider keys, models and the state of each rung are on the Ladder page.

Errors from a provider

Most provider failures never reach you: the router tries the next model. The exception is a provider that refuses the request itself, with a 4xx status other than 401, 402, 403, 408 and 429. That status and the provider's message are passed back to you.

  • In the chat completions shape, the provider's own type, code and param are kept where it sent them.
  • In the Messages shape, error.type is Anthropic's own when Anthropic sent the error, and otherwise the one that goes with the status.
  • When the provider's body holds no message that can be read, the message is "The model provider returned an error" with the status.

Codes in the request log

The Requests page shows a code under the status of every request that did not end well. Errors raised before routing, such as a bad key or a body that is not JSON, are not logged.

Error codes in the request log
CodeMeaning
provider_errorA provider refused the request itself. The status is the provider's.
stream_interruptedThe provider's stream broke off part of the way through. The status is 200.
client_closedYou closed the connection. The status is 200 when a stream was under way, and 499 when nothing had answered yet.
upstream_errorThe house model failed: status 502.
no_model_availableNothing could answer: status 503.
insufficient_creditOnly the house model was left and there was no credit: status 402.

Privacy

The request log

One row is written for every request that is routed. It holds:

  • an id for the row
  • the workspace
  • which of your API keys made the request
  • the time
  • the format of the request, chat completions or Messages
  • the tier it was routed as
  • the rung that answered
  • that rung's place on the ladder at the time
  • the provider
  • the model
  • the number of input tokens
  • the number of output tokens
  • the estimated cost at the prices you set
  • the credit charged
  • how long the request took
  • the HTTP status of the answer
  • whether it streamed
  • an error code, when there was one
  • the conversation fingerprint, which is a hash

That is all of it. Rows are kept until the operator deletes them; nothing removes them automatically.

What is never stored

  • Prompts, system prompts and replies.
  • Tool definitions, tool calls and tool results.
  • Images and other attachments.
  • The headers of your requests.
  • Your API keys. Only a hash of each is kept, with a masked form for display.

The grader and the conversation fingerprint read the text of a request in memory while it is being routed. The fingerprint that is stored is a hash, and the text cannot be recovered from it.

What else is stored

  • The workspace: its name, the address of a linked wallet, its credit balance and its routing settings.
  • API keys: the name you gave each, its hash and masked form, and when it was made, last used and revoked.
  • Provider keys, encrypted with AES-256-GCM, each with a masked form and, for a custom endpoint, its base URL.
  • The ladder, and for each rung and UTC day its spend, token counts and number of requests.
  • Conversation pins: a fingerprint and the rung it is pinned to. A pin expires after 6 hours and is deleted when the workspace next starts a new conversation.
  • Additions of credit: the amount, where it came from and a note.

A dashboard session is a signed cookie; no list of sessions is kept on the server. Signing in with a wallet signs a message in the wallet, which never hands over a key.

Where a request goes

A request is sent to the provider whose model answers it, with the key you stored for that provider, or to the house model's endpoint. What happens to it there is up to that provider. Baton sends the body of the request and, to Anthropic, the anthropic-version and anthropic-beta headers. Your Baton key, your cookies and your other headers are not sent on. Requests to OpenRouter also carry the name and address of this site, which is how OpenRouter attributes traffic to an app.

Self-hosting

These notes are for the operator of a deployment. Baton is a Next.js application that keeps its data in one Postgres database.

  • GET /api/health answers { "ok": true } when the database can be reached, and status 503 when it cannot.
  • Inference and MCP requests ask the platform for up to 300 seconds of run time. A model that writes for longer than the platform allows is cut off.
  • The API reads request bodies of up to 25 MB. Vercel limits a function's request body to 4.5 MB, so there the lower limit applies.
  • A truthy flag is true or 1. A number that cannot be read falls back to its default.
  • The product's name, the key prefix, the header names and the model names all come from one file, src/lib/brand.ts. A rename there changes them everywhere, these docs included.

Environment variables

"Production" below means NODE_ENV is production. How the house model variables work together is described under the house model.

Server environment variables
VariableWhat it doesWhen it is not set
DATABASE_URLThe Postgres connection string. The tables are created on the first connection.An embedded database in .data/pglite, for development. On Vercel, every request that needs the database fails without it.
SESSION_SECRETSigns session cookies and wallet sign-in messages. Use at least 32 characters. Changing it signs everyone out.Required in production. A fixed development value otherwise.
ENCRYPTION_KEYEncrypts stored provider keys. Any string; it is hashed to a 256-bit key. After a change, the keys stored before it cannot be read and have to be saved again.Required in production. A fixed development value otherwise.
ADMIN_SECRETThe bearer secret of the endpoint that grants credit.The endpoint answers 404.
STARTER_CREDIT_USDCredit, in US dollars, given to every new workspace.0 in production, 1 in development.
REQUIRE_CREDITWhen true, the API answers 402 to a workspace without credit, even for requests its own provider keys could answer.false
ALLOW_PRIVATE_UPSTREAMSWhen true, custom endpoints may use http and may point at localhost and private networks.false in production, true in development.
HOUSE_MOCKWhen true and HOUSE_BASE_URL is not set, a built-in mock answers as the house model.false in production, true in development.
HOUSE_BASE_URLThe house model's endpoint, in the chat completions format, without /chat/completions.No house model.
HOUSE_API_KEYSent to the house model's endpoint as a bearer token.No key is sent.
HOUSE_MODELThe model id sent to that endpoint and reported in the model response header.house
HOUSE_LABELThe house model's name in the dashboard.The value of HOUSE_MODEL, or "House model".
HOUSE_PRICE_IN_USD_PER_MTOKWhat the house model charges to credit per million input tokens, in US dollars.0.2
HOUSE_PRICE_OUT_USD_PER_MTOKThe same per million output tokens.0.8
RATE_LIMIT_REQUESTS_PER_MINUTEAPI requests one workspace may send in a clock minute before it gets a 429. 0 switches the limit off.120
RATE_LIMIT_SIGNUPS_PER_HOURWorkspaces one address may create in an hour by signing in without a wallet. 0 switches the limit off.5
RATE_LIMIT_WALLET_CHALLENGES_PER_HOURWallet sign-in messages one address may ask for in an hour. 0 switches the limit off.30

Other variables

Other environment variables
VariableWhat it doesWhen it is not set
NEXT_PUBLIC_SITE_URLThe public address of the deployment. Every URL in these docs and in the dashboard's samples, and the site named in the wallet sign-in message, comes from it. Read when the app is built.http://localhost:3000
NEXT_PUBLIC_X_URLA link to an X account, shown in the site header. Read when the app is built.No link.
NEXT_PUBLIC_GITHUB_URLA link to the source repository, shown in the site footer. Read when the app is built.No link.
PROVIDER_BASE_URL_<ID>Replaces a provider's base URL for the whole deployment, for example to send its traffic through a gateway. <ID> is one of ANTHROPIC, OPENAI, GOOGLE, OPENROUTER, CUSTOM.The provider's own address.
PGLITE_DIRWhere the embedded development database keeps its files. Used only without DATABASE_URL..data/pglite