Keep your best model for the hard parts.
One API key for every model. Baton keeps your best model on the hard requests, sends easy ones to your cheapest, moves the rest down as budgets run down, and takes over itself when all are spent.
https://www.batonapi.com/v1A day on a three-model ladder, run by the same policy as the API.
An animation of one day on a ladder of three models, routed by the same policy as the API. In the morning, hard and medium requests go to the top model and easy ones to the small model. When the top model has spent half of its daily budget, medium requests start one step down, on the mid-size model, and when that one has spent half of its own, on the small model. When the top model's budget is spent, hard requests step down too. Late in the day every budget is spent and the house model answers what is left, paid from prepaid credit. At midnight UTC the budgets reset and the day starts over.
Simulated time: 00:00 UTC
Morning. Hard and medium requests go to the top model, easy ones to the small model.
Top model
0% of budget spent
Mid-size model
0% of budget spent
Small model
0% of budget spent
House model
Paid from credit
Requests today
- Hard0
- Medium0
- Easy0
What moves, and when
Baton sends every request down your ladder of models. Where it starts depends on how hard it is, and how long it stays depends on how much of that model's daily budget is left.
Hard
- Starts
- At the top of the ladder.
- Leaves
- When that model's daily budget is spent.
- Stays: under half
- Stays: past half
- Leaves: spent
Medium
- Starts
- At the top of the ladder.
- Leaves
- When half of that model's daily budget is gone.
- Stays: under half
- Leaves: past half
- Leaves: spent
Easy
- Starts
- At the bottom of the ladder.
- Leaves
- When that model's daily budget is spent, for the house model.
- Stays: under half
- Stays: past half
- Leaves: spent
Each track is the daily budget of the model a request is on, and a lit cell means requests of that tier stay. The half mark is a preference, not a limit: the last model is used until it is spent, and after that a medium request takes the first model that still has budget.
A rate limit moves work down
A rate-limited model is skipped, and the request goes to the next one down. The model stays out for as long as its provider asks (60 seconds when it does not say, at least 5 seconds, at most 60 minutes); one whose provider is down or cannot be reached stays out for 30 seconds.
Nothing climbs past its start
A request only falls: if the model it is sent to is rate limited or unavailable, the next one down takes it, and after the last one the house model does. An easy request starts at the bottom, so it reaches your best model only in a conversation that is pinned to it.
A conversation stays on one model
A conversation keeps the model that answered it last, however its later messages are graded, until that model runs out of budget or is rate limited. The pin lapses 6 hours after the conversation's last answer, and pinning can be switched off in Settings.
Budgets reset at midnight UTC
Each model can have a daily budget in dollars. Spend is counted per UTC day, from token usage and the prices you set on the model, so every budget is whole again at midnight UTC.
Graded by rules, not by a model
Before a request goes out, Baton reads what was asked: how long it is, the words that signal hard or easy work, whether tools are attached, whether you asked for reasoning, and how much context came along. No second model reads your prompt.
Type a request to see where it would start on a ladder of three models. This is the same rule set the API runs.
- Grade
- --
- Starts
- --
- Why
- Nothing to grade yet. Type a request or pick an example.
The grade can be overridden per request with the x-baton-tier header set to hard, medium or easy.
Something always answers
When every model on your ladder is out of budget or rate limited, the request goes to the house model instead of failing. It needs no provider key.
- x-baton-model:
- <model id>
- The model that answered.
- x-baton-provider:
- house
- Who served it.
- x-baton-rung:
- -1
- Its place on your ladder. 0 is the top; -1 is the house model, which is not on it.
- x-baton-tier:
- hard
- How the request was graded.
The house model is paid from prepaid credit. Requests that go to your own provider keys never touch it.
Change one URL
Create a key in the app, then point the base URL of Claude Code, the OpenAI or Anthropic SDK, or any client that takes one at Baton.
Create a key
In the app. It is shown once, so copy it when it appears.
Point your client at it
Give it the base URL and the key, as in the samples below.
Add your own models
Optional. Add a provider key and put models on your ladder; without them the house model answers.
export ANTHROPIC_BASE_URL="https://www.batonapi.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_BATON_KEY"
export ANTHROPIC_API_KEY=""
claude# pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://www.batonapi.com/v1",
api_key="YOUR_BATON_KEY",
)
completion = client.chat.completions.create(
model="baton",
messages=[{"role": "user", "content": "Explain what a mutex is in one sentence."}],
)
print(completion.choices[0].message.content)// npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://www.batonapi.com/v1",
apiKey: "YOUR_BATON_KEY",
});
const completion = await client.chat.completions.create({
model: "baton",
messages: [{ role: "user", content: "Explain what a mutex is in one sentence." }],
});
console.log(completion.choices[0].message.content);# pip install anthropic
import anthropic
client = anthropic.Anthropic(
base_url="https://www.batonapi.com",
api_key="YOUR_BATON_KEY",
)
message = client.messages.create(
model="baton",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain what a mutex is in one sentence."}],
)
for block in message.content:
if block.type == "text":
print(block.text)curl https://www.batonapi.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_BATON_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "baton",
"messages": [{"role": "user", "content": "Explain what a mutex is in one sentence."}]
}'Questions
- Which models does it work with?
- Anything you can reach with an Anthropic, OpenAI, Google or OpenRouter key, plus any OpenAI-compatible endpoint you run yourself. You put the models in order and Baton works down the list.
- How does it decide how hard a request is?
- By rules, not by another model reading your prompt: the length of what you asked, words that signal hard or easy work, whether tools are attached, whether you asked for reasoning, and how much context came along. You can override the grade per request with a header.
- What counts against a budget?
- Each model on your ladder can have a daily budget in dollars. Spend is counted from the token usage the provider reports and the prices set on that model, and resets at midnight UTC. A rate-limit error also moves work down until the provider's retry time has passed.
- Does a conversation stay on one model?
- Yes, by default. Prompt caches and thinking blocks belong to one model, so a conversation stays where it is until its model runs out of budget or is rate limited. You can turn this off in settings.
- What happens to my provider keys?
- They are encrypted before they are stored and used only to call that provider for you. The request log keeps the model, token counts, cost and timing of each request, never the prompt or the reply.
- What does it cost?
- Requests that go to your own provider keys are billed by the provider, as before. The house model is paid from prepaid credit.
- Is a router always the right call?
- No. If one model at a lower reasoning effort gets you through the day, that is simpler and keeps a single prompt cache. Baton is for when it still runs out.