Model proxy and budgets
SpanDock can sit between your AI tools and the model providers. With the model proxy on, the server decides which model answers each call and records what it cost. You can swap models with rules, let Jev pick a cheaper model when it is good enough, and keep each person within a budget. Everything is configured on the server under Model routing.
What the proxy does
Under Model routing, the server can:
- Swap models with rules. For example, send Claude Code's Opus requests to Sonnet.
- Let Jev pick a model for each new conversation. Jev is SpanDock's model advisor: it looks at the start of the conversation, the person's budget and the candidate models, and chooses a cheaper model when it is good enough.
- Keep each person within a budget.
For each call, the server decides in this order:
- The budget. Over the limit, the call is handled the way you chose (see below).
- The conversation's earlier choice. A conversation keeps its model for 6 hours, so the model doesn't change mid-task.
- The first matching rule.
- Jev, if it is on.
- The requested model, unchanged.
Budgets
Budgets are per person. You assign clients to people; a client that isn't assigned counts as its own person. Each budget has a period (a day, a week or a month) and a limit.
- Near the limit (80% of it), or when spending is ahead of pace, only models cheaper than the requested one are used.
- Over the limit, the administrator chooses what happens: calls are blocked, sent to the cheapest model, or just flagged.
Spend is the total cost of the calls that went through the proxy, including Jev's own cost.
Turn it on
The proxy runs on the server, and each computer opts in:
- On one computer, turn on Settings → Model proxy.
- For clients, turn it on from the server under Clients → Client defaults, or select clients and choose Edit settings.
Claude Code is then pointed at SpanDock automatically. Codex can be pointed at it with the snippet SpanDock shows.
Calls use the server's provider keys, set under Settings → AI assistant. The keys stay on the server, in the OS credential store; tools on the clients authenticate only with a local proxy key that SpanDock sets up for them.
Only API-key traffic is routed. Claude Pro and Max sign-ins aren't used while the proxy is on.
What is recorded
Every call through the proxy records the requested model and the model that answered, how the choice was made (rule, Jev, budget, or unchanged), tokens, cost, savings, and Jev's cost. These calls appear in Activity, Sessions, Models and the Dashboard, and the Model routing page summarises the month.
Cost estimation
Many tools report tokens but not cost. In that case, SpanDock estimates the cost from a price catalogue and marks it as estimated. A cost the tool reports itself always wins.
Prices are taken from, highest priority first:
- your overrides, under Settings → Model prices;
- built-in Claude prices;
- OpenRouter's public model list, refreshed daily.
Model names are matched exactly after removing provider prefixes and date suffixes, so a smaller model is never priced as its larger sibling. When the catalogue changes, older records that have tokens but no cost are priced too.
The model cost leaderboard
The Models page is a leaderboard of the models your tools used. It ranks them by spend, tokens, calls, speed, reliability, or value (cost per million tokens), and shows how each model moved since the previous period. Only calls that used tokens or cost are counted, and each model's list price is shown for comparison. On a server, the leaderboard covers every client or one of them.