sagerouter is a metered, OpenAI-compatible API for open-weight models — GLM, Qwen, Kimi and friends. We publish our rate and the prevailing market rate side by side on every model we list, including the ones where the win is small. Then we let you set a hard cap so the bill can never surprise you.
No card. Prepaid credits only, so there is no invoice to be shocked by.
# the request you already send OpenAI curl https://sagerouter.ai/v1/chat/completions \ -H "Authorization: Bearer sk-sr-..." \ -d '{"model":"glm-5.2","messages":[...]}' # what comes back — metered, itemised "usage": { "prompt_tokens": 1842, "completion_tokens": 214, "cost_usd": — }
Same request you send OpenAI. Different base_url. The cost
above is that call priced at glm-5.2's rate in the table
below — computed on this page from the same numbers, not typed in.
Per 1M tokens, USD. "Market" is the prevailing published rate for the same model on the major aggregators. We are cheaper on almost every line — sometimes by a lot, sometimes by a little. We show you which is which, because a vendor who hides the small numbers is hiding something.
Prices are prepaid and metered per token. No minimum, no monthly platform fee, no seat licence. Models we cannot currently serve are not listed — we would rather show you a shorter list than sell you a name that returns an error. The roster grows as we validate capacity, and every addition ships with its market comparison on day one.
Two numbers to move. The answer updates as you drag. This is the same arithmetic that runs on your real invoice — there is no second, worse price waiting after signup.
The gap between the two bars is what sagerouter saves you. The gap to the frontier number is what open-weight models save you — that one is not our doing, and we are not going to pretend it is.
The failure everyone has a story about is a runaway agent loop burning a month of budget over a weekend. Set a hard monthly cap and we stop serving at the cap — not a warning email, not an overage line, a refusal with the reason in it.
HTTP 402Credits are prepaid, so the cap is a second floor under a floor: we cannot bill you for money you have not already put in. The error names the limit rather than the balance, so nobody gets sent to add funds that would not have helped.
Cheap inference with no explanation should worry you. Here is ours, in plain language.
Most of the saving isn't ours — it's the model class. GLM-, Qwen- and Kimi-class weights cost a fraction of a closed frontier model to serve. Any router will give you that. We say so instead of billing you for it as if we invented it.
This is the part that is actually ours. We source each model on the cheapest channel we can verify by billing it — one real metered call per model, then read the coefficient off the supplier's own billing log, rather than trusting a rate card. That is where the rest of the gap comes from.
Your price is our measured cost times a fixed multiplier. It does not vary by model, by customer, or by how much you spend — there is no volume tier you are failing to qualify for. When our cost falls, the table above falls with it and you renegotiate nothing.
You're routing client work through a company you'd never heard of ten minutes ago. You should get straight answers before you paste a brief.
Free credits on the house when your account is created — enough to run your real workload against the models above and check the arithmetic yourself. No card, no contract, no sales call. Alpha seats open in small batches.
One email when your seat opens, with a login and a ready-to-run curl. No drip campaign, unsubscribe anytime.