Route to the model that can actually do it
Most prompts do not need your most expensive model. Smouter grades each one and sends it to the cheapest model that still nails it — or you name the model yourself and skip this entirely.
AI gateway · one key · 50+ models
Smouter is an OpenAI- and Anthropic-compatible AI gateway. Every request is shopped across a benchmarked pool of providers serving the same weights, then bought at the floor that actually holds — same models, same code, up to 50% under OpenRouter.
from openai import OpenAIclient = OpenAI(- base_url="https://api.openai.com/v1",+ base_url="https://api.smouter.ai/v1", api_key="sk-smouter-…",)Scene one · the spread
Prices illustrate a real spread. The two struck quotes are the reason Smouter is not a price-sorted list: one node’s weights no longer match the reference, the other advertises a price it has never served.
Every node is fingerprinted against the reference weights and benchmarked for latency and error rate — continuously, not at onboarding.
A quote nobody honours is not a price. Nodes that advertise low and fail to serve are demoted before they can anchor yours.
The floor moves through the day. Smouter re-quotes on every call, so you ride it down instead of pinning one contract.
Scene two · the stack
Most prompts do not need your most expensive model. Smouter grades each one and sends it to the cheapest model that still nails it — or you name the model yourself and skip this entirely.
Prompts are trimmed and compressed before they reach a model. You pay for the tokens that carried meaning, not the ones that carried habit.
For whichever model you land on, the request goes to the lowest-cost node that passed its probe. This is the layer the order book above describes.
Exact and near-exact repeats return in milliseconds at a fraction of the price. Send one header when you need a genuinely fresh call.
Scene three · one key
openai · anthropic
Streaming, tool calls, JSON mode. Both SDKs, unchanged.
generate & edit
Text-to-image and edits on the same key and the same wallet.
speech ⇄ text
Speech synthesis and transcription, billed by the second.
vectors
Batch-friendly embedding models for search and retrieval.
grounded
Ground any model in fresh results with one flag. No second vendor.
async · cheaper
Queue large jobs at a lower rate and collect them when they land.
your servers
Attach Model Context Protocol servers so models call your tools.
one model id
Compose several models into one id your code already calls.
Scene four · programmable routes
model="@triage"one id your code already callsone receiptevery stage metered per token — no platform fee on topStages compose like chips. No orchestration framework, no second vendor, nothing new to deploy.
Each stage is a normal per-token request and the route settles as one receipt. No platform fee on top.
Your code keeps calling the same id. Swap the judge or add a model and nothing ships on your side.
Scene five · connect with Smouter
app.yourproduct.comtheir Smouter walletThe “sign in with” flow your users already trust. Your code keeps using any OpenAI SDK, unchanged.
Inference bills the user who made the call. Set an optional margin and you earn on every one.
They set a spend cap at consent and revoke your access in one click, instantly.
Two lanes
The sourcing pool — many nodes serving the same weights, picked per request on price and health. This is where the discounts live.
x-smouter-lane: valueThe model maker’s own endpoint under a zero-retention contract. Nothing stored, cached or replayed; every response carries cache-control: no-store.
x-smouter-lane: verified50+
models behind one key
1
line of code to change
50%
up to, under OpenRouter
0
subscription, minimum or seat fee
Pricing
$20/mo
Billed monthly
$50/mo
For heavy builders
Per token
No subscription
base_url="https://api.smouter.ai/v1"No card to start. Your first credit is on us, and nothing renews on its own.