Skip to content

AI gateway · one key · 50+ models

Every model hasa market price.You’ve been paying list.

Smouter is an OpenAI- and Anthropic-compatible AI gateway. Every request is shopped across a benchmarked pool of providers serving the same weights, then bought at the floor that actually holds — same models, same code, up to 50% under OpenRouter.

  • OpenAI & Anthropic SDKs
  • 50+ models, one key
  • chat · image · audio · embeddings
  • free to start
app.pyone line
from openai import OpenAIclient = OpenAI(-  base_url="https://api.openai.com/v1",+  base_url="https://api.smouter.ai/v1",   api_key="sk-smouter-…",)
gpt-5.5$1.25$0.61/ 1M in
scroll

Scene one · the spread

The floor is not the cheapest quote.
It’s the cheapest one that’s real.

gpt-5.5USD / 1M input tokens9 quotes
model makerlist0.9s1.25
node A-171.1s1.10
node C-041.0s0.94
node K-221.4s0.88
node F-091.2s0.72
node B-310.9s0.69
node M-061.0s0.61
node Q-132.8sweights drifted0.58
node Z-02—never served0.54

Prices illustrate a real spread. The two struck quotes are the reason Smouter is not a price-sorted list: one node’s weights no longer match the reference, the other advertises a price it has never served.

01

Probed, not trusted

Every node is fingerprinted against the reference weights and benchmarked for latency and error rate — continuously, not at onboarding.

02

Bait priced out

A quote nobody honours is not a price. Nodes that advertise low and fail to serve are demoted before they can anchor yours.

03

Re-shopped per request

The floor moves through the day. Smouter re-quotes on every call, so you ride it down instead of pinning one contract.

See every model and what it costs

Scene two · the stack

Four layers between your prompt and the bill.

01optional−20%

Route to the model that can actually do it

Most prompts do not need your most expensive model. Smouter grades each one and sends it to the cheapest model that still nails it — or you name the model yourself and skip this entirely.

02optional−11%

Send fewer tokens for the same answer

Prompts are trimmed and compressed before they reach a model. You pay for the tokens that carried meaning, not the ones that carried habit.

03always on−27%

Buy at the floor

For whichever model you land on, the request goes to the lowest-cost node that passed its probe. This is the layer the order book above describes.

04on by default−13%

Never pay twice for the same answer

Exact and near-exact repeats return in milliseconds at a fraction of the price. Send one header when you need a genuinely fresh call.

Scene three · one key

Everything you’d otherwise buy from four vendors, on one wallet.

Chat & completions

openai · anthropic

Streaming, tool calls, JSON mode. Both SDKs, unchanged.

Images

generate & edit

Text-to-image and edits on the same key and the same wallet.

Audio

speech ⇄ text

Speech synthesis and transcription, billed by the second.

Embeddings

vectors

Batch-friendly embedding models for search and retrieval.

Web search

grounded

Ground any model in fresh results with one flag. No second vendor.

Batch

async · cheaper

Queue large jobs at a lower rate and collect them when they land.

MCP tools

your servers

Attach Model Context Protocol servers so models call your tools.

Programmable routes

one model id

Compose several models into one id your code already calls.

Scene four · programmable routes

Wire several models into one id.
Your code never learns about it.

  1. model="@triage"one id your code already calls
  2. optimizetrims the prompt before it costs you
  3. guardrailstops what should never reach a model
  4. fan outruns three models on the same inputgpt-5.5claude opus 4.8gemini 3.1 pro
  5. judgea blind judge returns the strongest answer
  6. one receiptevery stage metered per token — no platform fee on top

Built in the dashboard

Stages compose like chips. No orchestration framework, no second vendor, nothing new to deploy.

Metered like any call

Each stage is a normal per-token request and the route settles as one receipt. No platform fee on top.

Rewired without a release

Your code keeps calling the same id. Swap the judge or add a model and nothing ships on your side.

Scene five · connect with Smouter

Ship an AI app without funding everyone else’s inference.

Your appapp.yourproduct.com
Sign in with Smouter
You pay for inference
$0.00
You earn optional margin
+$0.00
Your usertheir Smouter wallet
Requests this month
0
They pay
$0.00

OAuth 2.0 + PKCE

The “sign in with” flow your users already trust. Your code keeps using any OpenAI SDK, unchanged.

You pay nothing

Inference bills the user who made the call. Set an optional margin and you earn on every one.

Users stay in control

They set a spend cap at consent and revoke your access in one click, instantly.

Set it up in your tool — every guide

Two lanes

Cheapest by default. First-party on demand.

Value

default

The sourcing pool — many nodes serving the same weights, picked per request on price and health. This is where the discounts live.

  • Up to 50% under OpenRouter
  • Caching and optimization eligible
  • Nodes probed continuously
x-smouter-lane: value

Verified

when it has to be first-party

The model maker’s own endpoint under a zero-retention contract. Nothing stored, cached or replayed; every response carries cache-control: no-store.

  • Official supply only
  • Zero retention
  • Flip per key or per request
x-smouter-lane: verified

50+

models behind one key

1

line of code to change

50%

up to, under OpenRouter

0

subscription, minimum or seat fee

  • Global edge routing, sub-50ms gateway overhead
  • Automatic failover the moment a node degrades
  • Per-key spend caps, model allowlists and rate limits
  • Guardrails on prompts and responses, inbound and outbound

Pricing

Subscribe, or just pay for the tokens.

Most popular

Plus

$20/mo

Billed monthly

  • Every frontier model — GPT, Claude, Gemini, Grok, DeepSeek
  • More usage per dollar than a single-vendor plan
  • Runs in any coding agent or IDE
  • Chat, image and research included
Subscribe

Pro

$50/mo

For heavy builders

  • Everything in Plus
  • About 3× the usage
  • Priority routing on busy models
  • Verified-lane access
Subscribe

API

Per token

No subscription

  • OpenAI- and Anthropic-compatible
  • Chat, image, audio, embeddings and search
  • No minimums, no seats
  • Build your own routes
See model pricing

Change one line. Keep the rest.

base_url="https://api.smouter.ai/v1"

No card to start. Your first credit is on us, and nothing renews on its own.