Skip to content

Quickstart

  1. Create an account, verify your email, and add credits in Billing.
  2. Create a key in API keys. The plaintext sk-smouter-… value is shown once.
  3. Call GET /v1/models, select an available model of the right kind, and export its id as MODEL_ID.
  4. Set the OpenAI base URL to https://api.smouter.ai/v1.
curl https://api.smouter.ai/v1/chat/completions \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "'"$MODEL_ID"'",    "messages": [{"role":"user","content":"Say hello."}]  }'

Connect a coding tool

Using Smouter from Cursor, Claude Code, OpenCode, Codex, or another agent instead of your own code? Each tool has a step-by-step guide — the exact fields, a copyable config, and the failure modes we actually see.

All of them reduce to the same three values — base URL, key, model id — collected on the connect hub.

Authentication & per-key controls

Every API route accepts Authorization: Bearer sk-smouter-…, and every route also honors the Anthropic-style x-api-key header as a fallback when no bearer is present (Bearer wins if both are sent). Keys are hashed at rest. Delegated Connect tokens (smd_…) authenticate the chat-shaped surfaces only — embeddings, images, and audio return 403 app_scope_forbidden for them.

  • Scopes: optional model and provider allowlists plus an expiry. Every fallback and guardrail model is checked against the allowlist.
  • Spend controls: monthly cap, daily cap, and warning threshold are independent positive integer micro-USD values. Clearing one does not clear the others; the warning threshold is advisory.
  • TPM: an integer from 1 through 2,147,483,647, or null for unlimited. An org TPM limit is an aggregate ceiling across its keys; the stricter key or org limit wins.
  • Rotation: rotate in place from the dashboard. The key id, scopes, caps, and history stay attached; the old secret stops authenticating immediately and the replacement is shown once.
  • Revocation: revoking a key takes effect on its next request.

Zero retention

Enable it on the key or send x-smouter-zero-retention: on. Smouter does not retain full prompts or responses for replay or caching, and suppresses customer-content safety snippets and Anthropic signed-thinking continuity. Usage and billing are still metered, compliance events can retain non-content metadata, and an extended-thinking tool conversation started under zero retention cannot later recover its discarded reasoning context. The effective state is reflected as x-smouter-zero-retention: on.

Chat completions

POST /v1/chat/completions accepts OpenAI-style messages, tools, structured output, sampling parameters, and HTTPS or base64 image content when the selected model supports them. Unsupported provider-specific fields can still fail at model dispatch.

Streaming

Set stream:true for canonical server-sent events. The final charge for a live stream is settled after consumption and is visible in Usage. A cache replay is returned as the same stream shape.

stream = client.chat.completions.create(    model=os.environ["SMOUTER_MODEL_ID"],    messages=[{"role": "user", "content": "Write a haiku."}],    stream=True,)for chunk in stream:    print(chunk.choices[0].delta.content or "", end="")
  • n must be 1.
  • The maximum completion size is a deployment setting, not a fixed number — handle the 400 rather than hardcoding a ceiling.
  • On reasoning models, the tokens spent reasoning are drawn from the same max_tokens allowance as the visible answer. A budget that is too tight for the task returns finish_reason:"length" with truncated — or empty — content, which is the OpenAI-protocol behaviour rather than an error. When that happens the response carries an x-smouter-warning header naming the budget and how much of it reasoning consumed. Raising max_tokens costs nothing extra: you are billed on tokens generated, not on the ceiling you set.
  • store and provider cache-control extensions are stripped. Use Smouter cache headers instead.
  • On the verified lane, images must be inline base64 data: URLs — remote image_url references and audio content parts return 400. The value lane accepts HTTPS image URLs.

Models, kinds & pricing units

GET /v1/models is public. Availability is computed from live routable capacity and is the source of truth at dispatch time. Every row carries kind and pricing_unit; a defensive client can still treat a missing kind as chat.

curl https://api.smouter.ai/v1/models {  "object": "list",  "data": [{    "id": "claude-opus-4.6",    "name": "Claude Opus 4.6",    "owned_by": "Anthropic",    "kind": "chat",    "pricing_unit": "mtok",    "lanes": {      "verified": {"available":false,"input_per_mtok_usd":0,"output_per_mtok_usd":0},      "value": {"available":true,"input_per_mtok_usd":2.5,"output_per_mtok_usd":12.5}    }  }]}
KindUnitInterpret the lane price as
chat, embeddingmtokUSD per 1 million input/output tokens; embeddings are input-only.
imageimageUSD per generated image.
ttscharUSD per 1 million input characters.
sttrequestUSD per transcription request.

The legacy lane field remains named input_per_mtok_usd for compatibility; use pricing_unit to interpret it for image and audio rows. Image, embedding, and audio rows advertise the value lane only. A token-metered image model, when one is live, advertises kind: "image" with pricing_unit: "mtok" plus an image_price_estimate_usd estimate — branch on pricing_unit, not on kind.

Embeddings

POST /v1/embeddings accepts one string or an array of up to 2,048 strings (1,000,000 total characters). encoding_format may be float or base64; dimensions is passed through when the chosen embedding model supports it. Embeddings bill input tokens and enforce model/provider scopes, spend caps, moderation, and an input-only TPM reservation. They do not use response caching, idempotent replay, affinity, auto-routing, or programmable routes.

curl https://api.smouter.ai/v1/embeddings \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{"model":"'"$EMBEDDING_MODEL_ID"'","input":"Search this sentence.","encoding_format":"float","dimensions":256}'

Images

POST /v1/images/generations bills per generated image. prompt is required and n defaults to 1; it must be an integer from 1 through 10 (above 10 returns 400 invalid_request_error). A configured safety policy moderates the prompt before any spend, and a provider returning more images than requested is clamped to n — you are never billed above it.

The body is validated fail-closed before a hold is opened:

  • Accepted keys: model, prompt, n, size, quality, style, response_format, user. Any other key returns 400 naming the parameter.
  • size may be 256x256, 512x512, or 1024x1024; omitting it sends 1024x1024 explicitly, so upstream defaults never apply.
  • quality: "hd" is not supported and returns 400.
  • prompt is capped at 4,000 characters (4,096 UTF-8 bytes).

Editing an image

POST /v1/images/edits takes multipart/form-data — a prompt plus the reference image(s) to transform — and returns the same { created, data } shape as generation, billed at the same per-image price. Not every image model accepts a reference image, so check the catalog rather than assuming: rows that do carry supports_image_edit: true in GET /v1/models, and any other model returns 400 before a hold is opened.

  • Send one file as image, or several as image[] — not both spellings in the same request. Repeated parts are all used.
  • PNG and JPEG only, decided by the file's actual bytes rather than its declared Content-Type. Truncated or non-image uploads return 400 before any spend.
  • Accepted fields: model, prompt, n, response_format, user. n must be 1.
  • mask and size return 400 rather than being accepted and ignored — the serving model edits from your prompt and reference, and renders at its own resolution.
  • Upload size, per-file size, and image-count ceilings are deployed configuration; oversized uploads return 413, and a burst of concurrent uploads may return 503 with retry-after.
curl https://api.smouter.ai/v1/images/generations \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{"model":"'"$IMAGE_MODEL_ID"'","prompt":"A paper bird on a blue desk","n":1}'
Image requests are not deduplicated: Idempotency-Key is not honored on /v1/images/*, so a retry generates — and bills — a new image.

Upstream faults are never passed through: a provider's auth, quota, or throttle failure surfaces as Smouter 503 service_unavailable (with retry-after when the upstream throttled), so retry logic only handles Smouter statuses.

Audio

The gateway implements POST /v1/audio/speech for TTS and POST /v1/audio/transcriptions for multipart STT. Audio rows are published only when an enabled model has a positive price, a routable provider, a live secret, and positive margin.

Audio availability is the catalog: a model serves when its row (kind of tts or stt) appears in GET /v1/models — check the catalog rather than assuming a roster.
curl -s https://api.smouter.ai/v1/models | jq '  [.data[] | select(.kind == "stt" or .kind == "tts") |   {id, kind, pricing_unit, lanes}]' # [] means audio is not currently provisioned for customer traffic.

TTS bills per character of input text. STT bills per second of audio, read from the uploaded file's duration — a file whose duration cannot be determined is rejected with 400. TTS voice accepts the OpenAI voice names on OpenAI-named models (unrecognized voices pass through to the upstream and may be rejected there). File-size, text-length, and response-size ceilings are deployed configuration, not fixed client contracts. TTS input is moderated; textual STT output can be scanned before delivery.

Responses API

POST /v1/responses translates Responses input, instructions, function tools, and output limits into the canonical routing path. Non-streaming calls return a Responses object. Streaming calls emit response.created, response.in_progress, output item/content/delta events, then response.completed; an interrupted generation ends with response.failed.

curl https://api.smouter.ai/v1/responses \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{"model":"'"$MODEL_ID"'","instructions":"Be concise.","input":"Explain a Merkle tree."}'

The endpoint is stateless. A non-empty previous_response_id returns 400 invalid_request_error: resend the full conversation context in input. store and metadata are accepted and ignored. A truncated generation ends with status: "incomplete" — streams terminate with response.incomplete carrying incomplete_details — and every stream event has a monotonic sequence_number.

Batch API

The Batch API asynchronously processes JSONL-style chat-completion lines. POST /v1/batches takes a JSON envelope whose input field is an array of line objects or a single raw JSONL string. Give each line a unique custom_id to correlate results (an omitted id defaults to req-<n>); method and url are optional but must be POST and /v1/chat/completions when present, and every line body needs model plus messages.

curl https://api.smouter.ai/v1/batches \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "completion_window":"24h",    "input":[      {"custom_id":"line-1","method":"POST","url":"/v1/chat/completions",       "body":{"model":"'"$MODEL_ID"'","messages":[{"role":"user","content":"Hello"}]}}    ]  }'
  • GET /v1/batches lists the 20 most recent jobs; GET /v1/batches/:id polls one. Jobs expire after the completion window.
  • Results are application/x-ndjson, include a trailing newline, and return 409 batch_not_ready until ready. Lines always run non-streaming.
  • Cancellation is best effort. Status is checked before each line; the line already in flight may finish, while later lines do not start.
  • Each line rechecks the key's scopes, spend caps, zero-retention state, and org safety policy at execution time. Async batch work has no minute-window TPM gate.
  • A zero-retention key is rejected with invalid_request_error: “Batch storage is incompatible with zero-retention keys: batch inputs and results are persisted for retrieval.”
  • Delegated Connect tokens cannot use batch (403 app_scope_forbidden), and org-billed keys are not yet supported (400 org_batch_unsupported).

Limits and discount

Line and byte ceilings are deployment settings — read the current values from GET /account/batches/config. Each job snapshots an integer discount_bps, and the displayed completion discount, discount_bps / 100 percent, is a ceiling: the realized credit is additionally capped by each line's proven routing margin, so thin-margin models can refund less than the headline rate. The credit is finalized only after a completed job, never a cancelled one. Open the batch dashboard.

MCP tools

Register account-owned HTTP or SSE MCP servers in the MCP dashboard. Secrets are write-only: leaving the secret blank during an edit keeps it; clearing it atomically switches authentication to none. Refresh caches the server's tool list, and disabled servers cannot refresh or run tools.

curl https://api.smouter.ai/v1/mcp/tools/call \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -d '{"server":"weather","tool":"get_weather","arguments":{"city":"Berlin"}}'

POST /v1/mcp/tools/call resolves the named server inside the authenticated account and returns {server, tool, content, is_error}. The proxy call itself is not model-token billing; model use of its output is metered by the later model request. Server URLs must be publicly routable — private, loopback, and cloud-metadata hosts are rejected. A tool that runs but reports failure returns 200 with is_error: true; transport and protocol faults surface as typed 502s.

Metadata, tags, fallbacks & guardrails

curl https://api.smouter.ai/v1/chat/completions \  -H "Authorization: Bearer $SMOUTER_API_KEY" \  -H "Content-Type: application/json" \  -H "x-smouter-tags: production,checkout" \  -H "x-smouter-models: $FALLBACK_MODEL_ID" \  -H "x-smouter-guardrail-policy: Block requests containing secrets" \  -H "x-smouter-guardrail-model: $GUARDRAIL_MODEL_ID" \  -H "x-smouter-guardrail-on-block: error" \  -d '{    "model":"'"$MODEL_ID"'",    "metadata":{"customer":"acme","flow":"checkout"},    "messages":[{"role":"user","content":"Summarize this."}]  }'

Metadata and tags

  • metadata must be a JSON object with at most 16 keys and at most 4,096 serialized JavaScript characters/code units. It is never forwarded to the provider — it stays in Smouter as request labeling, and it does not affect response-cache matching.
  • x-smouter-tags is split on bare commas, trimmed, and deduplicated — at most 10 tags of 64 characters each. There is no escaping, so a comma cannot appear inside one tag.
  • Exceeding any of these limits returns 400 — nothing is silently truncated.
  • Tags echo in x-smouter-tags. Metadata echoes in a successful non-stream JSON body; both remain queryable in account logs.

Fallback models

Provide models:[…] in the body, or x-smouter-models when the body field is absent. IDs are canonicalized and deduplicated in order; more than 16 returns 400. auto and @route are not allowed. Smouter selects the first viable candidate at dispatch; this is not a retry after a model has begun generating. All candidates obey the key/grant model and provider allowlists, and the fallback list is never forwarded upstream.

Guardrail headers

  • x-smouter-guardrail-policy enables the classifier; it is limited to 2,000 characters. Each classification is a metered model call charged to your wallet — if it cannot be funded, the request returns 402 insufficient_quota.
  • x-smouter-guardrail-model selects a concrete classifier. A model outside the caller's allowlist returns 400 before dispatch. Multimodal content cannot be classified: under on-block:error it returns 400 without a billed check.
  • x-smouter-guardrail-on-block is error by default: block or unparseable verdicts return 400. fallthrough is advisory — a flagged response is served with x-smouter-guardrail-flag: output; flagged input is served without a customer-visible marker.
  • x-smouter-guardrail-output:on also classifies the completed answer. Output generation has already been charged if that answer is withheld.
  • stream:true conflicts with output guardrail + on-block:error, and with a blocking/fail-closed model moderation policy. Retry non-streaming. Regex-only output policies can scan a live stream.

Stored safety policies

Configure a policy on an API key or organization from the dashboard. A key policy takes precedence; an org policy applies to its keys that have no key-specific row. Policies can combine moderation categories, injection/jailbreak detection, secret-leak output scanning, block/flag mode, and fail-closed behavior. Blocking model-based output policies require non-streaming calls. Policy writes and deletions are audited.

Lanes, auto-routing & programmable routes

Send x-smouter-lane: value|verified; value is the default, and any unrecognized value is treated as value. The catalog reports availability per lane. A response reflects the lane in x-smouter-lane. There is no per-key default lane — the header decides per request.

  • Auto: model:"auto" resolves to one concrete available model before pricing. It is account-gated, not cacheable, and unavailable to scoped credentials that cannot authorize every possible resolution.
  • Routes: invoke a dashboard-defined pipeline with model:"@route-name". Stages can optimize, guardrail, fan out, judge, and return; billed sub-calls settle under one route execution. Response headers include x-smouter-route-name and x-smouter-route-stages. Route invocations are non-streaming (stream:true returns 400) and unavailable to delegated app tokens.

Caching & reliability

Exact response caching is account-isolated and applies only when the request is eligible — deterministic sampling, meaning an explicit temperature: 0 or a pinned seed. Prefix savings use compatible upstream warm-prefix accounting. Semantic caching is explicit opt-in via x-smouter-cache-semantic: on and excludes tool, structured-output, and streaming calls.

# Skip cache lookup AND storage for this request-H "x-smouter-cache: no-store" # Skip the lookup but still store the fresh answer-H "x-smouter-cache: no-cache" # Opt this request into semantic caching-H "x-smouter-cache-semantic: on" # Five-minute reuse window-H "x-smouter-cache-ttl: 300" # Partition cache entries inside your account-H "x-smouter-cache-namespace: user-8f2c"
  • idempotency-key replays a completed response without charging again — a replay carries x-smouter-idempotent-replay: true. Reusing a key while the original is still running returns 409 conflict; reusing it with a different body returns 422 invalid_request_error. Replay applies to the chat-shaped surfaces only.
  • Zero retention intentionally disables replay-body storage; a later retry with the same key then returns 502 idempotency_error instead of a replay.
  • x-request-id supplies your correlation id.
  • x-smouter-conversation-id supplies a stable affinity key for eligible multi-turn traffic.

Billing & receipts

Smouter reserves against prepaid account credit, then settles the delivered work. Spend caps run before admission, but output policy checks and some scoring failures can happen after billable work. Therefore an error status is not by itself proof that the charge is zero; inspect the receipt and Usage.

x-smouter-lane: valuex-smouter-cache: missx-smouter-charge: $0.0004x-smouter-baseline-openrouter: $0.0031x-smouter-saved: $0.0027

Receipt headers are present when their values are known; a live stream's final charge is settled after the response body drains. Alongside x-smouter-charge and x-smouter-saved, a response can carry x-smouter-prefix-saved and x-smouter-tokens-saved when those savings applied, and x-smouter-warning when the completion ended against its max_tokens ceiling rather than finishing on its own. The dashboard shows the same usage, tags, key attribution, and audited account changes.

Logs, audit, safety events & scoring

  • Usage is backed by GET /account/logs and supports applied filters for an exact model id, lane, status, key, tag, and [from,to) time range. Rows include status, tokens, cost, key attribution, metadata, and tags.
  • Audit log records sensitive account, key, batch, route, Connect, MCP server, organization, and policy changes without plaintext secrets. It filters by time, exact action, and target type.
  • Safety events filters block/flag events by time, endpoint, stage, action, category, and key. Under zero retention, evidence is replaced with non-content redaction metadata.
  • The Playground can score the latest answer from 0–10 or compare A/B answers, including a tie. Scoring is a metered judge call. A charged parse failure can return 502 while preserving x-smouter-charge and x-smouter-lane.

Organizations

Organizations provide role-based members, org-scoped keys, aggregate usage, per-member spend roll-up, and seat capacity. When capacity is unlimited, the first seat action sets a finite capacity rather than adding to an invisible number, and seat removal cannot cut capacity below current membership. Org owners/admins can also set an aggregate TPM ceiling and an org safety policy that applies to org keys without their own policy.

Organization billing

An organization can carry its own wallet. The default billing mode, member_personal, leaves every member paying from their own wallet. An owner can switch the org to org_wallet mode (step-up verified and audited, effective immediately): requests on org-scoped keys then reserve and settle against the organization wallet instead of the member's. Admins fund the wallet by card checkout or by transferring their own personal credit; every member can view the wallet balance, state, and recent ledger. Org-billed usage is always pay-as-you-go — it never consumes a member's subscription quota.

  • An empty org wallet fails the request with 402 insufficient_quota; a frozen wallet returns 403 org_frozen.
  • If org billing is administratively unavailable, enrolled org keys fail closed with 403 org_billing_unavailable — they never silently fall back to billing members.
  • Batch is not yet available on org-billed keys (400 org_batch_unsupported), and delegated app calls always bill the connected user, never the org.

Errors & rate limits

Errors use {"error":{"message":"…","type":"…"}}; the Anthropic-compatible /v1/messages surface re-shapes the same errors into the Anthropic error envelope. Branch on the status and the returned type — two 429s or two 502s can have different causes — and treat 500 api_error as the unexpected-failure backstop.

StatusError typesMeaning
400invalid_request_error, model_not_found, org_batch_unsupportedMalformed or unsupported input, a blocked guardrail, an incompatible option, or — on the chat surfaces — an unavailable model.
401authMissing, invalid, expired, or revoked credential.
402insufficient_quota, per_key_cap_exceededWallet (personal or org), guardrail budget, or per-key spend admission failed before the requested model call.
403account_frozen, key_model_not_allowed, org_frozen, org_billing_unavailable, app_scope_forbidden and other app_* typesAccount state, credential scope, org billing state, or delegated-app limits forbid the operation.
404not_found, model_not_foundThe owned resource is missing, or — on embeddings, images, and audio — the requested model is not available.
409conflict, batch_not_readyThe same idempotency key is still in flight, or batch results were requested before they are ready.
413payload_too_large, invalid_request_errorThe request body exceeds the deployed size ceiling.
422invalid_request_errorAn idempotency-key was reused with a different request body.
429rate_limit, rate_limit_exceeded, tokens_per_minute_exceededRequest, subscription, or TPM window exhausted. Honor retry-after.
502api_error, idempotency_errorAn upstream or scoring operation failed, or a zero-retention replay had no stored body. A settled scoring attempt can still carry x-smouter-charge.
503service_unavailableNo healthy capacity, including a provider throttle masked as Smouter unavailability.

TPM headers

Chat completions, Messages, Responses, and embeddings reserve against key/org TPM. A rejection returns 429 tokens_per_minute_exceeded with retry-after, x-ratelimit-reset, x-ratelimit-remaining-tokens, and x-ratelimit-limit-scope. Embeddings reserve input only. Images, audio, and async batch are not TPM-gated.

The requests-per-minute ceiling is a deployment setting, not a published quota — treat any 429 rate_limit as the signal and honor retry-after.

Security & privacy

  • Usage records store model, outcome, token/unit counts, charge, attribution metadata, and timestamps—not full prompt or completion bodies.
  • Eligible response bodies may be retained for account-isolated caching and idempotent replay unless zero retention is effective.
  • Fallback controls, request metadata, account ids, and Smouter credentials are never forwarded as provider parameters — the chosen upstream receives only the content required to perform the request.
  • Safety policies can record categorized events and, without zero retention, short redacted evidence. Customer zero retention replaces that evidence with non-content redaction metadata.

See the Privacy policy and Data processing page for the broader policy terms.

Coding agents & frameworks

Coding tools (Cursor, Claude Code, OpenCode, Codex, Cline, Aider, Zed, …) each have a dedicated setup guide above. Frameworks configure like any OpenAI-compatible endpoint: base URL, key, and a live model id. Tool-specific support changes independently of the gateway, so verify the tool accepts a custom OpenAI base URL.

import { createOpenAI } from "@ai-sdk/openai"; const smouter = createOpenAI({  baseURL: "https://api.smouter.ai/v1",  apiKey: process.env.SMOUTER_API_KEY,});const model = smouter(process.env.SMOUTER_MODEL_ID);
Need help reproducing a typed error or choosing the right surface? Visit the support center — help articles and tickets — or email support@smouter.ai.