How response caching works
Exact and prefix cache hits: lower latency and cost, visible per request.
Smouter can serve repeated work from cache: identical requests (and requests sharing a long identical prefix) may hit a cached response instead of a fresh upstream call, which lowers both latency and cost. The cache tier of every request is visible per row on Usage.
Caching respects your data settings and can be controlled per request — see the docs for the exact headers and defaults.
Still stuck? Open a ticket — we reply by email — or write to support@smouter.ai.