Skip to content

How we keep the 50% floor honest

A look at how we benchmark provider prices in the open, how we pick the cheapest path per request, and how we'd spot it ourselves if the floor slipped.

Caching that's actually safe for chat

Response caching is the easiest way to drop a bill in half — and the easiest way to ship a bug. Here's the line we draw, and why.

Auto-routing vs. picking your own model

When letting Smouter choose the model saves more than fixing it, and when you really do want to lock in gpt-5.

We're heads-down on the alpha. Real posts will land here as soon as there's something worth reading.