Blog

Routing between LLM providers to balance cost and quality

3 min read

  • llm
  • cost
  • evals

One backend runs the conversational AI in two of our products: Cabin, a companion for thinking decisions through, and Mo, the AI inside QuesMo. Both stream replies to people on phones, and both have to stay affordable at consumer prices.

That puts two pressures on every request. The reply should be good enough that the person comes back, and cheap enough that we can afford it when they do. No single model wins on both for every kind of request, so we route.

Not every request needs the same model

A conversation has many kinds of work inside it. Some turns need care: a person is weighing a hard choice, and the reply has to be thoughtful and short. Others are routine: a title for the conversation, a summary for memory, a label for what the message is about.

Sending all of that to the most capable model wastes money. Sending all of it to the cheapest one hurts quality exactly where it matters most. Routing lets each kind of task go to a model that is good enough for it, and no more expensive than it needs to be.

Make routing a setting, not a deploy

Our routing lives in runtime configuration. Which provider and model handle which task is a setting the backend reads, not a constant in the code. We can move a task to a different model without shipping a new build, and the change can be rolled back the same way.

Two design rules make this practical. First, the code talks to models through one internal interface, so a task does not care which provider answers it. Differences in request formats, streaming events and error types are handled in one adapter per provider. Second, plan for failure. If a provider times out or rate-limits you, the request should go somewhere sensible instead of failing in front of a user.

Cache what repeats

Much of every request is the same from turn to turn: the system prompt, the product’s instructions, and long-lived context about the person. Providers now offer prompt caching, which charges much less for input the model has seen recently.

To benefit, the prompt has to be built for it. Stable content goes first, in a fixed order. Anything that changes each turn, like the latest message or the current time, goes at the end. A small change near the top, such as a timestamp, quietly breaks the cache for everything after it. We treat prompt layout as part of cost design, not just wording.

Evaluate offline before anything changes

A cheaper model that looks fine in a few manual tries can still be worse in ways that matter: longer replies, a softer opinion, a missed detail from memory. So no prompt or model change reaches users until it has been through offline evaluations.

The idea is simple. Replay a fixed set of realistic conversations through the current setup and the proposed one, then compare. Replies are judged against what each product promises. Cabin, for example, promises short answers and a clear opinion instead of “it depends”. Some checks are plain rules, like reply length. Others need judgement, which is where a written rubric and reading samples by hand come in.

Because two products share this backend, every change is checked against both. A tweak that helps one can easily hurt the other, and we would rather find that in an eval than in reviews.

What we would tell another team

Route by task, not by product. Keep the choice in configuration, behind one interface, with a plan for when a provider fails. Lay out prompts so the stable parts can be cached. Build the eval set before you need it, because the day a new model launches is a bad day to find out you have no way to compare it.

None of this needs a large team. It needs the discipline to change one thing at a time and measure it before it ships.

Work with usNeed something like this built? Say hello