Everything else · Published October 4, 2026

Preventing budget excess by routing ai calls differently

Choosing a new path to prevent budget overspending

The situation

They had to stop public ai calls from exceeding a shared account budget while keeping existing public page behavior. The public path was calling ai before counting the spend, which made it hard to refuse calls based on budget. The pool's model runner already checked admission, ran inference, counted usage, and reported shared account spend. They needed to find a way to close the gap in the public path and prevent budget excess.

What they weighed

  • Route every public model call through the shared pool's budgeted model runner, with purposes as data.

    Chosen
  • Keep counting public model calls only after inference.

What they decided

Route every public model call through the shared pool's budgeted model runner, with purposes as data.

Why

They chose this option because the existing pool runner already checked the budget before inference and reported usage. By routing every public model call through the pool's budgeted model runner, they could close the gap in the public path with one generic solution. This way, they could ensure that the account budget was checked before each call and prevent excess spending.

The first step

They would edit the system and pool, add a focused account budget test, run it, and commit without deployment.

How they would know it worked

The test would prove that a reached shared account budget refuses the next call before the ai binding runs.

How it turned out

Not known yet. The result is added here when it comes in.

Shared by the person who made this decision. Names, places, dates and numbers were removed before it was published.