Everything else · Published October 4, 2026
Preventing budget excess by routing ai calls differently
Choosing a new path to prevent budget overspending
The situation
They had to stop public ai calls from exceeding a shared account budget while keeping existing public page behavior. The public path was calling ai before counting the spend, which made it hard to refuse calls based on budget. The pool's model runner already checked admission, ran inference, counted usage, and reported shared account spend. They needed to find a way to close the gap in the public path and prevent budget excess.
What they weighed
Route every public model call through the shared pool's budgeted model runner, with purposes as data.
Keep counting public model calls only after inference.
What they decided
Route every public model call through the shared pool's budgeted model runner, with purposes as data.
Why
They chose this option because the existing pool runner already checked the budget before inference and reported usage. By routing every public model call through the pool's budgeted model runner, they could close the gap in the public path with one generic solution. This way, they could ensure that the account budget was checked before each call and prevent excess spending.
The first step
They would edit the system and pool, add a focused account budget test, run it, and commit without deployment.
How they would know it worked
The test would prove that a reached shared account budget refuses the next call before the ai binding runs.
How it turned out
Not known yet. The result is added here when it comes in.
Shared by the person who made this decision. Names, places, dates and numbers were removed before it was published.