Everything else · Published October 2, 2026
Restarting agents after usage limit issues
Agents stopped due to usage limits
The situation
A person instructed to restart agents that stopped working due to a usage limit issue. The agents were previously running but stopped due to a strange out-of-usage or quota issue. Other messages in the conversation reported that some agents had been restarted, but the issue was recurring. The person wanted to ensure that all affected agents were restarted and that a solution was put in place to prevent the issue from happening again. A watchdog was suggested to auto-retry any agents that hit the usage limits again.
What they weighed
Act on the message and verify the result
Keep it as context when no action is called for
What they decided
Restart all affected agents and put in a watchdog
Why
The person explicitly asked to start all agents that died due to an out-of-usage issue. Messages in the same conversation reported that agents were being restarted but also noted that rate-limit cutoffs were recurring, so both restarting and adding a retry watchdog were necessary. The goal was to get all the previously running agents up and working again and to prevent the issue from happening again in the future.
The first step
List all running and stalled workflow runs and agents, restart the stalled ones, and schedule an hourly check that restarts any that report rate-limit stoppage.
How they would know it worked
A current roster shows zero stalled agents due to usage limits, and the watchdog run reports at least one successful check in the next hour.
How it turned out
Not known yet. The result is added here when it comes in.
Shared by the person who made this decision. Names, places, dates and numbers were removed before it was published.