How do you keep AI agent costs under control?

Use organization-managed provider keys, set daily, weekly, or monthly allowances, and review usage by model, team, or person. In OpenWork, use cheaper models for routine tasks and reserve frontier models for work that needs them.

Updated

  • Spend limits per day, week, or month, in USD
  • Limits for everyone, a team, or one person
  • Usage by model, team, or person for the last 31 days
  • Pause, warn, or approve requests for 25% more

Why AI agent costs get out of hand

Agents use far more tokens than chat. One task can read dozens of files, call tools, and retry, and every step is billed. When people pay with personal keys, the company can't see the total until the invoices arrive.

Most of the waste comes from three places: no limits, no visibility, and using a frontier model for questions a small model could answer.

Set spend limits

Spend limits apply to priced organization-key providers; they do not cap every external bill or personal provider key. In the AI Gateway, open Limits and choose who a limit applies to: everyone in the organization (including people who join later), a team, or one person. Turn on any mix of a daily, weekly, and monthly amount. Whichever runs out first applies.

What happens when someone reaches a spend limit
ChoiceWhat happens
Pause their modelsRequests stop until the limit resets or an admin gives more.
Only warnRequests keep working and the limit shows as over.
Let them ask for 25% moreThe person can request more, and admins approve or deny it in Limits.
  • In the desktop app, the chat says which limit was reached and when it resets.
  • Members see what's left in the account menu and in Settings › Usage.

See who spends what

  • The AI Gateway overview shows who spends the most and which models people use.
  • Usage charts show tokens or cost for the last 31 days, grouped by model, team, or person.
  • Each person's page shows what they can use, their limit, and their recent spend.
  • Costs come from token counts and published model prices, including cached input and reasoning tokens.

Use cheaper models where they're good enough

  • OpenWork Models gives your team hand-picked open models, such as GLM, Kimi, and DeepSeek, for $10 per user per month, with no API keys to manage.
  • Local models through Ollama or LM Studio cost nothing per token.
  • Grant expensive models only to the teams that need them, and turn on Only models you provide so nobody adds a personal key.
  • Already run LiteLLM? Connect it to OpenWork and keep its budgets.

Next steps

Frequently asked questions

Can I set a monthly AI budget per team?
Yes. In the AI Gateway, add a spend limit for a team with a monthly amount. You can also add daily or weekly amounts; whichever runs out first applies.
What happens when someone hits their limit?
You choose: pause their models, only warn, or let them ask for 25% more for an admin to approve.
Are the cost numbers exact?
They're close estimates from token counts and published prices. Your provider's invoice is the final number.

Know what AI costs before the invoice.

Spend allowances and usage reporting for organization-managed providers.

Free and open source. macOS, Windows, and Linux.