Cost of agents
Most of what agents cost, they waste.
Before you attribute the bill, it is worth making it smaller. Nearly all avoidable agent spend comes from three places, and all three are governance settings rather than model choices.

Where it goes
Three ways to pay for nothing.
None of these show up as a line item. They show up as a total that is larger than the work would suggest.
Choosing between a hundred tools
Every tool an agent can see costs tokens to read and attention to discard. Most of a bloated context is capability that was never going to be used.
Guessing, then retrying
An agent that picks the wrong tool does not stop. It tries again, and each attempt bills exactly like a successful one.
Doing it the long way
Without a playbook, a model reconstructs your process from first principles on every run — and takes a more expensive route than the one your team would take.
What to change
Four settings, in the order that matters.
Each one reduces the work the model has to do. The third is a hard stop rather than an improvement — it is what stops a bad run becoming an expensive one.
-
Narrow
Show fewer tools
Spaces expose only what an identity should see. A smaller surface is a shorter context and a smaller chance of choosing wrong.
-
Instruct
Carry the process in a Skill
The sequence and the judgement calls come from a Skill rather than from the model improvising, so fewer steps are spent working out what to do.
-
Bound
Cap the retry loop
Rate limits stop a failing agent from billing indefinitely, at the gateway, before anything downstream notices.
-
See
Find the expensive patterns
Usage per space, per agent and per tool shows which work is costly and which is simply repeated — the first is a decision, the second is a bug.
The same task
One run, two ways of paying for it.
The work is identical. The difference is how much the model has to figure out before it can start.
Everything visible, nothing written down
- Seventy tools in context, most irrelevant to the task
- A wrong tool chosen, then retried
- The process reconstructed from scratch on every run
- Failed attempts billed exactly like successful ones
- Spend visible only as a monthly total
Narrowed and instructed
- Only the tools this identity should see
- The sequence carried by a Skill, not inferred
- Retry loops capped at the gateway
- Failures counted alongside successes
- Usage broken down per space, agent and tool
What you get back
A number that means something.
Cutting waste changes the total. Measuring it properly changes what you do next.
Cost per completed task, not per call
A call is not a unit of value. What matters is what a finished piece of work costs, including everything that failed on the way to it.
Failures counted, not hidden
Denied, rate-limited and errored calls appear alongside successes, so the real denominator is visible.
Attribution when you need it
Once waste is under control, the question becomes who should pay for the rest. That is what the FinOps work is for.
Book a demo
See what governed AI agents look like.
A 20-minute demo on your stack. We'll show Palma working with the agents, tools and identity provider you already run.
- Enterprise security
- Role-based access
- Instant integration
Latest Blog Posts
Stay up-to-date with the latest in enterprise AI, MCP servers, and secure integration strategies.