Skip to content
← Back to the blog

Fewer tools, better names: the hidden cost of giving an agent everything

Every tool you add gets paid for twice: in tokens and in hesitation. 50 available tools do not make an agent more capable, they make it more indecisive.

Category
Tooling
Date
Reading
3 min
Menos herramientas, mejores nombres: el costo oculto de darle todo al agente

There is a common reflex when building an agent: if I give it more tools, it will be able to do more. It feels intuitive and it is wrong.

Tool definitions live inside the context window, on every single call. A typical definition runs between 100 and 500 tokens. A 5-server setup with 58 tools eats roughly 55,000 tokens before the agent reads a single word from the user. Some popular integrations run 17,000 on their own.

But the token cost is the smaller of the two.

The cost that matters is hesitation

If the agent sees 50 tools, it has to read and rank 50 descriptions to decide which one to use. That inflates the prompt, slows the response, and worst of all, nudges it toward the broad tool when the narrow one would have been safer.

We saw this with a generic "save file" tool. It existed for legitimate cases, but the agent started reaching for it constantly, including for things that had their own validated path. It never did anything catastrophic. It simply took the wide road because the wide road was there.

The fix was not a better prompt. It was removing the generic tool from that agent's catalog.

How we design ours

The name does half the work. A tool called delegate_asset gets used differently than one called process_file. The name tells the model what kind of action this is before it reads the description.

The description says when NOT to use it. This is what helped us most. Explaining what a tool does is not enough; you have to say in which cases something else is a better fit. "Use this only if the file is not already in the account storage; if it is, use X instead."

Each agent gets its own catalog. The agent that talks to clients and the one that executes internal tasks do not share tools. They share infrastructure, not capabilities.

Tools are tested. We have a test that validates the schema of every tool and fails if one ends up with a wrong type or an empty description. That sounds obvious until an empty description ships to production and the agent starts picking blind.

Lazy loading

There is an elegant way out when the catalog genuinely has to be large: do not load every definition upfront, fetch them when needed. The published numbers are striking: one case where definitions consumed 134,000 tokens dropped to roughly 5,000 with lazy loading, an 85 % reduction.

That is the right direction as a catalog grows. But the other half deserves saying out loud: if your agent needs 50 tools, maybe you do not need lazy loading. Maybe you need 2 agents.

The question we ask before adding one

Before adding a new tool to a catalog, the question is not "would this be useful?" Almost everything would be useful. The question is: what concrete task is impossible today, and why does none of the existing tools cover it?

If there is no clear answer in one sentence, the tool does not get in.

Do you want to bring AI into your company?

We know how to do it. Tell us how things work today and we will tell you which part can run on its own, which needs an agent and which is better left to a person.