The loop, the tools, and the moment one breaks
Agentic AI is what you get when a language model is given a way to act rather than only to answer. A runtime hands the model a catalogue of permitted actions: run this query, fetch that page, write a file, call an internal service. The model replies with a request to use one of them. The runtime, not the model, carries the action out and feeds the result back in. Then the cycle repeats. Notice what the model itself never touches. It opens no connection, holds no credential and sees no file system; it emits structured text naming a tool and its arguments, and ordinary code decides whether to honour that request. Everything a demo presents as planning or initiative is that cycle running several times, with each result shaping the next request. Which is why the engineering that decides whether an agent works sits outside the model entirely.
- The tool surface is the real design work. Every action needs a name, an argument list and a description precise enough that a model under pressure picks the right one.
- Return values matter more than prompts. A tool that answers an impossible request with a blank string has just taught the model that the request succeeded.
- Permissions and budgets set the size of the accident. Reading a production database and writing to it are different products with different approval paths.
- Stop conditions have to be explicit: a cap on steps, a cap on spend, and a rule for what a repeated identical action means.
- Traces, meaning a readable record of every call and everything it returned, are the only way to debug a run. Without them you are reading tea leaves.
Failure is where this category earns or loses its money. One wrong answer from a chat assistant costs the reader a minute; one wrong step inside a loop becomes the input to the step after it. The model has no way to distinguish a tool that returned nothing because the correct answer is nothing from a tool that returned nothing because a credential expired overnight. So the useful work is defensive and dull. Make every tool describe its own failures in words the model can act on. Require a human signature for anything destructive. Halt the run when the same call repeats. Log enough that somebody can reconstruct afterwards what happened and why. Do that and agents pay off on bounded work with a verifiable result, where a person can look at the output and say whether it is right. Skip it, or apply agents to open-ended work with no such check, and what you have built is an expensive generator of plausible activity.
