Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 24 min read

Tool Calling Fundamentals - Complete Deep Dive

Stage 5 - Agents Lesson 21 of 36

Prerequisites: LLM APIs and SDKs, Structured Outputs, Tracing and Observability for LLM Apps Used in: Agent Architectures, Model Context Protocol, Prompt Injection and the OWASP LLM Top 10 Build it: Lesson 3 - Tools - Schemas, Validation, Dispatch implements this as runnable, tested code you can execute offline.


What is Tool Calling?

Tool calling is the mechanism that lets a model request an action in the outside world. You describe the functions available, the model replies with a structured request naming one and supplying arguments, your code executes it, you append the result to the conversation and call the model again so it can use what came back.

Read that sequence again and notice the part everyone skips: the model never executes anything. It emits a JSON object that says โ€œI would like get_order_status called with order_id equal to this.โ€ That object is a request, not an invocation. Nothing happens until your executor decides to run it. That boundary is not an implementation detail โ€” it is the entire security architecture of the feature, and every control worth having lives on your side of it.

Real-world analogy: a requisition form. A new employee needs a part from the warehouse. They do not walk into the warehouse; they fill out a form naming the part and the quantity. A clerk checks the form is filled in correctly, checks the employee is entitled to that part, fetches it, and hands it over with a note if something was wrong. The employee can ask for anything at all. What they actually get is whatever the clerk approves. The model is the employee, your executor is the clerk, and a system where the employee has warehouse keys is not a different design โ€” it is a missing clerk.


The Mechanic - Two Calls, Not One

flowchart LR
    A[Client request] --> B[Assemble prompt and tool schemas]
    B --> C[Model call one]
    C --> D[Model emits a structured tool call request]
    D --> E[Gate one - validate against the schema]
    E --> F[Gate two - authorise arguments against the caller permissions]
    F --> G[Gate three - confirm if irreversible]
    G --> H[Your executor runs the tool]
    H --> I[Append the tool result to the message list]
    I --> J[Model call two]
    J --> K[Another tool call or a final answer]
    K -->|another tool call| E
    K -->|final answer| L[Respond to the client]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A,G client
    class B,E,F,H service
    class C,D,J,K async
    class I data
    class L edge

Three consequences fall out of the shape. Every tool use costs at least two model calls, and the message list grows each round including the full tool result. So a five-step loop costs far more than five single calls, because each one re-sends everything before it โ€” see Tokens, Context Windows, and Cost Math. The loop has no natural terminus, so it runs until you stop it, and a confused model will keep requesting tools indefinitely. And the tool result becomes untrusted input to the model, because whatever a tool returns is read as language, which makes any tool that fetches third-party content an injection vector by design โ€” the indirect path described in AI Security.
๐Ÿ’ก Whatever a tool hands back is read as language, so a sentence hidden in a fetched web page arrives looking exactly like an instruction from you.


Tool Schema Design Is the Highest-Leverage Work

Most teams treat tool definitions as plumbing and then blame the model for choosing badly. The definitions are prompt text. The model reads your names, descriptions, and parameter documentation and decides from them alone. A vague description is the single largest cause of wrong tool choice, ahead of model capability.
๐Ÿ’ก A tool schema is the name, the written description, and the typed argument list - all of it goes into the prompt as plain text the model reads before choosing.

{
  "name": "get_order_status",
  "description": (
    "Look up the current shipping status of one order belonging to the signed-in "
    "customer. Use when the customer asks where an order is, whether it shipped, "
    "or when it will arrive. Requires an order id obtained from search_orders - "
    "never guess an order id. Do not use this for refunds or cancellations."
  ),
  "parameters": {
    "type": "object",
    "properties": {
      "order_id": {"type": "string", "description": "Order id from search_orders"},
      "detail": {"type": "string", "enum": ["summary", "full_timeline"]}
    },
    "required": ["order_id"]
  }
}

What that description does that a one-liner cannot: it says when to use the tool, it says when not to, it names the tool that produces the required id, and it forbids inventing one. Four sentences of prompt engineering doing work people try to do with model upgrades.

The rest of the design rules:

Smell to symptom to fix

Smell in the tool definition Symptom you observe Fix
Description is one vague phrase like fetches data Tool never called - or called for unrelated questions Rewrite as when-to-use plus when-not-to-use guidance
Two tools with near-identical descriptions Model alternates between them across runs Merge them - or state the boundary explicitly in both
Twenty-plus tools always in scope Wrong-tool rate climbs and prompt cost rises Scope the list per task - group behind fewer higher-level tools
Free-form string where the value set is closed Invalid values that fail deep inside the tool Replace with an enum
Deeply nested object parameters Missing required fields and malformed arguments Flatten to typed top-level parameters
An id parameter with no stated source Model invents plausible ids that do not exist Document the producing tool - validate existence and ownership
One tool with a mode flag doing several jobs Correct tool chosen with the wrong mode Split into separate named tools
Tool returns a large unbounded payload Context bloat and rising cost each round Summarise - paginate - limit returned fields
Errors return an opaque stack trace Model retries the identical failing call Return an actionable message naming what to change
Tool name mirrors an internal service Model cannot map the question to the tool Rename for the user-facing action

Never Trust Model-Supplied Arguments

The arguments are generated text, influenced by the userโ€™s message, by retrieved documents, and by anything a tool returned earlier in the loop. So the rule is blunt: an argument that came from a model is user input, with exactly the trust level that phrase implies.

Two separate checks, in order, and they are not the same thing. Validation asks whether the arguments are well-formed โ€” present, correctly typed, in range, enum members, matching the expected id format. That is ordinary schema work, the same machinery as Structured Outputs, and it runs before anything executes. Authorization asks whether the real caller is permitted to have this action performed on this object. That is the check teams forget, because the arguments validated cleanly and looked reasonable.

def execute(tool_name, args, principal):          # principal from the verified session
    spec = TOOLS[tool_name]                       # unknown tool name is a hard stop
    payload = spec.arg_model.model_validate(args) # gate one - well formed
    order = orders.get(payload.order_id)          # gate two - authorise the real caller
    if order is None or order.customer_id != principal.customer_id:
        audit.deny(principal, tool_name, payload)
        return tool_error("no order with that id - call search_orders first")
    return spec.run(payload, principal)           # credential attached server side

Three details there are load-bearing. The principal comes from the authenticated session and never from the model โ€” if the model supplies a user id, you have handed it the ability to choose whose data to read. The failure message does not distinguish โ€œnot yoursโ€ from โ€œdoes not exist,โ€ because that difference is an enumeration oracle. And the credential is attached inside spec.run, server-side, so no token is ever in the context window.
๐Ÿ’ก An enumeration oracle is an error message that reads differently for โ€œdoes not existโ€ than for โ€œnot yoursโ€, which lets someone map what exists just by guessing ids.

The general principle: never let the model pick whose record to read. Scope the query by the verified caller and let the model choose only among things that caller may already touch. Identity mechanics in Authentication.


Read-Only and Mutating Tools Are Not the Same Class

ย  Read-only tools Mutating tools
Cost of a wrong call A wasted call and some context A real side effect in a real system
Retry safety Safe to retry freely Safe only if the tool is idempotent
Automatic execution Fine Fine for low-value and reversible - gated otherwise
Human confirmation Not needed Required for anything irreversible or high-value
Parallelism Safe to fan out Serialise unless provably independent
Blast radius of an injection Disclosure bounded by permissions Action taken on an attackerโ€™s behalf

Anything irreversible wants a confirmation step: money movement, deletion, external sends, permission changes, anything a customer would notice. The confirmation must display the resolved arguments โ€” the actual recipient, amount, and record โ€” never a model-written summary of what it is about to do, because that summary is text an attacker may have influenced.


Parallel and Sequential Calls

Providers increasingly let a model emit several tool calls in one turn. Fan them out when they are genuinely independent โ€” three lookups against different systems collapse three round trips into one, usually the largest latency win available in a tool-using loop. Give each parallel call its own timeout plus an overall deadline, or the slowest one sets your latency, and serialise parallel mutations by default, since the model has no view of your consistency requirements.

Sequential is required whenever one callโ€™s output is anotherโ€™s input, which the model cannot always tell. The classic failure is emitting get_order_status in parallel with the search_orders call that was supposed to produce the order id, then hallucinating the id. The defence is in the description (โ€œrequires an order id obtained from search_ordersโ€) plus validation that rejects an id which does not exist.


Error Handling - Recover or Fail Hard

A tool error is an opportunity, not just a failure. The model is in a loop and can correct itself if you tell it what went wrong in language it can act on. But some errors must not be handed back.

Situation Return to the model Why
Missing or malformed argument Yes - name the field and the expected form The model can fix it on the next turn
Value not found Yes - suggest the lookup tool that produces valid values Turns a dead end into a recovery
Enum value invalid Yes - list the permitted values Cheap and almost always corrected
Transient dependency failure Retry inside the tool with backoff - surface only on exhaustion The model should not manage your retry policy - see Retry and Backoff
Rate limited Retry inside the tool - then a plain wait-and-retry message Model-driven retry storms are a real cost incident
Authorization denied A neutral not-found message with no detail Detail becomes an enumeration oracle
Dependency down and the breaker is open Fail the turn - tell the user Looping against an open circuit breaker wastes tokens and time
Unknown tool name Fail hard and alert Indicates schema drift or an injection attempt
Same failing call repeated Break the loop - escalate The model is stuck and will not self-correct

The shape that works: recoverable errors are data, unrecoverable errors are exceptions. A recoverable error returns a normal tool result whose content is an actionable message. An unrecoverable one ends the turn with a plain explanation to the user.


Idempotency and Side Effects

The uncomfortable property of the loop: a retried model call can legitimately re-request the same action. You retry because the response was truncated or was invalid JSON; the model, seeing the same state, asks to send the same email again. Nothing malfunctioned. The architecture simply gave you an at-least-once delivery channel for side effects.

So the idempotency guarantee belongs in the tool, not in the loop โ€” not in orchestration state that a crash or a retry will lose.

Key design and storage in Idempotency.


Timeouts and Per-Tool Latency Budgets

A turnโ€™s latency is the sum of every model call plus every tool call, and the user waits through all of it. Give each tool a budget derived from the turnโ€™s budget rather than one global timeout copied from your HTTP defaults โ€” a five-second database lookup is broken, a five-second document conversion is normal. Then decide per tool what a timeout means: sometimes a recoverable error the model can route around, sometimes a hard failure of the turn because proceeding without that result yields a confidently wrong answer. Write the decision down instead of letting it default to whatever the exception handler does, and put a circuit breaker on each tool so a dead dependency fails fast rather than consuming the whole budget in timeouts.


Terminating the Loop

The loop has no natural end, so termination is your job. Use several limits together, because each catches a different pathology.

Every termination path needs a defined user-visible outcome. Silently returning nothing after hitting a limit is worse than saying the request could not be completed. Each should also emit a metric โ€” loops that hit limits are a leading quality signal.


Observability and Evaluating Tool Use

Every tool call gets its own span, as a child of the model call that requested it: tool name, arguments as sent, result or error, duration, retry count, idempotency key, and the authorization decision. Without per-call spans you cannot answer โ€œwhich model call caused this write,โ€ and you cannot tell a flaky integration from a model calling it wrongly. Mechanics in Tracing and Observability for LLM Apps.

Then score tool use as its own eval dimension, because end-to-end answer quality hides tool errors the model talked its way around. Three questions, three metrics. Right tool? โ€” including โ€œno toolโ€ when the question needed none, scored against labelled inputs with expected selections. Correct arguments? โ€” well-formed, complete, and drawn from the right source rather than invented. Stopped when it should? โ€” one call and done versus eleven steps for a one-step question, and stopping cleanly when the answer is genuinely unavailable.

Add the failure classes as permanent cases: a question needing no tool, two tools that look similar, a required id that must come from a lookup first, a tool that returns an error, and an injected instruction inside a tool result. This slots into the golden set described in Evals.
๐Ÿ’ก A golden set is a fixed list of inputs paired with the answers you have agreed are correct, re-run on every change so a regression shows up as a number.


Bad to Good to Great

Bad - dispatch on the modelโ€™s output

Parse the tool name and arguments and call the function โ€” getattr on a module, or a dictionary lookup, then straight through with whatever came back.

An unknown tool name crashes or, worse, reaches something you never meant to expose. Arguments are unvalidated, so a malformed one fails deep inside a dependency and you hand the stack trace back to the model. Nothing authorises the object being touched, so an id from the model selects any record in the table โ€” a data-access bug with a language model in the middle of it. There is no loop bound, so a confused model spends your budget, and no idempotency, so a retry sends the email twice.

Good - validate arguments and cap the loop

Typed argument models per tool, registry-based dispatch that rejects unknown names, a maximum iteration count, and errors returned to the model as text.

This is a real system and it stops the loud failures. What it does not stop is the quiet one: validation proves the arguments are well-formed, not that the caller is entitled, so an order_id of the correct shape belonging to a different customer passes every check. Retries still duplicate side effects, because an iteration cap bounds count and not cost and nothing distinguishes a repeated identical call from progress. And mutating tools run on the same path as read-only ones, so a wrong choice among them is a wrong write.

Great - validated, authorised, idempotent, bounded, traced

  1. Tool list scoped to the task, with descriptions written as when-to-use and when-not-to-use guidance, flat typed parameters, enums on closed sets, and no two overlapping definitions.
  2. Validation then authorization as separate gates, the principal taken from the verified session and never from the model, denials returned as a neutral not-found and audited.
  3. Idempotency inside the tool, keyed by caller plus tool plus canonicalised arguments plus step index, with the key generated by the executor.
  4. Mutating tools on a separate path from read-only ones, with human confirmation on anything irreversible, displaying resolved arguments rather than a model-written summary.
  5. Per-tool timeouts, retries, and circuit breakers inside the tool, plus a written decision per tool about whether a timeout is recoverable or fatal.
  6. Layered termination โ€” iteration cap, token and spend budget, wall-clock deadline, repeated-call detection, no-progress detection โ€” each with a user-visible outcome and a metric.
  7. A span per tool call carrying arguments, result, duration, retries, and the authorization decision.
  8. Tool use scored in the eval suite on selection, argument correctness, and stopping behaviour, with injection-inside-a-tool-result as a permanent case.

The difference between Good and Great is not validation depth. It is that in Great, a wrong tool call by a persuaded model is a contained, logged, non-duplicated, bounded event.


When to Use

โœ… Use tool calling when:

โŒ Do not use tool calling when:


Common Interview Questions

Q1: Walk me through what actually happens in a tool call.

You send the prompt along with schemas describing the available tools. The model replies not with prose but with a structured request naming a tool and supplying arguments. Your code validates those arguments and authorises them against the real callerโ€™s permissions. Then it executes the tool, appends the result to the message list, and calls the model again so it can use what came back - which may produce another tool call or a final answer. The critical point is that the model never executes anything; it emits a request, and your executor decides whether to honour it. That boundary is the security architecture of the whole feature, which is why every control - validation, authorization, idempotency, confirmation, egress restriction - lives on your side of it. It also means every tool use costs at least two model calls, with the message list growing each round.

Q2: Your model keeps picking the wrong tool. How do you debug it?

I start at the tool definitions, because they are prompt text and a vague description is the most common cause, ahead of model capability. First I check for overlapping descriptions - if a human reading two definitions cannot state the boundary in one sentence, the model is effectively guessing, so I merge them or write the boundary into both. Then I rewrite descriptions as explicit when-to-use plus when-not-to-use guidance, naming which tool produces any required id. Then I look at list size, since selection accuracy degrades as tools accumulate, and scope the list to the task instead of exposing a global registry. Structural fixes follow: split any tool with a mode flag into distinct named tools, replace free-form strings with enums, flatten nested parameters, and rename anything named after an internal service. And I score tool selection as its own eval dimension against labelled inputs, including cases where the right answer is no tool at all, so I can tell whether a change actually helped.

Q3: The model returns an account id and asks you to read that account. What do you do?

Treat the id as user input, because that is exactly what it is - generated text influenced by the userโ€™s message, retrieved documents, and earlier tool results. Two distinct checks in order. Validation confirms it is well-formed and of the right type and shape. Authorization confirms the verified caller from the session is permitted to read that specific account, which is the check that gets skipped because the argument validated cleanly and looked plausible. The principal comes from the authenticated session and never from the model - if the model supplies the user id, I have handed it the ability to choose whose data to read. On denial I return a neutral not-found rather than distinguishing not-yours from does-not-exist, since that distinction is an enumeration oracle, and I audit the denial. The toolโ€™s credential is attached server-side inside the executor, so no token is ever in the context window.

Q4: A retry causes the same email to be sent twice. Where does the fix belong?

In the tool, not in the loop. The retry is legitimate - you retry a model call because the response was truncated or unparseable, and the model, seeing the same state, asks to send the same email again. Nothing malfunctioned; the architecture simply gives you at-least-once delivery for side effects. So the executor generates an idempotency key from the caller, the tool name, and the canonicalised arguments, plus a step index so a deliberate second send is distinguishable from a duplicate of the first. The tool looks that key up and returns the prior result unchanged on a repeat instead of acting again. The model never supplies the key, because it will produce a fresh one for a duplicate request. State transitions and deletes are made conditional on current state so a replay is a no-op. Remembering what already ran in orchestration state does not work, since a crash or a retry loses it.

Q5: How do you stop an agent looping forever?

Several limits together, because each catches a different pathology. A hard cap on model calls per turn bounds the obvious runaway. A token and spend budget across the turn bounds cost, which iteration count does not, since one step can be enormous. A wall-clock deadline checked before each call bounds user-visible latency. Repeated-call detection - hashing tool name plus canonicalised arguments - catches the most common real case, a model re-requesting the same failing lookup, and that should break the loop rather than keep paying for it. No-progress detection catches turns of text with neither a tool call nor an answer. Each path needs a defined user-visible outcome, because silently returning nothing is worse than saying the request could not be completed, and each should emit a metric, since loops hitting limits are a leading quality signal. Underneath all of it, per-tool timeouts and circuit breakers stop a dead dependency from consuming the whole budget in timeouts.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access