Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 17 min read

Lesson 21 - Latency and Cost Engineering

Code: agentic-course/agentic/trace.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Latency Engineering for LLM Apps covers the theory and the interview framing, without code.


What you will build


The idea

An agent is structurally worse than a single model call on both axes, and it is worth being precise about why, because the shape of the problem tells you which levers exist.

Latency is additive. A loop makes N model calls one after another, each waiting on the previous one’s tool result. You cannot overlap step 3 with step 2, because step 3’s prompt contains step 2’s output. Four steps at 1.5 seconds each is six seconds before the first useful word, and no infrastructure removes that dependency chain.

Input tokens compound. The loop resends the entire conversation every step: system prompt, task, every assistant turn, and every tool result so far. The context does not stay the same size, it grows monotonically, and the last call pays for everything before it.
πŸ’‘ Input tokens are the text you send and output tokens are what comes back, and providers bill the two at different rates.

The analogy that sticks: a single model call is one phone call. An agent is a chain of phone calls where, before each one, you read the entire previous transcript aloud. Adding a step costs one call plus a re-reading of everything to date. Both effects are properties of the loop rather than of the model, which is why β€œuse a faster model” is rarely your biggest lever.


Measure before you optimize

A real run through the library: three tools, one of which fetches a 2.7 KB policy document. Prices are assumed unit prices, not any provider’s real numbers.

tracer = Tracer()
run = Agent(
    FakeModel([
        tool_call("fetch_policy", {"topic": "refunds"}),
        tool_call("get_order", {"order_id": "4471"}),
        tool_call("check_delivery", {"order_id": "4471"}),
        "Order 4471 was delivered 12 days ago, inside the 30-day refund window.",
    ]),
    Registry([fetch_policy, get_order, check_delivery]),
    system="You are a support agent. Answer only from retrieved documents.",
    budget=Budget(max_steps=6, price_per_1k_input=0.5, price_per_1k_output=1.5),
    tracer=tracer,
).run("Can I still refund order 4471?")
stop answered  steps 4   tokens in/out 2181/20   cost 1.120500
executed ['fetch_policy', 'get_order', 'check_delivery']

run:agent.run  0ms
  model:model.call.0  0ms  tok 22/1
  tool:tool.fetch_policy  0ms
  model:model.call.1  0ms  tok 707/1
  tool:tool.get_order  0ms
  model:model.call.2  0ms  tok 723/1
  tool:tool.check_delivery  0ms
  model:model.call.3  0ms  tok 729/17

Read the numbers off it rather than guessing. 2181 input tokens against 20 output tokens - input is a hundred times the output, on the axis that is priced lower per token. total_cost is 1.1205, and of_kind("model") breaks it down as [0.0125, 0.355, 0.363, 0.39]: the first call is essentially free and the last three cost about the same as each other.

The per-call input column is where the diagnosis lives: 22, 707, 723, 729. One jump of 685 tokens, then two jumps of about 15. The 685 arrived immediately after tool.fetch_policy returned and was re-sent on every remaining call, so that one tool result is responsible for roughly 1900 of the run’s 2181 input tokens.

Durations read 0ms because FakeModel is a local scripted stand-in with no network in the way. Against a real provider the model spans carry their seconds and the tool spans carry theirs, and you subtract to see which side owns the wall clock. This is where teams reliably guess wrong. The instinct is to optimize the model call, because that is the part that feels expensive and has a vendor attached. Very often a tool is the real cost: a document fetch with no projection, a query that returns every column, a search tool that returns full pages instead of the passage that matched. The tracer settles the argument in about ten seconds, which is why lesson 9 put tokens and cost on every span rather than only on the run.


Why a fat tool result is charged many times

flowchart LR
    S1[Step 1 sends 22 input tokens] --> T1[fetch_policy returns 2.7 KB]
    T1 --> S2[Step 2 sends 707 input tokens]
    S2 --> T2[get_order returns a small record]
    T2 --> S3[Step 3 sends 723 input tokens]
    S3 --> T3[check_delivery returns one line]
    T3 --> S4[Step 4 sends 729 input tokens]
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class S1,S2,S3,S4 service
    class T1,T2,T3 data

Green is a model call carrying the whole transcript; yellow is a tool result appended to it and re-sent from then on. A 700-token result on step 1 of a four-step run costs roughly 2100 tokens, not 700.

The levers, in the order they actually pay

1. Cut steps

Each step is a whole round trip plus a bigger context, so removing one step removes two costs. Dropping check_delivery from the run above - its answer was already in the order record - takes it from 4 steps to 3 and from 2181 input tokens to 1755, with the fat document still in place. Concretely: merge tools that are always called together, let a tool return the joined result rather than making the model join it, and write a genuinely good description on each tool so the model does not spend a step discovering it picked the wrong one. A vague tool description is a latency bug - that is the argument lesson 3 makes about descriptions being prompt text.

2. Trim what tools return

The highest-leverage lever and the least discussed. Same run, same trajectory, same answer - the only change is fetch_policy returning the relevant sentence instead of the whole section:

before   input 2181   output 20   cost 1.120500   per-call input [22, 707, 723, 729]
after    input  165   output 20   cost 0.112500   per-call input [22,  35,  51,  57]

92% fewer input tokens and 90% less cost, with run.output byte-identical and run.trajectory() unchanged. Nothing about the model, the prompt, or the loop moved.

The rule: a tool returns what the model needs to decide, not everything it knows. Project the columns you use, and return the matched passage rather than the document. Cap list results and say how many were omitted. Where you can summarise deterministically on the tool side, do that instead of paying the model to re-read the whole thing every step.

3. A cheaper model for the early or easy steps

The seam makes this a one-line change, because Agent takes any object implementing complete. Early steps in a research-then-answer run are usually routing decisions - which tool, what query - and a smaller model does them fine; keep the expensive one for the final synthesis, where quality is visible. Route by task class too, so an easy case a small model answers correctly never reaches the big one. Measure it as a mixed-model pass rate on your eval suite rather than assuming, because a cheap model that costs you an extra step has made both axes worse.
πŸ’‘ An eval suite is a fixed set of test cases with agreed-correct answers, re-run on every change so a quality drop shows up as a number instead of a complaint.

4. Caching, then 5. parallel independent tool calls

Caching gets the section below, because its failure modes are specific. Parallelism ranks last: when the model requests several tools in one turn and none depends on another’s output, they can run concurrently and the step’s latency becomes the slowest tool rather than the sum. tool_calls(("a", {...}), ("b", {...})) scripts that shape, and the library dispatches sequentially because there is no async anywhere in this course, on purpose, so the control flow stays readable. It is last because it only helps the step that has independent calls, while cutting a step helps every step after it.

Exact-match response caching

Cache on a hash of everything that determines the output:

def _key(self, messages, tools, max_output_tokens, temperature) -> str:
    payload = {
        "messages": [m.to_dict() for m in messages],
        "tools": tools or [],
        "model": getattr(self.inner, "name", "unknown"),
        "max_output_tokens": max_output_tokens,
        "temperature": temperature,
        "prompt_version": self.prompt_version,
        "tenant": self.tenant,
    }
    return hashlib.sha256(
        json.dumps(payload, sort_keys=True, default=str).encode()
    ).hexdigest()

def complete(self, messages, tools=None, max_output_tokens=512, temperature=0.0):
    if temperature != 0.0:      # caching a sample caches one draw from a distribution
        return self.inner.complete(messages, tools, max_output_tokens, temperature)
    key = self._key(messages, tools, max_output_tokens, temperature)
    if key not in self.cache:
        self.cache[key] = self.inner.complete(
            messages, tools, max_output_tokens, temperature)
    return self.cache[key]

Every component is there because leaving it out is a bug. The tool schemas are in the key because the same prompt with a different tool set is a different question. prompt_version is in the key so shipping a new system prompt does not serve answers generated under the old one - see lesson 22 for the release identity this belongs to.

And tenant is in the key because of this - the same cache with the tenant left out, which is the most common version of the mistake:

leaky acme  : Refunds run 30 days.
leaky globex: Refunds run 30 days.        <- cache hit, wrong tenant

Globex asked its own question, hit acme’s entry, and was served acme’s answer. Nothing errored, nothing logged, and the two tenants have different refund policies. Forgetting a key component does not degrade the cache, it serves one tenant’s answer to another - so the key must include the caller’s permission scope and every identity that can change a correct answer. The general treatment is in caching; the agent-specific version is that your cache key and your authorization model have to agree.

Where caching pays most in an agent is not the final answer, which is usually unique. It is the repeated internal work: the same retrieval query on step 2 of a thousand runs, the same classification of a common intent, the same tool result for a stable record. Cache the tool call as well as the model call. Provider-side prompt caching on a long stable prefix is a complementary win, covered in prompt and semantic caching.
πŸ’‘ Provider-side prompt caching is the provider charging you less to re-read the opening chunk of a prompt it has seen before, which only works while that chunk stays unchanged.

Why semantic caching is specifically dangerous for agents

Semantic caching serves a stored answer when the new question is similar to an old one, by embedding distance. That is a reasonable trade for a FAQ and a bad one for an agent, for a reason beyond the usual β€œsimilar is not identical”. Similarity is not equivalence: β€œCan I refund order 4471” and β€œCan I refund order 4472” sit almost on top of each other in embedding space and have different correct answers.
πŸ’‘ Embedding distance is a single number for how close two pieces of text are in meaning, which is not at all the same as them meaning the same thing.

The agent-specific part is worse. A near-miss can change which tool the agent picks. A cached turn that is a tool request for get_order, served against a question that needed check_eligibility, does not return a slightly stale sentence. It sends the run down a different branch and executes the wrong tool, with arguments derived from the wrong question. The damage is an action, not a paragraph. If you use semantic caching at all, use it on leaf read-only steps with a tight threshold, never on a turn that can select a mutating tool, and put the hit rate and the false-hit rate in your eval suite.

Time to first token, and the honest version of a long run

Total time and time-to-first-token are different numbers with different owners. Streaming improves the second and does nothing to the first: a four-step run still takes four round trips, streaming only means the user watches step 4’s words arrive instead of waiting for it to finish.
πŸ’‘ Time to first token is how long the user stares at nothing before the first word appears, which is a separate number from how long the whole answer takes. So for an agent the useful thing to stream is not tokens, it is progress. A run that shows nothing for eight seconds reads as broken, and the user reloads - which starts a second run, doubling cost and producing a duplicate action if any tool mutated something. The library gives you the hook directly:

run = Agent(model, registry, on_step=lambda s: notify(s.index, s.succeeded_names)).run(task)

on_step fires once per completed step with the tools that ran, which is enough to render β€œlooking up your order” and then β€œchecking the refund policy”. People tolerate a slow operation they can see advancing far better than a fast one that looks dead. Which latency number to promise is covered in latency engineering, and the percentile vocabulary in performance metrics.

Set budgets from measured percentiles

You have the instrument now, so stop guessing the ceilings. Run your eval suite and collect step_count, cost and duration per case. Then set the budget from the distribution of the cases that passed: max_steps a little above the p95 step count of passing runs, and max_cost from their p99. Take the user-facing timeout from measured p95 total time rather than from what feels acceptable. A budget picked from a distribution is a statement about what your system does; a budget picked from a round number is a statement about what you hoped. The ceilings themselves are lesson 8, and max_cost is a check in the eval suite because a right answer that cost ten times its budget did not succeed.

Exercise

Instrument a two-tool run, find the largest contributor from the trace, cut it, and show the measured before-and-after on tokens and cost with the answer held constant. Save as tests/test_exercise_cost.py inside agentic-course/. Success criterion: python3 -m unittest tests.test_exercise_cost -v reports OK, having asserted which span the largest contributor followed, that the answer and trajectory are unchanged, and that input tokens and cost both fell by more than 5x.

Worked solution ```python import unittest from agentic import Agent, Budget, FakeModel, Registry, Stop, Tracer, tool, tool_call FULL = ("SECTION 1. Refunds are accepted within 30 days of delivery. " "Items must be unused and in their original packaging. ") * 24 SLIM = "Refunds: within 30 days of delivery, items unused." PRICES = dict(price_per_1k_input=0.5, price_per_1k_output=1.5) # assumed unit prices SCRIPT = [ tool_call("fetch_policy", {"topic": "refunds"}), tool_call("get_order", {"order_id": "4471"}), "Order 4471 is 12 days old, inside the 30-day window.", ] def traced_run(policy_text: str): @tool(description="Fetch the refund policy section.", topic="A policy topic.") def fetch_policy(topic: str) -> str: return policy_text @tool(description="Look up an order by id.", order_id="The order id.") def get_order(order_id: str) -> dict: return {"id": order_id, "status": "delivered", "delivered_days_ago": 12} tracer = Tracer() run = Agent(FakeModel(list(SCRIPT)), Registry([fetch_policy, get_order]), system="You are a support agent.", tracer=tracer, budget=Budget(max_steps=5, **PRICES), ).run("Can I still refund order 4471?") return run, tracer def input_token_jumps(tracer): """Growth in prompt size between consecutive model calls.""" tok = [s.attributes.get("input_tokens", 0) for s in tracer.of_kind("model")] return [tok[i + 1] - tok[i] for i in range(len(tok) - 1)] class TestTrimTheToolResult(unittest.TestCase): def test_trimming_the_worst_span_cuts_cost_and_keeps_the_answer(self): before_run, before = traced_run(FULL) jumps = input_token_jumps(before) # [685, 16] worst = jumps.index(max(jumps)) # 0 self.assertEqual(before.of_kind("tool")[worst].name, "tool.fetch_policy") after_run, after = traced_run(SLIM) self.assertEqual(before_run.output, after_run.output) self.assertIs(after_run.stop, Stop.ANSWERED) self.assertEqual(before_run.trajectory(), after_run.trajectory()) self.assertLess(after.total_input_tokens, before.total_input_tokens / 5) self.assertLess(after.total_cost, before.total_cost / 5) # measured: input tokens 1425 -> 81, cost 0.7350 -> 0.0630 ``` The load-bearing trick is `input_token_jumps`. A tool span records its arguments and its duration, not the size of what it returned, so you cannot read a tool result's cost off the tool span. You read it off the **growth in the next model call's input tokens** - the right number anyway, because that growth is what you pay on every remaining step. The assertions on `output` and `trajectory()` keep the exercise honest. A cost optimization that changes the answer is a regression with a nice number attached.

Checkpoint

Why is an agent structurally worse than a single model call on both latency and cost?

Latency is additive because step N’s prompt contains step N-1’s output, so the calls cannot overlap. Cost compounds because the loop resends the whole transcript every step, so the context grows monotonically and the last call pays for everything before it.

A run costs more than you expected. Which trace column do you read first?

The per-model-call input_tokens series. A single large jump tells you which tool result entered the transcript and is now being re-sent every step. Teams tend to blame the model call; the tracer usually points at a tool.

Why is trimming a tool result higher-leverage than it looks, and what breaks if you leave the tenant out of a cache key?

A tool result is charged on every subsequent model call, not once - returning a sentence instead of a section took the measured run from 2181 input tokens to 165 with a byte-identical answer. And a cache missing a key component does not degrade, it silently serves one tenant’s answer to another, with no error and no log line.

Why is semantic caching worse for an agent than for a FAQ?

Because a cached turn can be a tool request. A near-miss does not return a slightly stale sentence, it sends the run down a different branch and executes the wrong tool with arguments derived from the wrong question.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access