Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 14 min read

Lesson 7 - Failure Handling - Retries, Timeouts, Idempotency

Code: agentic-course/agentic/model.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: LLM APIs and SDKs covers the theory and the interview framing, without code.


What you will build


The idea

A kitchen sends an order ticket to the grill and the line goes dead before the acknowledgement comes back. The order may have been cooked. It may have been cooked and billed. Or the ticket may never have arrived. Re-sending is free if nothing happened, and expensive if the steak is already on the pass.

That is the whole shape of failure handling for agents. Retries are cheap for reads, dangerous for writes, and the dangerous case is not β€œit failed” - it is β€œyou do not know”.

The part specific to agents: model endpoints fail transiently at rates that would be treated as an incident for an internal service, and you are calling them several times per user request because a loop makes N calls where a chat makes one. So retries are ordinary plumbing here, not paranoia, and the failure probability compounds across steps.


The real retry loop

HTTPModel.complete is the whole thing, and it is short enough to read:

for attempt in range(self.retry.attempts):
    req = urllib.request.Request(self.endpoint, data=body, headers={...}, method="POST")
    try:
        with urllib.request.urlopen(req, timeout=self.timeout) as resp:
            return self._parse(json.loads(resp.read()))
    except urllib.error.HTTPError as exc:
        # 4xx other than 429 will fail identically forever - do not retry.
        if exc.code != 429 and 400 <= exc.code < 500:
            raise
        last_error = exc
    except (urllib.error.URLError, TimeoutError) as exc:
        last_error = exc

    if attempt < self.retry.attempts - 1:
        # Full jitter. Lesson 7 explains why the jitter is not optional.
        import random
        time.sleep(random.uniform(0, min(delay, self.retry.max_delay)))
        delay *= 2

raise RuntimeError(f"model call failed after {self.retry.attempts} attempts") from last_error

The policy is a three-field dataclass, so it is configuration rather than scattered constants:

from agentic.model import RetryPolicy

RetryPolicy(attempts=3, base_delay=0.5, max_delay=8.0)   # the shipped defaults

Retry only what can succeed differently

The one branch worth memorising is the 4xx check. A 400 means your request body is malformed, a 401 means your key is wrong, a 404 means the path does not exist. Retrying any of them re-sends a request that will fail identically forever - you pay three times the latency to reach the same exception, and your dashboard shows three failures instead of one.

429 is the exception, because it is not β€œyou are wrong”, it is β€œnot now”. Retry 429, retry 5xx, retry connection errors and timeouts. Fail fast on the rest.

Jitter is not optional

Exponential backoff spreads retries over time. Jitter spreads them over callers. Without it, a fan-out of twenty parallel tool calls that all hit one rate limit will all sleep 0.5s, all wake together, all retry together, and rebuild exactly the burst the backoff was meant to disperse. The queue does not drain, it oscillates.

random.uniform(0, ceiling) - full jitter - breaks the correlation. Each caller waits a different amount, the herd smears out, and the downstream sees load it can actually absorb. This is the same mechanism as retry with backoff anywhere else in distributed systems; nothing about a model endpoint makes it special.

Honour the provider’s retry_after over your own arithmetic

Your computed delay is a guess. A Retry-After header or a retry_after field in an error body is the provider telling you when the limit resets. The shipped loop computes its own delay, so this is the change to make when you write a real adapter:

import random
from agentic.model import RetryPolicy

def delay_for(attempt: int, policy: RetryPolicy, retry_after: float | None = None) -> float:
    if retry_after is not None:
        return retry_after                       # the provider knows, you do not
    ceiling = min(policy.base_delay * (2 ** attempt), policy.max_delay)
    return random.uniform(0, ceiling)            # full jitter

Ignoring retry_after and retrying sooner is how an account gets throttled harder, or blocked.


The timeout case that should worry you

A connection error is honest: nothing happened. A timeout is ambiguous, and ambiguity is the expensive kind.

Your client gave up at 30 seconds. The provider may have finished generating at 31 seconds, billed you for the tokens, and written the response to a socket nobody was reading. From where you stand the two cases look identical, and a naive retry means:

  1. You pay twice for one logical request.
  2. You get a different answer. Generation is non-deterministic, and even at temperature=0.0 you are not guaranteed the same tokens across calls. So the retry is not β€œthe same request again” - it is a second, independent attempt at the same question.

That second point is what makes model retries unlike retrying a database read. There is no stable value to re-fetch. If some part of your system already acted on the first response - a logged decision, a queued message, a partially streamed answer - the retry can contradict it.

Raising the timeout does not fix this; it moves the boundary. What fixes it is making the request identifiable.


Idempotency: make the retry cost nothing

The pattern is four steps and it predates agents by decades:

  1. Derive a key from the inputs. Same inputs, same key - a hash of the prompt, model, and parameters.
  2. Persist the key before you call, with a status of in-flight.
  3. On retry, check the key. Completed means return the stored result. In-flight means wait or fail, do not fire again.
  4. Store the result under the key when it lands.

Idempotency is the general treatment. The agent-specific rule is sharper:

If a tool has side effects, the idempotency guarantee belongs in the tool.

Not in the retry loop, because the loop is not the only thing that repeats a request. A model can legitimately ask for the same action twice - it forgot the result was already in the transcript, the tool result was ambiguous, or a step got retried upstream. A retry wrapper around the HTTP call does nothing about that, because the second charge is a new, valid request.

CHARGES: dict[str, str] = {}

@tool(description="Charge a card once for an order.", mutating=True,
      order_id="The order id.", amount_paise="Amount in paise.")
def charge(order_id: str, amount_paise: int) -> str:
    key = f"{order_id}:{amount_paise}"
    if key in CHARGES:
        return f"already charged {CHARGES[key]}"      # idempotent, and says so
    CHARGES[key] = f"txn-{len(CHARGES) + 1}"
    return f"charged {CHARGES[key]}"

Point a confused model at it twice and the money moves once:

model = FakeModel([
    tool_call("charge", {"order_id": "4471", "amount_paise": 249900}),
    tool_call("charge", {"order_id": "4471", "amount_paise": 249900}),
    "Payment complete.",
])
run = Agent(model, Registry([charge])).run("pay 4471")

run.trajectory()   # ['charge', 'charge']  - two calls executed
CHARGES            # {'4471:249900': 'txn-1'}  - one charge

Both calls succeeded, which is the right outcome: the second one truthfully reports that the work is already done, so the model is not misled into trying a third time. And note that the second identical call was tolerated rather than blocked - the loop’s repeat_limit allows a bounded number of repeats before calling the agent stuck, which is lesson 8.


An exception taxonomy that maps to behaviour

tools.py defines three exception types. The distinction is not tidiness - each one produces a different, deliberate outcome in the loop:

Exception Who sees it What happens Use for
ToolError the model returned as a tool result, loop continues bad arguments, not found, validation, transient tool failure
FatalToolError nobody run ends with Stop.TOOL_FATAL authorization denial, missing credentials, open circuit
ToolConfigError you, as a crash propagates out of run() a registry wired up wrong

ToolError is the one that makes agents feel intelligent. The error text becomes a tool message, the model reads it, and it corrects - test_tool_error_is_returned_to_the_model_which_can_retry asks for order 9999, gets no order '9999', asks for 4471, and answers.

FatalToolError is deliberately hidden from the model. test_fatal_tool_error_aborts_instead_of_looping asserts model.call_count == 1: the model never learns that authorization failed, so it cannot go looking for a way around it.

An unexpected exception is a fourth case, handled at the dispatch boundary. It becomes Error: <tool> failed internally (RuntimeError) - no message, no stack trace - because test_internal_exception_does_not_leak_details_to_the_model asserts a secret in the exception text never reaches the transcript.


When to stop retrying and open a circuit

Retries assume the failure is transient. Once a dependency is down, every retry adds load to something already struggling and latency to a request that will fail anyway. A circuit breaker counts failures, opens after a threshold, fails fast while open, and lets a trial request through after a cooldown.

In an agent, an open circuit is not a ToolError - there is nothing for the model to correct, and letting it keep trying just burns steps:

BREAKER = {"open": True}      # your real breaker tracks failures and a cooldown

@tool(description="Call the pricing service.", sku="The sku to price.")
def price(sku: str) -> str:
    if BREAKER["open"]:
        raise FatalToolError("pricing circuit is open; stopping instead of retrying")
    return f"{sku}: 4999"

run = Agent(FakeModel([tool_call("price", {"sku": "abc"}), "unused"]),
            Registry([price])).run("price abc")
run.stop          # Stop.TOOL_FATAL

The run ends immediately with a named reason you can alert on, instead of three retries times four steps of paid latency ending in a vague apology.


Exercise

Write a tool that fails on its first call and succeeds on its second, then prove the agent self-corrects because the ToolError text reached the model.

Success criterion: run.stop is Stop.ANSWERED, requested_tools() has two entries, trajectory() has one, and the error text appears in run.messages.

python3 -m unittest tests.test_loop -v
Worked solution ```python from agentic import Agent, FakeModel, Registry, Stop, tool, tool_call from agentic.tools import ToolError ATTEMPTS = {"n": 0} @tool(description="Read the shipment tracker.", awb="The airway bill number.") def track(awb: str) -> str: ATTEMPTS["n"] += 1 if ATTEMPTS["n"] == 1: raise ToolError("tracker timed out; retry the same awb") return f"{awb}: out for delivery" def reply(messages): tool_msgs = [m for m in messages if m.role == "tool"] if not tool_msgs: return tool_call("track", {"awb": "AWB123"}) if "timed out" in tool_msgs[-1].content: # the model reads the error return tool_call("track", {"awb": "AWB123"}) return "Your parcel is out for delivery." run = Agent(FakeModel([reply, reply, reply]), Registry([track])).run("where is AWB123") run.stop # Stop.ANSWERED run.output # 'Your parcel is out for delivery.' run.requested_tools() # ['track', 'track'] run.trajectory() # ['track'] - only the successful call run.blocked_tools() # ['track'] - the failed one "timed out" in run.messages[2].content # True ``` The callable reply is the point of the exercise. It branches on the *content* of the last tool message, which is only possible because `Registry.dispatch` turned the `ToolError` into a `tool` message and the loop appended it. Swallow that error, or log it server-side and return a generic `"error"`, and the model has nothing to act on - it will either repeat the identical call until `repeat_limit` fires or apologise to the user about a failure it could have recovered from. Also note the honest accounting: two requests, one execution. `trajectory()` counts what ran, `requested_tools()` counts what was asked, and keeping them separate is what stops a retry from looking like two successful shipments.

Checkpoint

Which failures should you never retry, and why? 4xx other than 429. The request itself is wrong, so an identical retry fails identically - you buy three times the latency and three times the error-rate noise for the same outcome.

Why is jitter not optional? Backoff spreads retries over time; jitter spreads them over callers. Without it a bounded fan-out retries in lockstep and rebuilds the burst the backoff existed to disperse.

Why is a timeout worse than a connection error? A connection error means nothing happened. A timeout means you do not know: the call may have completed and billed, so a naive retry risks paying twice and getting a different answer, since generation is not deterministic.

A retried model call re-requests the same refund. Where does the fix belong? In the tool. The loop’s retry wrapper cannot help, because the second request is new and valid. Derive a key from the inputs inside the tool and make the repeat a no-op that says so.

When do you stop retrying? When the failure stops being transient. Open a circuit breaker, fail fast, and in an agent raise FatalToolError so the run ends with a named stop reason instead of burning steps on a dependency that is down.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access