Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 17 min read

Lesson 19 - Durable Execution - Checkpoints and Resumption

Part 4 - Scale the Pattern Lesson 19 of 24

Code: agentic-course/agentic/memory.py Tests: agentic-course/tests/test_memory_evals_guard.py Run it: python3 -m unittest tests.test_memory_evals_guard -v Concept: Agent Memory, State, and Durable Execution covers the theory and the interview framing, without code.


What you will build

There is no Checkpoint class and no workflow engine here. Durability is a pattern built from run.messages plus a place to put it β€” the library gives you the transcript and the resume hook, you supply the store.
πŸ’‘ The transcript is the whole conversation so far - your instructions, the model’s turns, and every tool result - and a checkpoint is that transcript written somewhere a restart can read it back from.


The idea

Everything so far assumed a run fits inside a request. Real agent work does not. A tool waits on a slow upstream. A step needs a human to approve a refund, and that human is asleep. A migration task walks two thousand records. The wall-clock span is minutes to days, and across that span your process will restart β€” a deploy, an OOM kill, a spot instance reclaimed. If a restart loses the run, you have not built an agent, you have built a demo. And the expensive failure is not the lost work, it is the repeated work: a resumed run that replays a payment charges twice.

The analogy. A long-running agent is a form you fill in over several sittings. If the tab crashing loses everything, nobody finishes. If it saves after every field, you come back and continue. And the field that already submitted a payment must not submit again when you reload β€” the harder half of the problem, and the reason lesson 7 comes before this one.


What goes in a checkpoint

Four things, each available on Run:

What Where it comes from Why it must be there
the message list run.messages the agent’s entire state β€” everything it knows
the step index run.step_count so the resumed run knows how far along it is
budget consumed run.input_tokens, run.output_tokens, run.cost otherwise resumption resets the budget and the ceiling means nothing
the plan, if any your planner from lesson 17 the route, so a resume does not re-plan

The third row is the one people skip and the one that costs money. A run that crashes at step seven of eight and resumes with a fresh Budget(max_steps=8) has silently been granted fifteen steps. Carry the consumption forward and subtract it.
πŸ’‘ Tokens are the chunks of text a provider bills you for, so a count that resets on resume is a cost ceiling that no longer means anything.

checkpoint = {"task": "Settle invoice inv-9",
              "messages": [m.to_dict() for m in first.messages],
              "step_index": first.step_count,
              "input_tokens": first.input_tokens,
              "output_tokens": first.output_tokens,
              "cost": first.cost}
[m["role"] for m in checkpoint["messages"]]   # ['system','user','assistant','tool']

Message.to_dict() is real and ships in model.py. from_dict does not β€” the library serialises and you write the rehydration, which must restore tool_calls and tool_call_id or the transcript is malformed:

from agentic import Message, ToolCall

def rehydrate(rows: list[dict]) -> list[Message]:
    return [
        Message(role=row["role"], content=row.get("content", ""),
                tool_calls=tuple(ToolCall(id=c["id"], name=c["name"],
                                          arguments=c["arguments"])
                                 for c in row.get("tool_calls", ())),
                tool_call_id=row.get("tool_call_id"))
        for row in rows
    ]

Idempotency is the precondition, not a nicety

Before any resumption story, the tools have to be safe to call twice. A resumed run replays a transcript, and a model reading that transcript may well re-issue a call it already made β€” it is being cautious, and it has no way to know the first one landed. So the mutating tool carries a caller-supplied key and checks it first, exactly as in lesson 7 and /concepts/idempotency/. EFFECTS is the assertion surface β€” it appends only when a charge genuinely happens, so its length is the number of real side effects across every run below.

from agentic import Agent, Budget, FakeModel, Registry, Stop, tool, tool_call

LEDGER: dict[str, str] = {}
EFFECTS: list[str] = []

@tool(description="Charge an invoice. Idempotent on the key.",
      invoice="The invoice id.", key="A caller-supplied idempotency key.",
      mutating=True)
def charge(invoice: str, key: str) -> str:
    if key in LEDGER:
        return f"already charged {invoice}, ref {LEDGER[key]}"
    LEDGER[key] = f"ref-{len(LEDGER) + 1}"
    EFFECTS.append(key)
    return f"charged {invoice}, ref {LEDGER[key]}"

billing = Registry([charge])

Crash and resume, for real

Run it with a one-step budget so the process β€œdies” right after the charge, before an answer. Then resume from the saved transcript β€” with one detail that matters: Agent.run prepends its own system message and appends the task, so pass history without the system message or you get it twice.
πŸ’‘ The system message is the standing instruction that sits at the top of every prompt, which is why a saved transcript still carrying one ends up with two.

def billing_reply(msgs):
    if not any(m.role == "tool" for m in msgs):
        return tool_call("charge", {"invoice": "inv-9", "key": "inv-9:attempt-1"})
    return "Invoice inv-9 is settled."

first = Agent(FakeModel([billing_reply]), billing, system="Settle invoices.",
              budget=Budget(max_steps=1)).run("Settle invoice inv-9")
first.stop          # Stop.STEP_BUDGET - no answer, but the charge landed
EFFECTS             # ['inv-9:attempt-1']

history = [m for m in rehydrate(checkpoint["messages"]) if m.role != "system"]
resumed = Agent(FakeModel([billing_reply, billing_reply]), billing,
                system="Settle invoices.", budget=Budget(max_steps=4)).run(
    "Continue. Do not repeat completed work.", history=history)

resumed.stop          # Stop.ANSWERED
resumed.output        # 'Invoice inv-9 is settled.'
resumed.trajectory()  # [] - it needed no tools, the result was already in context
EFFECTS               # ['inv-9:attempt-1'] - still exactly one

history= is the real resume mechanism and test_history_is_replayed in tests/test_loop.py pins the replay ordering: history lands before the new task, in order. The resumed model reads the tool result already in its context, needs no tools, and answers.

That is the happy path. The one you must survive is a resumed model being less clever and re-issuing the charge anyway β€” the exercise scripts exactly that, and the outcome is retry.trajectory() == ["charge"] with len(EFFECTS) still 1. The tool executed and the side effect count did not move. That gap between β€œran” and β€œhad an effect” is what idempotency buys. It is also why the key is caller-supplied rather than generated inside the tool: a key the tool invents is different on every call, which is a counter with extra steps. Note too that the loop’s repeat guard tracks (name, stable_arguments) per run β€” seen_calls is local to Agent.run, so an identical call in run one and run two is not a repeat. Across resumptions your tool’s idempotency is the only defence.

flowchart LR
    T[Task] --> L[Agent loop step]
    L --> TL[Idempotent tool call]
    TL --> CP[Checkpoint written]
    CP --> L
    CP --> ST[Durable store]
    X[Process restart] --> RS[Resume]
    ST --> RS
    RS --> RH[Rehydrate messages then pass as history]
    RH --> L
    L --> A[Answer]
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    classDef async fill:#b4f,stroke:#333,color:#000
    class T,A client
    class L,TL,RH service
    class CP,ST data
    class X,RS async

The durable memory half

A transcript is this run’s state. Facts that must outlive it go in a FactStore, which ships with JSON durability:
πŸ’‘ The transcript is what this one run knows and then throws away, while a fact store is the small set of things you want remembered the next time the same customer shows up.

from agentic.memory import Fact, FactStore

facts = FactStore()
facts.write(Fact(key="preferred_courier", value="BlueDart", source="ticket T-1"))
restored = FactStore.load(facts.dump())      # dump is JSON, safe in a row or a blob
restored.get("preferred_courier").value      # 'BlueDart'
restored.get("preferred_courier").source     # 'ticket T-1'

test_round_trips_through_json pins that. Two things carry over from lesson 6, and both matter more once storage is durable. source is what lets you call forget_source() when a document turns out to be poisoned. And FactStore.PROTECTED refuses writes to keys like is_admin and price, because a durable store the model can influence is one an injection can influence permanently.
πŸ’‘ An injection is a hostile instruction planted inside content the agent reads, and the model has no way to tell it apart from an instruction you wrote.


Human-in-the-loop is just slow durability

A confirmation gate is a checkpoint whose resume trigger is a person. Registry(confirm=...) is the real hook: a callable taking (name, args) returning a bool, consulted for any tool with requires_confirmation=True.

APPROVALS: dict[str, bool] = {}

@tool(description="Refund an order.", order_id="The order id.",
      mutating=True, requires_confirmation=True)
def refund(order_id: str) -> str:
    return f"refunded {order_id}"

gated = Registry([refund],
                 confirm=lambda n, a: APPROVALS.get(f"{n}:{a['order_id']}", False))

pending = Agent(FakeModel([tool_call("refund", {"order_id": "4471"}),
                           "I need approval first."]),
                gated, budget=Budget(max_steps=3)).run("Refund order 4471")
pending.blocked_tools()   # ['refund']
pending.trajectory()      # []  - nothing happened

# Checkpoint, page a human, resume when they answer - minutes or days later.
APPROVALS["refund:4471"] = True
after = Agent(FakeModel([tool_call("refund", {"order_id": "4471"}), "Refunded 4471."]),
              gated, budget=Budget(max_steps=3)).run(
    "Continue. The refund is approved.",
    history=[m for m in pending.messages if m.role != "system"])
after.trajectory()        # ['refund']

One shipped limitation, stated plainly: dispatch returns a single denial string, "Error: the user declined this action.", whether the human said no or has not answered yet. The loop has no pending state. Distinguishing declined from awaiting is your approvals store’s job, checked before you resume β€” and worth getting right, because resuming a declined action is a much worse bug than failing to resume an approved one.


When a step cannot complete

Timeouts. Tool.timeout is a declared field on every tool and the shipped dispatch does not enforce it β€” no async, no threads, so a slow tool blocks. Treat the field as documentation of intent and enforce the deadline in your tool body or at your HTTP client. A timeout you believe in and do not have is worse than none.

Compensation. Step three fails after steps one and two committed real effects, and no transaction spans three services. You need compensating actions applied in reverse β€” refund the charge, release the hold β€” the saga pattern: /concepts/saga-pattern/. For an agent that means each mutating tool needs a named inverse, and an abandoning run walks its trajectory() backwards invoking them. A deliberate design, not a retry.

When to stop building this. A checkpoint table, history=, and a hop counter is genuinely enough for a run spanning minutes across one or two systems. Once you need durable timers, exactly-once orchestration across many services, versioned long-lived workflows, and visibility into thousands of in-flight runs, you are rebuilding a workflow engine badly. Reach for Temporal, Cadence or Step Functions, and read /concepts/durable-execution/ first. The reverse mistake is as common: adopting an engine for a nine-second run buys a deployment dependency and a new failure mode for nothing.


Exercise

Checkpoint a run mid-flight, resume it, and assert the side effect happened exactly once β€” including when the resumed model re-requests the same charge. Success criterion: len(EFFECTS) == 1 and len(LEDGER) == 1 after all three runs, while retry.trajectory() == ["charge"] proves the tool did execute on the resume.

python3 -m unittest tests.test_memory_evals_guard -v
Worked solution Reuse `charge`, `billing`, `LEDGER`, `EFFECTS`, `billing_reply`, `checkpoint` and `rehydrate` from above: ```python import json # 1. Crash: the charge lands, the answer never arrives. assert first.stop is Stop.STEP_BUDGET and EFFECTS == ["inv-9:attempt-1"] # 2. Resume through a real JSON round trip, system message stripped. saved = json.loads(json.dumps(checkpoint)) history = [m for m in rehydrate(saved["messages"]) if m.role != "system"] assert history[1].tool_calls[0].name == "charge" # the call survived assert history[2].tool_call_id is not None # so did the pairing warm = Agent(FakeModel(["Invoice inv-9 is settled."]), billing, system="Settle invoices.", budget=Budget(max_steps=4)).run( "Continue. Do not repeat completed work.", history=history) assert warm.stop is Stop.ANSWERED and len(EFFECTS) == 1 # 3. The hostile case: a resumed model that re-issues the charge. same_call = tool_call("charge", {"invoice": "inv-9", "key": "inv-9:attempt-1"}) retry = Agent(FakeModel([same_call, "Already settled."]), billing, budget=Budget(max_steps=3)).run( "Settle invoice inv-9", history=history) assert retry.trajectory() == ["charge"] # it ran assert len(EFFECTS) == 1 and len(LEDGER) == 1 # and changed nothing ``` **Why step three is the real test.** Steps one and two prove resumption works when the model behaves; step three proves it is *safe* when it does not, and you do not choose which one production gives you. The assertion pair is the point: `trajectory() == ["charge"]` says the tool executed, `len(EFFECTS) == 1` says it had no second effect. Durability without that gap is a duplicate-charge generator with good logging. Note the JSON round trip in step two, too. Serialising through `to_dict` and back is where a checkpoint usually breaks, because a dropped `tool_calls` tuple or a missing `tool_call_id` produces a transcript that looks fine and reads wrong. Assert on the rehydrated shape, not only on the final answer.

Checkpoint

What are the four parts of a checkpoint, and which is most often forgotten? The message list, the step index, budget consumed, and the plan if there is one. Budget consumed is the one people skip β€” a resumed run with a fresh Budget has silently been granted a second full allowance.

Why must tools be idempotent before you can resume at all? Because a resumed run replays a transcript, and the model may re-issue a call it already made without knowing the first one landed. A caller-supplied idempotency key makes the second execution a no-op. The loop’s repeat guard cannot help: seen_calls is local to one Agent.run.

What does passing history= actually do, and what does the confirm gate not tell you? history= inserts those messages between the system prompt and the new task, in order β€” pinned by test_history_is_replayed. Strip the system message from a saved transcript first, or the resumed run carries it twice. The confirm gate cannot distinguish declined from not-yet-answered: dispatch returns one denial string for both, so your approvals store holds that distinction.

When is a workflow engine the right answer? When you need durable timers, exactly-once orchestration across many services, versioned long-lived workflows, and visibility into thousands of in-flight runs. Below that, a checkpoint table plus history= is enough β€” and adopting an engine for a nine-second run buys a dependency and a new failure mode for nothing.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access