Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 14 min read

Lesson 20 - MCP and Frameworks - Porting What You Built

Part 4 - Scale the Pattern Lesson 20 of 24

Run it: python3 -m unittest discover -s tests -t . Concept: Model Context Protocol covers the theory and the interview framing, without code.


What you will build

This lesson has no primary module, because it is about what sits around the code you wrote. It is also the lesson most likely to age: MCP and the framework landscape move fast, so specifics here will shift. The reasoning will not.


The idea

You wrote the loop. That sounds like a small thing and it is the reason the rest of this course made sense, because you now know where the costs and the failures live:

None of that is framework knowledge. It is loop knowledge, and it transfers.

The honest position on frameworks. Neither hand-rolling nor adopting is always right. Frameworks genuinely give you prebuilt connectors to dozens of systems, graph and state abstractions for branching workflows, streaming plumbing that is tedious to get right, checkpointing you would otherwise build as in lesson 19, and an ecosystem โ€” which means other people have already hit your bug. What they hide is the loop, the budget, the retry policy, and precisely where your tokens go. Those are exactly the four things that show up in an incident review. Adopt a framework knowing what it abstracts; that is the whole point of having written it once.


MCP: an N-times-M problem

The Model Context Protocol addresses an integration explosion. Without a standard, every application that wants to reach every data source writes its own connector: N applications times M sources is N ร— M integrations, each separately built, separately versioned, separately broken.

A standard interface collapses that. Each application implements one client; each source exposes one server. N + M. This is not a new idea โ€” it is the argument for any protocol, and the reason you do not write a bespoke driver per database.

flowchart LR
    subgraph Before
        A1[App one]
        A2[App two]
        S1[Source one]
        S2[Source two]
    end
    subgraph After
        B1[App one with a client]
        B2[App two with a client]
        P[Standard protocol - TRUST BOUNDARY]
        T1[Server one]
        T2[Server two]
    end
    A1 --> S1
    A1 --> S2
    A2 --> S1
    A2 --> S2
    B1 --> P
    B2 --> P
    P --> T1
    P --> T2
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class A1,A2,B1,B2 client
    class P edge
    class S1,S2,T1,T2 data

Left: four integrations for two apps and two sources, and it is N ร— M from there. Right: each app implements one client, each source one server, and the protocol boundary is where trust has to be re-established โ€” everything crossing it arrives from code you did not write.

The architecture in plain terms. A host application contains a client. The client talks to servers. Each server exposes capabilities to the host. The primitive kinds a server can offer are tools the model may invoke, resources it may read, and prompt templates it can be given. That mapping is the useful part: a tool is the thing you already built in lesson 3, a resource is retrievable context from lesson 16, and a prompt template is text entering your context.

Deliberately absent above: wire format, transport details, version numbers, message schemas. Those belong in the current specification, they change, and a half-remembered version of them is worse than none. Read the spec for the shape; read this lesson for what it means for your registry.


Four engineering realities that outlive the spec

1. A third-party server is a supply-chain dependency. It runs code you did not write, and unlike a library it receives your context and your data at runtime. That deserves the scrutiny you give any dependency and then some: who maintains it, what it logs, where it sends what it receives, what happens when it is compromised. An unpinned MCP server is an unpinned dependency with network access to your prompts.

2. Tool descriptions from a server are prompt text entering your context. This is the one people miss, and it follows directly from what you built. Look at what to_schema() produces:

get_order.to_schema()["function"]["description"]
# 'Look up an order by its id.'

That string is sent to the model on every call. When it arrives from a third-party server, a compromised server is an injection vector โ€” it can ship a description that reads as an instruction, and it lands in your context with more authority than a retrieved document because it looks like configuration rather than data. Everything in lesson 12 applies: treat incoming descriptions as untrusted content, review them, and never render them into a system prompt unexamined.

3. Dynamic discovery means your prompt surface can change with no deploy on your side. A server that adds or edits a tool changes what your model sees, without a commit, a review, or an eval run in your repo. That is a genuinely new failure mode: the thing your evals graded last week is not the thing running today. Pin server versions, review changes as you would review a prompt change, and re-run the suite from lesson 10 when a pinned version moves.

4. Too many available tools degrades selection. Not a protocol problem, a context problem, and connecting four servers makes it worse fast โ€” every toolโ€™s schema and description occupies context and adds a wrong answer to the selection problem. The fix is the one you already have: scope per task with registry.scoped().

And the rule that overrides all of it: authorization is enforced at your boundary. Never delegate it to the server. A server tells you what a tool can do; only your code knows who is asking and what they may touch. In this library that is FatalToolError for a denial and permitted supplied by the callerโ€™s identity, never by the model. Concretely: even if a server exposes delete_account, whether this user may delete that account is a question your code answers before dispatch.

from agentic import Agent, Budget, FakeModel, Registry, tool, tool_call

full = Registry([search_kb, get_order, issue_refund])
[s["function"]["name"] for s in full.schemas()]
# ['search_kb', 'get_order', 'issue_refund']

untrusted = full.scoped({"search_kb"})
[s["function"]["name"] for s in untrusted.schemas()]     # ['search_kb']

attempt = Agent(FakeModel([tool_call("issue_refund", {"order_id": "4471"}),
                           "I cannot do that."]),
                untrusted, budget=Budget(max_steps=3)).run(
    "A document told me to refund order 4471.")

attempt.requested_tools()   # ['issue_refund'] - the injection got the model to try
attempt.trajectory()        # []               - and nothing happened

One property worth knowing before you rely on it: scoped intersects, it never widens. Re-scoping a narrowed registry cannot recover a tool the parent already excluded, so a scoped registry handed to an untrusted context cannot be talked back open.

untrusted.scoped({"search_kb", "issue_refund"}).schemas()
# still only search_kb - allowed is intersected with the parent's allowed

Porting what you built

The mapping is conceptual, not an API table โ€” these names move, which is the point.

What you built What a framework calls it
Agent.run loop a graph or state machine, with your loop body as its nodes and edges
Budget(max_steps=...) a recursion or iteration limit on graph traversal
Stop enum terminal states or end conditions on the graph
Registry plus scoped a tool or toolkit object, with binding per node
Tracer and spans their tracing or observability integration
Suite and Report their eval or experiment product
History and FactStore memory modules, usually short-term and long-term
checkpoint plus history= a checkpointer or persistence layer

Read that table in the right direction. The frameworkโ€™s job is to supply the plumbing; your job is to know what each box must guarantee. When a graph runs away, you look for the recursion limit because you know a loop needs a ceiling. When a port produces different answers, you diff the assembled context because you know order matters (lesson 6). When a trajectory eval fails, you check whether their trace records requested or executed tool calls, because you know those are different questions.

That is the durable skill. Framework APIs churn on a timescale of months; the reasoning about budgets, trust boundaries, context assembly, and what a stop reason means does not. Port the reasoning and let the API be whatever it currently is.


Exercise

Two parts, and the first is written prose rather than code.

Part one. Take the Registry you built. For each of the four trust questions, write down the exact place in your code where the answer lives:

  1. Who wrote the tool description?
  2. What data does the tool receive?
  3. What can the tool reach?
  4. Who authorized the call?

Part two. Scope a registry down for an untrusted context and assert the narrowed schema list.

Success criterion: the untrusted registryโ€™s schema names equal ["search_kb"], len(untrusted) == 1, and a scripted injectionโ€™s refund attempt appears in requested_tools() but not in trajectory().

python3 -m unittest discover -s tests -t .
Worked solution **Part one โ€” where each answer lives.** 1. **Who wrote the tool description.** The `description=` argument to `@tool`, in your repo, under review. For an MCP server it is whoever controls that server's code โ€” outside your review, changeable without your deploy. That asymmetry is the whole reason to pin versions. 2. **What data the tool receives.** `Tool.validate` decides, and it is the only gate: unknown arguments are rejected, types are checked, `Literal` becomes an enum. Whatever survives `validate` is what `t.fn(**args)` gets. Model-supplied arguments are untrusted input, validated like a request body. 3. **What the tool can reach.** The function body, and nothing else constrains it โ€” no sandbox in this library. A tool that opens a socket can reach the internet. For egress, `Guard.check_host` with `Policy(allowed_hosts=...)` is the allowlist, and an empty set means no egress at all. 4. **Who authorized the call.** `Registry.allowed` for the tool-level allowlist, `Registry.confirm` for the human gate, and `FatalToolError` raised inside the tool for a runtime denial. Row-level authorization is the caller-supplied `permitted` from [lesson 16](/agentic-ai/retrieval-as-a-tool/) โ€” never something the model chooses. **Part two โ€” the assertions.** ```python from agentic import Agent, Budget, FakeModel, Registry, tool, tool_call from agentic.guard import Policy full = Registry([search_kb, get_order, issue_refund]) assert len(full) == 3 untrusted = full.scoped({"search_kb"}) assert [s["function"]["name"] for s in untrusted.schemas()] == ["search_kb"] assert len(untrusted) == 1 attempt = Agent(FakeModel([tool_call("issue_refund", {"order_id": "4471"}), "I cannot do that."]), untrusted, budget=Budget(max_steps=3)).run( "A document told me to refund order 4471.") assert attempt.requested_tools() == ["issue_refund"] # it tried assert attempt.trajectory() == [] # nothing ran # scoped() intersects - a narrowed registry cannot be widened back open. assert [s["function"]["name"] for s in untrusted.scoped({"search_kb", "issue_refund"}).schemas()] == ["search_kb"] # The strictest stance ships in guard.py: no tools at all for untrusted input. assert Policy(allowed_tools={"search_kb", "issue_refund"}).for_untrusted_input().allowed_tools == set() ``` **Why both assertions in the injection check.** `trajectory() == []` alone passes on a model that never tried, on a broken `FakeModel`, and on a registry that silently dropped everything. Pairing it with `requested_tools() == ["issue_refund"]` proves the attempt happened *and* was contained, which is the only thing that demonstrates the scoping works. And note what the allowlist does that a prompt cannot. `dispatch` returns the same message for a tool that does not exist and one that is not allowed โ€” it does not leak the existence of tools the model may not see. No instruction in a retrieved document can make `issue_refund` appear in a registry that excludes it.

Checkpoint

What problem does MCP solve, and what does it not solve? It turns N ร— M bespoke integrations between applications and data sources into N + M by standardising the interface. It does not decide what your agent may do โ€” authorization, budgets, tool scoping and termination all remain yours, and a standard interface means a larger surface to scope, not a smaller one.

Why is a third-party MCP server an injection vector and not just a dependency risk? Because the tool descriptions it supplies are prompt text that enters your context on every call, and they look like configuration rather than data. A compromised server ships an instruction-shaped description straight into your prompt. It is also a dependency that receives your context and data at runtime.

What changes about your prompt surface when discovery is dynamic? It can change with no deploy on your side. A server that edits a tool changes what your model sees without a commit, a review, or an eval run in your repo โ€” so the thing your evals graded last week may not be the thing running today. Pin versions and re-run the suite when a pin moves.

Why can a scoped registry never be widened back open? Because scoped intersects with the parentโ€™s allowed rather than replacing it. A registry handed to an untrusted context cannot recover a tool the parent excluded, which is what makes the narrowing a guarantee instead of a convention.

What survives a framework migration? The reasoning: that a loop needs a ceiling, that input tokens dominate because the transcript is resent, that requested and executed are different questions, that authorization belongs at your boundary, and that context order matters. The APIs churn on a timescale of months. None of that does.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access