Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 21 min read

Model Context Protocol - Complete Deep Dive

Stage 5 - Agents Lesson 23 of 36

Prerequisites: Tool Calling, Agent Architectures, API Design Used in: Agent Memory and State, Prompt Injection and Jailbreaks, Service Discovery Build it: Lesson 20 - MCP and Frameworks - Porting What You Built implements this as runnable, tested code you can execute offline.


What Is the Model Context Protocol?

The Model Context Protocol, MCP, is an open standard for how a model-facing application connects to external capabilities β€” tools it can invoke, data it can read, and prompt templates it can reuse. It is an integration interface, not a model feature and not a framework. Its whole purpose is to let one implementation of β€œhere is how to talk to our ticketing system” be reused by any application that speaks the protocol, instead of being rewritten per application.

Real-world analogy: cargo before and after the shipping container. Every commodity used to need bespoke handling at every port β€” a crane, a crew, and a procedure per pairing of cargo type and dock. Standardising the box did not make ships faster; it made every box compatible with every crane, so the number of things anyone had to build collapsed. MCP is the box. It does not make models smarter. It makes capabilities portable.

Two framings to hold onto, because they are what the rest of the page is about. Upside: a capability implemented once becomes available everywhere, and discovery becomes dynamic instead of compiled in. Downside: a capability implemented by someone else, discovered at runtime, is a dependency that reads your context and receives your data. Both are consequences of the same design.
πŸ’‘ A model’s context is the single block of text it reads on every call - your instructions, the conversation so far, and whatever a tool handed back.


The Problem It Addresses

Before a standard, every integration was bespoke per application. Your chat product wrote a connector to the document store; your internal agent wrote a different connector to the same document store; a third team wrote a third. With N applications and M data sources, you are looking at N times M integrations, each separately authored, separately authenticated, separately tested, and separately broken when the upstream API changes.

A standard interface between an application and a capability collapses that to N plus M. Each application implements the client side once. Each capability implements the server side once. Any client works with any server.

flowchart LR
    subgraph Bespoke["Before - N times M bespoke integrations"]
        A1[App one] --> X1[Bespoke adapter]
        A2[App two] --> X2[Bespoke adapter]
        X1 --> S1[Document store]
        X2 --> S1
        X1 --> S2[Ticket tracker]
        X2 --> S2
    end

    subgraph Standard["After - N plus M with one protocol"]
        B1[App one] --> P[Protocol client]
        B2[App two] --> P
        P --> M1[Server for document store]
        P --> M2[Server for ticket tracker]
    end

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A1,A2,B1,B2 client
    class X1,X2 edge
    class P,M1,M2 service
    class S1,S2 data

This is the same argument that produced ODBC for databases and the Language Server Protocol for editors, and the same argument gets made for every integration standard that sticks. It is not a new idea β€” it is a well-worn idea applied to model applications, which is a reason to take it seriously rather than a reason to discount it.


The Architecture in Plain Terms

Three roles, and the names matter because they map onto a trust boundary.

  1. Host β€” your application. The chat product, the IDE, the agent runtime. It owns the model connection, the user relationship, the policy layer, and the decision about which capabilities are available for a given task.
  2. Client β€” the protocol-speaking component inside the host. It maintains a connection to a server and mediates between the host and that server. One host typically runs several clients, one per connected server.
  3. Server β€” a process that exposes capabilities over the protocol. It may wrap a local filesystem, a database, an internal service, or a third-party SaaS API.
    πŸ’‘ The names run backwards from web habit: the client is a small component inside your own application, and a server is often just a helper process on the same laptop.

The model itself is not in that list, and that is the important part. The model never talks to a server. The host decides what to put in the model’s context, the model proposes an action, and the host β€” through a client β€” decides whether to execute it. Every enforcement point you care about lives in the host.

flowchart LR
    subgraph Trusted["Your host application - the trust boundary"]
        U[User] --> H[Host runtime and policy layer]
        H --> AZ[Authorization and audit and rate limits]
        H --> CL[Protocol clients - one per server]
        H --> LM[Model call]
    end

    CL --> LS[Local server as a child process]
    CL --> RS[Remote server over the network]
    LS --> FS[Local files and developer tooling]
    RS --> TP[Third party or internal service]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class U client
    class H,AZ service
    class CL,LM edge
    class LS,RS async
    class FS,TP data

Everything outside the Trusted box is code you did not write, running on data you did send.


What a Server Can Offer

Servers expose a small number of primitive kinds. The useful distinction between them is who initiates and what the thing does, not the wire details.

Primitive What it is Who drives it Engineering note
Tools Callable operations with a described input shape, which may have side effects The model proposes, your host executes This is the risky one. Treat every invocation as untrusted intent needing authorization
Resources Readable content the host can pull in as context - documents, records, files The host or user selects Read-only by nature, but the content lands in your prompt so it can carry injected instructions
Prompt templates Reusable parameterised instruction blocks a server offers for common tasks The user or host chooses Server-authored text that becomes your system instructions. Review it like a dependency, not a config value

A server typically advertises what it offers when a client connects, so the host learns the available capabilities at runtime rather than having them hardcoded. That discovery step is the feature and the hazard, and it gets its own section below.

On wire-format specifics: this page deliberately does not give them. The protocol uses a conventional JSON request-and-response message scheme, the exact envelope and method names have changed across revisions, and any article that pins them ages badly. Read the current specification for the wire contract. The concepts above are the durable part.


Local and Remote Servers

Conceptually there are two deployment shapes, and the difference is mostly about trust and lifecycle rather than capability.

A local server runs on the same machine as the host, usually launched by the host as a child process and talked to over its standard input and output streams. No network hop, no inbound port, and the server inherits the machine’s identity β€” which makes it convenient for developer tooling and filesystem access, and means a compromised local server has whatever access the user has.

A remote server runs elsewhere and is reached over the network via HTTP. Now you have the usual distributed-systems concerns and they are not optional: transport security, authentication of the host to the server and of the user to the underlying data, multi-tenancy, timeouts, retries, and rate limits. The naming and details of the HTTP-based transport have been revised more than once as the standard has matured, so treat the specific mechanism as a lookup rather than something to memorise.

The decision rule is simple enough: local for machine-scoped capabilities where the user’s own identity is the right identity; remote for shared, multi-user, or centrally governed capabilities where you need an audited boundary. Remote servers then get the treatment any service dependency gets β€” Rate Limiting, Retries and Backoff, and a timeout on every call.


Why a Standard Matters Operationally

Beyond the integration count, three practical wins.

One implementation, many hosts. A team that owns a data source can publish one server and have it work in every compatible application β€” the assistant, the IDE, the internal agent β€” without knowing about any of them. That inverts the usual ownership problem, where the application team writes integrations for systems they do not understand.

Discovery becomes dynamic. Capabilities are advertised at connection time rather than compiled in, so a server can gain a tool and every connected host can use it without a client release.

A shared vocabulary for governance. Once integrations have a uniform shape, the controls around them can be uniform too: one place to enumerate what is connected, one place to apply per-capability authorization, one place to log invocations. That is genuinely hard to retrofit onto a pile of bespoke connectors, and it is the argument that matters most at organisational scale.

Aspect Bespoke per-application integration Protocol-based integration
Build cost Rewritten per application - N times M Client once per app plus server once per source - N plus M
Who owns it The application team, for a system they do not own The team that owns the data source
Capability changes Every consuming app needs code changes Servers advertise changes and hosts pick them up
Reuse across hosts None Any compatible host
Trust surface Known at build time and reviewed in your own code Third-party code and third-party text entering at runtime
Discovery Compiled in and explicit Dynamic and therefore harder to pin
Governance Per-integration and inconsistent Uniform enumeration and policy point
Failure mode Breaks visibly at build or deploy time Can change behaviour with no deploy on your side

Notice that the last three rows are not wins. The same properties that make the standard valuable are the ones that create the risks below, and pretending otherwise is how teams get surprised.


The Engineering Realities

This is where the page earns its place. Everything above is the sales pitch and it is mostly true. Here is the part that determines whether your deployment is defensible.

A third-party server is a supply-chain dependency

An MCP server you did not write does two things simultaneously: it receives the data you send it β€” arguments, and often surrounding context β€” and it returns content that enters your model’s context. That is a strictly larger blast radius than a normal library dependency, because it is code running with network access on your users’ data.

Apply dependency discipline, not plugin enthusiasm: know who publishes it, pin a specific version, read the source when it is available, review what network egress and credentials it wants, run it with least privilege, and prefer servers you or your organisation operate for anything touching sensitive data. β€œIt was one line in a config file” is how supply-chain incidents start, and a config file is exactly how these get installed.

Tool descriptions are prompt text

A server tells your host what its tools do, in prose, and that prose goes into the model’s context so the model can choose between tools. Server-supplied descriptions are therefore untrusted input reaching your model as instructions. So are resource contents, and so are prompt templates.

A malicious or compromised server can put text in a description that redirects the model β€” instructions to prefer its tool, to pass along data from other parts of the conversation, to ignore a policy in your system prompt. The user sees a tool list; the model sees a paragraph of attacker-controlled text sitting alongside your instructions. This is indirect prompt injection with an unusually convenient delivery channel.
πŸ’‘ Indirect prompt injection is when the hostile instruction arrives inside content the model was told to read, not from the person typing.
The defences are the ones in Prompt Injection and Jailbreaks. Treat all server-supplied text as data. Enforce authorization outside the model, constrain what any single tool call can reach, and log the descriptions you actually sent so an incident is reconstructible.

Dynamic discovery means your prompt surface can change without a deploy

If a server advertises its capabilities at runtime, then the text in your prompt and the set of actions your model can take are determined by something outside your release process. A server update can add a tool, reword a description, or change what a tool does, and your system’s behaviour changes with no commit on your side. Every eval you ran was against a prompt surface that no longer exists.
πŸ’‘ Your prompt surface is all the text that reaches the model on a call - your own instructions plus every tool description and document the host pulled in.

Mitigations, in rough order of importance. Pin server versions and treat an upgrade as a code change that goes through review and re-evals. Snapshot the advertised capability set and diff it on connect, alerting on unexpected change rather than silently accepting it. Then maintain an allowlist of permitted tools rather than exposing whatever appears, and record the exact descriptions and schemas in your traces so you can tell which surface produced a given run.

Too many tools degrades selection quality

Tool selection is a discrimination task, and it gets harder as the candidate set grows. Connect six servers exposing sixty tools and you get worse choices than with the eight tools the task actually needs. You also get overlapping tools whose near-identical descriptions the model cannot distinguish. And you pay for all sixty definitions on every model call in a loop, which compounds badly per Agent Architectures.
πŸ’‘ Tool definitions are not registered once and remembered - the full text of every one is resent on every model call, so sixty tools means sixty descriptions billed each time round the loop.

Scope the tool set per task, not per installation. Available servers and exposed tools are different things. Group tools by workflow and expose the relevant group, or use a cheap routing step to select the group before the main call. Fewer, clearly differentiated, non-overlapping tools beat comprehensive coverage every time.

Authorization is yours to enforce

The most common serious mistake: assuming the server handles permissions. It might, for its own resources, under its own credentials β€” which are frequently the host’s credentials, not the end user’s. That means a server can happily perform an action the requesting user was never entitled to, and it will look like a successful tool call.

Enforce at your boundary. The host knows who the user is; the server generally does not. Check the user’s entitlement for every invocation in your code, before the client dispatches. Scope credentials per user where the underlying system supports it rather than handing a broadly-privileged token to a shared server. Gate destructive and irreversible operations behind human confirmation. Audit every invocation with user, server, tool, arguments, and outcome. And never let the model’s proposed action be the authorization decision β€” it is a request, subject to the same checks as any other request. Authentication and API Design apply unchanged.


Bad to Good to Great

Bad - connect every interesting server and let the model figure it out

Install a handful of community servers from config snippets, expose all their tools, run on a broadly-privileged token, no version pinning, no audit trail.

You now have unreviewed third-party code receiving your users’ data, and attacker-controllable text sitting in your system context. The action surface can change without a deploy, and tool selection degrades from sheer count. Worst of all, you cannot answer β€œwhich tool did what, for whom, with what arguments” after an incident.

Good - a reviewed and pinned set with tool-level allowlisting

Servers are reviewed before adoption and pinned to a version. Only named tools are exposed. Destructive operations require confirmation. Invocations are logged. Remote servers have timeouts and rate limits.

This is a defensible position and enough for most products. What it still misses: tool sets are scoped per installation rather than per task, so selection quality erodes as the catalogue grows. Server-supplied descriptions get trusted implicitly once the server is approved, and credentials are usually host-scoped rather than user-scoped. Capability drift within a pinned version goes unnoticed.

Great - servers as governed dependencies with a per-task capability surface

  1. Servers in a dependency inventory with an owner, a pinned version, a review record, and an upgrade path that triggers re-evals.
  2. Self-operated servers for sensitive data. Third-party servers are for third-party data, not for your primary records.
  3. Per-task tool scoping. A routing step or workflow stage selects the small tool group relevant to the current task instead of exposing the union of everything connected.
  4. Capability snapshot and diff on connect, alerting on any change to the advertised tool set, schemas, or descriptions.
  5. Server text handled as untrusted data β€” descriptions, resources, and prompt templates are bounded, recorded in traces, and never granted instruction authority over your system prompt.
  6. Authorization in the host, per user, per invocation, with user-scoped credentials wherever the underlying system supports them and human confirmation on irreversible actions.
  7. Full audit and tracing of every invocation, sufficient to replay a run and reconstruct which capability surface produced it.
  8. Evals that cover the integration, including tool-selection accuracy with the real exposed set and a regression run after any server upgrade.

The difference between Good and Great is that Great treats connected capabilities as part of the deployed system β€” versioned, scoped, audited, and evaluated β€” rather than as configuration someone added once.


A Note on Pace

This area is moving quickly. The primitive kinds, the transport mechanisms, the authorization story, and the surrounding tooling have all changed over successive revisions. Anything written down about specific methods, envelopes, or version numbers has a short shelf life - including this page, which is why it stays at the level of concepts.

The durable skill is not the current specification. It is the integration and trust reasoning. Recognise the N times M problem and why standardising the interface collapses it. Know that a connected capability is a dependency with a supply chain, that text arriving from it is untrusted input to your model, that dynamic discovery moves part of your prompt surface outside your release process, and that authorization belongs at your boundary. Those transfer to whatever the ecosystem looks like in two years, and to whatever competing standard shows up alongside it.


When to Use

βœ… Adopt a protocol-based integration when:

❌ Stay with a direct integration when:


Common Interview Questions

Q1: What problem does MCP actually solve?

The integration explosion. Before a standard, every model application wrote its own bespoke connector to every data source, so N applications and M sources meant N times M integrations - each separately authored, authenticated, tested, and separately broken when an upstream API changed. Standardising the interface between a model application and an external capability turns that into N plus M: each application implements the client side once, each capability implements the server side once, and any client works with any server. The secondary win is ownership - the team that owns a data source can publish one server instead of application teams writing integrations for systems they do not understand.

Q2: Walk through the architecture and say where the trust boundary sits.

A host application - your product - contains one or more clients, each holding a connection to a server that exposes capabilities. Servers offer a few primitive kinds: tools the model can invoke, resources the host can read as context, and reusable prompt templates. The key structural point is that the model never talks to a server. The host decides what enters the model’s context, the model proposes an action, and the host decides whether to execute it. So the trust boundary is the host: everything outside it is code you did not write, running on data you did send. Every enforcement point that matters - authorization, audit, rate limiting, tool scoping - has to live inside the host, because that is the only component that knows who the user is.

Q3: What are the security implications of adding a third-party MCP server?

Three distinct ones. First, supply chain: the server receives your data and runs code you did not write, so it needs dependency-grade scrutiny - known publisher, pinned version, least privilege, reviewed egress. Second, injection: tool descriptions, resource contents, and prompt templates from that server become text in your model’s context, which makes a malicious or compromised server an indirect prompt-injection vector that can steer the model or try to exfiltrate conversation data. Third, authorization: the server often acts under the host’s credentials rather than the end user’s, so it will happily perform an action the requesting user was never entitled to. The mitigation for the third is structural - check entitlement in your own code before dispatching, never treat the model’s proposed action as the authorization decision.

Q4: Dynamic tool discovery sounds convenient. What is the catch?

Your prompt surface and your action surface stop being controlled by your release process. A server update can add a tool, reword a description, or change a tool’s behaviour, and your system behaves differently with no commit on your side - which also means every eval you ran was against a surface that no longer exists. So I pin server versions and treat an upgrade as a reviewed code change with a re-eval. I snapshot the advertised capability set and diff it on connect, so unexpected changes alert rather than pass silently. I also keep an explicit allowlist of permitted tools instead of exposing whatever appears, and record the exact descriptions and schemas in traces so any run is reconstructible.

Q5: You have connected six servers exposing sixty tools and quality dropped. Why?

Tool selection is a discrimination task and it degrades as the candidate set grows, especially with near-duplicate tools whose descriptions the model cannot tell apart. You are also paying for all sixty definitions on every model call, which in an agent loop compounds on each iteration. The fix is to separate available from exposed: scope the tool set per task rather than per installation, grouping tools by workflow and surfacing only the relevant group, or adding a cheap routing step that picks the group before the main call. Eight clearly differentiated non-overlapping tools beat sixty comprehensive ones, and tool-selection accuracy against the real exposed set belongs in the eval suite so this shows up as a measured regression rather than a vibe.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access