The substrate that connects AI to a customer’s systems — MCP’s host/client/server model and primitive vocabulary, why it reframes integration from N×M to N+M, the consent/exec security boundaries, and the gateway pattern (Runlayer’s $11M bet on 18,000+ servers) that collapses the OAuth sprawl.
How the AI reaches the customer’s systems
An enterprise AI solution is only as useful as the systems it can reach — the CRM, the ticketing system, the data warehouse, the internal API. The Model Context Protocol (MCP) is now the default substrate for that connector work, defined as “an open protocol that enables seamless integration between LLM applications and external data sources and tools.” For a solutions engineer, MCP is both the architecture you propose when a customer asks “can it talk to our systems?” and a deep-probe topic in the architecture round.
The architecture is a host / client / server model. The MCP Host is the AI application (e.g. Claude Desktop, your agent runtime); the MCP Client lives inside the host and maintains a dedicated connection to one server; the MCP Server exposes context and capabilities. Communication runs as JSON-RPC 2.0 over one of two transports: stdio (local process-to-process) or streamable HTTP (client-to-server via HTTP POST, with optional server-sent events for streaming). That’s the whole topology — and being able to draw it cleanly is the first signal you’ve built with it.
code
1MCP TOPOLOGY (host / client / server over JSON-RPC 2.0)23 +-----------------------------------------+4 | MCP HOST (the AI app / agent runtime) |5 | +-------------+ +-------------+ |6 | | MCP Client | | MCP Client | ... | one client per server7 | +------+------+ +------+------+ |8 +----------|-----------------|------------+9 | stdio | streamable HTTP (POST + optional SSE)10 +----v-----+ +-----v------+11 | Server A | | Server B | each exposes Resources/Prompts/Tools12 | (files) | | (CRM API) |13 +----------+ +------------+1415 Transports: stdio = local process; streamable HTTP = remote (auth: bearer/OAuth).
The primitive vocabulary is small and standardized, which is the whole point. Servers expose Resources, Prompts, and Tools (data, templates, executable functions). Clients expose Sampling, Elicitation, and Logging (the ability to ask the LLM to complete, to ask the user a clarifying question, and to emit structured logs back to the server). Experimental “Tasks” add durable-execution wrappers for multi-step operations. Because the vocabulary is fixed, an integration author learns one protocol and one set of primitives, and every new connector is mechanically the same shape — that uniformity is what makes MCP worth standardizing on.
The catalyst for MCP is the N×M integration problem. Without a standard, every (agent, tool) pair needs custom glue: Speakeasy’s gateway analysis puts the math plainly — scaling to five MCP servers and three agents means managing 15 OAuth implementations, 15 credential stores, and 15 audit trails. With MCP, every connector becomes a tool or resource behind one protocol, so the integration surface collapses from N×M custom integrations to N+M conforming endpoints — you write each server once and each host speaks to all of them. Interview angle. “Why MCP instead of just writing API clients?” → it reframes integration from N×M bespoke glue to N+M standard endpoints; the integrations “build themselves” because every connector is the same shape.
Transports and the failure modes at scale
The two transports have very different operational profiles, and choosing wrong is a documented source of production pain. stdio is a local process-to-process pipe — fast, no network, but it ties the server’s lifecycle to the host process and doesn’t cross a machine boundary, so it’s for local tools (a desktop file server, a local database) not a shared enterprise service. Streamable HTTP is the remote transport: a client POSTs JSON-RPC and the server may stream responses over server-sent events. It crosses the network, carries bearer/OAuth auth, and is what every multi-tenant enterprise connector uses — but it inherits all the usual distributed-systems failure modes (timeouts, partial reads, dropped SSE streams, retries) that stdio never had.
Two scale failure modes a senior solutions engineer names before the customer hits them. (1) Long-running tool calls vs request timeouts: a synchronous JSON-RPC call that takes 90 seconds (a big report, a slow downstream API) collides with gateway and load-balancer timeouts — which is exactly what MCP’s experimental Tasks primitive addresses by wrapping multi-step operations in a durable, pollable execution rather than one blocking call. (2) Connection fan-out: one client maintains a dedicated connection per server, so an agent talking to a dozen servers holds a dozen live connections, and an SSE stream that silently drops looks like a hung agent, not an error. The fix for both is the gateway (next): it terminates and pools connections, translates protocols, and turns long calls into managed jobs. Interview angle. “stdio or HTTP, and what breaks?” → stdio for local/lifecycle-bound tools; streamable HTTP for remote/multi-tenant, where timeouts on long tool calls and dropped SSE streams are the failure modes — and durable Tasks plus a gateway are how you contain them.
Security is in the spec, not bolted on
A tool call is arbitrary code execution on behalf of the user, so MCP names three binding security principles in the spec itself: user consent (users must explicitly consent to and understand all data access and tool operations), privacy and safety (hosts must protect user data and treat tool execution with caution), and sampling controls (users must approve any LLM sampling requests). For streamable HTTP, the spec adds bearer-token / API-key / custom-header auth and recommends OAuth. These are the minimum bar; enterprise deployments add IdP integration and zero-trust on top. Interview angle. “What’s the risk when an agent can call tools, and how do you contain it?” → tool exec is arbitrary code; contain it with explicit user consent, a per-agent allow-list of servers, IdP-checked permissions, and an audit log of every call — never an open tool registry.
The gateway pattern: collapsing the OAuth sprawl
When the connector count grows, the N×M OAuth problem comes back as an operational nightmare, and the answer is an MCP gateway: a central proxy in front of every server that holds credentials, brokers auth, logs every call, and enforces policy. Two flavors exist in practice: infrastructure-only (routing and protocol translation — you bring auth/observability) and full (auth, audit, rate limiting, and policy enforcement in one product). Tyk formalizes the gateway as “an intermediary layer that intercepts and manages all communication between AI agents and their tools.” The rule of thumb: reach for a gateway once an agent uses more than roughly five servers.
Case study: Runlayer’s enterprise-MCP bet. Runlayer — founded by Andy Berman, Tal Peretz, and Vitor Balocco — emerged from stealth with $11M in seed funding from Khosla Ventures and Felicis, explicitly to make enterprise AI “secure, compliant, and fully connected.” Its product runs enterprise MCP at scale, exposing 18,000+ MCP servers behind threat detection, fine-grained IdP-integrated permissions (Okta or Microsoft Entra), zero-trust architecture, and “complete observability.” The architectural tradeoff is stark: the alternatives — each agent holding its own OAuth credentials to each server, or a homegrown single-tenant proxy — either fail compliance or fail scale. Runlayer is the proof that enterprise buyers will pay for a purpose-built gateway rather than re-implement these primitives per team.
code
1WITHOUT vs WITH a gateway (5 servers x 3 agents)23 Without gateway With MCP gateway4 ------------------------------ ------------------------------5 15 OAuth implementations 1 credential broker6 15 credential stores 1 policy + allow-list point7 15 audit trails 1 unified audit log8 per-agent raw tokens (risk) IdP-checked, zero-trust calls910 Flavors: infra-only (routing only, you add auth/obs)11 full (auth + audit + rate-limit + policy in one)12 Rule of thumb: add a gateway once an agent uses > ~5 servers.
The build-vs-buy tension here is the same one from Lesson 1, sharpened: open-protocol-first stacks (MCP + a Speakeasy/Tyk-built gateway) win on portability and vendor optionality but make you compose auth + audit + policy yourself; vendor-covered stacks (Runlayer) win on time-to-compliance — auditability on day one — at the cost of optionality and adoption speed for new MCP features. Interview angle. “Would you use Runlayer or build the gateway?” → name the axis: a regulated buyer who needs audit logs and IdP integration on day one pays for the governance layer (time-to-compliance); a buyer who prizes portability composes the open primitives. The wrong answer is to pick one without naming the tradeoff.
The gateway is to MCP what an API gateway is to microservices: the one place credentials, policy, and audit live so no individual agent has to. Runlayer raising $11M to be exactly that — for 18,000+ servers behind Okta/Entra — is the market telling you enterprises will buy this primitive rather than rebuild it per team.
A note on how this composes with the rest of the stack: the gateway is also the natural enforcement point for the retrieval and tool guardrails from Lesson 1. Because every tool call passes through it, the gateway is where you apply the per-agent allow-list, the IdP-derived RBAC, the rate limit that prevents a denial-of-wallet loop, and the structured audit log that the observability layer (Lesson 3) ingests as spans. In a customer architecture diagram, the MCP gateway and the API gateway are the two choke points where governance is real — everything the model says about its own permissions is advisory; everything these two layers enforce is a control.
The operator checklist for an MCP rollout, which is exactly what you walk a customer through: (i) ship every internal connector as an MCP server with Resources/Prompts/Tools declared in a manifest; (ii) route all agent-to-tool traffic through a single gateway that holds credentials; (iii) wire the gateway into the customer’s IdP (Okta / Entra) so tool access inherits existing RBAC; (iv) emit a structured audit log for every tool call (model identity, tool name, args hash, outcome); (v) treat user consent and tool exec as security boundaries, not defaults; (vi) require an explicit allow-list of servers per agent in production.
A customer asks why you’d standardize on MCP instead of writing direct API clients for each system. Best answer?
AIt collapses integration from N×M bespoke glue to N+M standard endpoints — every connector is the same Resources/Prompts/Tools shape, so integrations stop explodingBMCP is faster at runtime than REST callsCMCP removes the need for authentication to backend systems
An agent in production needs to call eight internal MCP servers across two teams. What’s the right architecture for credentials and audit?
AGive the agent the OAuth tokens for all eight servers directlyBRoute all traffic through one MCP gateway that holds credentials, enforces a per-agent allow-list, inherits the customer’s IdP RBAC, and emits a unified audit logCSpin up a separate agent per server so each holds only one credential
A security reviewer asks what stops an agent from taking a destructive action via a connected tool. Strongest containment story?
AThe model is instructed to be careful with destructive actionsBRun the agent with broad admin credentials but log everythingCExplicit user consent for tool ops, a per-agent allow-list of servers, IdP-checked permissions, and an audit log of every call — consent and exec treated as boundaries
A regulated bank needs audit logs and IdP integration for AI tool access on day one. Build the gateway in-house or buy a governance product like Runlayer?
AAlways build in-house for maximum controlBBuy the governance layer for time-to-compliance (audit + Okta/Entra on day one); name the tradeoff that you give up some portability and new-feature adoption speedCAvoid MCP entirely and use direct integrations to dodge the question
You’re explaining MCP’s topology in an architecture round. Which description is correct?
AHost contains clients (one per server), each client connects to a server exposing Resources/Prompts/Tools, over JSON-RPC 2.0 via stdio or streamable HTTPBA single shared bus that all agents and tools publish to and subscribe fromCServers run the LLM and call out to clients for data
MCP and connector governance show up in the architecture round (“design an LLM-powered enterprise search system,” “connect the agent to these five systems”) and as deep probes on the security and integration axes. Strong answers draw the host/client/server topology, explain the N×M→N+M reframe, name the consent/exec security boundaries, and reach for a gateway with an allow-list and IdP integration past ~5 servers; weak answers wave at “it just calls the APIs” and ignore credential sprawl, consent, and audit.
01“What is MCP, architecturally?” → host (AI app) contains clients (one per server); servers expose Resources/Prompts/Tools; JSON-RPC 2.0 over stdio or streamable HTTP.
02“Why MCP over direct API clients?” → N×M bespoke glue → N+M standard endpoints; every connector is the same shape, so integrations stop exploding.
03“What are MCP’s security principles?” → user consent, privacy/safety (tool exec is arbitrary code), sampling control; enterprise adds IdP + zero-trust.
04“When do you add a gateway?” → past ~5 servers per agent; it holds credentials, brokers auth, enforces policy, and emits one audit log.
05“Infra-only vs full gateway?” → infra-only routes/translates (you add auth/obs); full bundles auth + audit + rate-limit + policy (e.g. Runlayer).
06“Build vs buy the gateway?” → buy for time-to-compliance (audit + Okta/Entra day one); build for portability — name the tradeoff, don’t just pick.
07“How do you contain a destructive tool call?” → consent + per-agent allow-list + IdP-checked permissions + audit log; never an open registry or raw tokens.
08“What goes in the tool-call audit log?” → model identity, tool name, args hash, outcome — per call, so every action is attributable.
Going deeper, expect: “stdio vs streamable HTTP — when each?” (stdio for a local process like a desktop file server; streamable HTTP for a remote service, with bearer/OAuth and optional SSE for streaming); “what’s Elicitation for?” (the client primitive that lets a server ask the user a clarifying question mid-flow — a consent/UX boundary, not just data); and “how do you stop indirect prompt injection through a tool’s output?” (treat tool/retrieved content as untrusted data, delimit it from instructions, and never let it expand the agent’s allow-list — the connector is an attack surface the observability layer must watch, tying back to Lesson 3).
Could you draw the MCP topology, explain N×M→N+M, name the security boundaries, and decide gateway build-vs-buy with the tradeoff?
New to itGetting thereConfident
Takeaways
MCP is host/client/server over JSON-RPC 2.0 (stdio or streamable HTTP) with a fixed primitive vocabulary.
Its value is uniformity: integration collapses from N×M bespoke glue to N+M standard endpoints.
Security is in the spec — user consent, tool-exec caution, sampling control; enterprise adds IdP + zero-trust.
Past ~5 servers, route through a gateway that holds credentials, enforces policy, and unifies the audit log.
Runlayer ($11M seed, 18,000+ servers) proves buyers pay for governance — buy for time-to-compliance.
Never let an agent hold raw tokens to an open tool set; scope an allow-list and audit every call.
Next: the other half of the job — running discovery and telling the value/ROI story that earns the right to demo.