OX Security resolved 5,095 hostnames behind 15,465 published MCP servers and found foreign infrastructure, home networks and abandoned domains. Here is the vetting, allowlist and gateway process that actually holds up.
Last Updated: September 28, 2026
OX Security resolved 5,095 unique hostnames out of 15,465 published MCP servers and found 15.6% sitting outside the US, roughly 2.3% no longer resolving at all, and six abandoned domains you could buy for $4–12 a year. The protocol has no concept of where a tool runs. That means your vetting has to happen at the hostname, the domain registration and the gateway — not at the marketplace listing. This is the process: resolve and re-resolve every endpoint, pin agents to a self-hosted or first-party gateway instead of a marketplace URL, scope approvals per tool rather than per server, and stop treating model refusal as a control.
MCP security coverage has been stuck on the protocol for a year: token passthrough, confused deputy, the odd CVE in a popular server. All real, all fixable in the spec, and none of them address what actually breaks a production agent stack — the server behind a published listing is an unverifiable moving target. You read the source on GitHub. The endpoint runs something else. Nobody checks.
On 24 September 2026, OX Security put numbers on that gap. Their research team analysed 15,465 published MCP servers drawn from three public registries — the official MCP registry, the Cline marketplace and the GitHub MCP registry — narrowing to 5,095 unique hostnames for infrastructure analysis. What they found is not a protocol bug. It is a supply chain with no chain of custody.
What 5,095 hostnames actually tell you
The headline figure is geography. The useful figure is resolution — and every check below is one you can run today against your own installed server list.
The abandoned-domain number should move your roadmap. 2.3% of 5,095 is roughly 117 hostnames still sitting in someone's mcp.json, still dialled by a client on startup, pointing at nothing. Six were free to claim. An attacker who registers one needs no vulnerability: they inherit the trust relationship, serve a tool catalogue of their choosing, and any client with a stored approval for that server name starts handing over arguments. It is the abandoned-package pattern with a worse blast radius, because the caller is an agent with credentials rather than a build step.
The code-to-runtime gap
For a remote MCP server, the published repository and the running binary are two different artefacts with no cryptographic link between them. You can audit the code perfectly and still be talking to a fork, a modified build, or a different operator entirely after a domain changes hands. Local stdio servers do not have this gap — you run what you pulled. Everything remote does.
Why "percentage outside the US" is the wrong residency question
The 15.6% framing is a US-centric read of a global dataset, and if you are procuring in the GCC or the EU it tells you almost nothing. Your requirement is not "minimise foreign infrastructure" but a positive statement you can evidence: personal data under this workflow stays in the UAE, or the EEA, or an approved adequacy jurisdiction. A server resolving to Frankfurt is "outside the US" and perfectly fine for a Dublin controller. A server in Virginia is inside the US and a problem for a bank in Riyadh with a localisation condition in its licence.
The operative point in the OX research is not the ratio. It is this: MCP provides no native protocol mechanism to enforce where connected tools run or where data may be processed. There is no region field in the handshake — and after the 2026-07-28 specification retired the handshake entirely in favour of a stateless request/response core, there is not even a handshake to put one in. Residency in MCP is an infrastructure property, enforced outside the protocol or not at all.
So the only defensible residency control is an egress control. You do not ask "is this server compliant?" You assert: agents in this environment may only reach these endpoints, which resolve to these regions, re-checked on a schedule. Everything else is a vendor questionnaire answered by someone with no visibility into where their own tunnel terminates.
The vetting workflow that actually survives an audit
Here is the sequence we use when scoping agent tooling. It is deliberately boring, and the order matters: three of these steps will disqualify a server before anyone reads its README.
MCP server due diligence, in order
Enumerate every MCP endpoint in client configs, CI runners and developer machines. This list is routinely far longer than the one the platform team maintains, because IDE-level installs never went through procurement.
Not the vendor's marketing page — the A/AAAA record. Capture ASN and geolocation. Flag consumer ISPs, residential blocks and tunnelling ranges. That is the 0.45% bucket: a hard fail for anything touching production data.
WHOIS the registrable domain, record expiry, and monitor it alongside your TLS certificates. A tool endpoint whose domain lapses in 60 days is a scheduled incident, not a risk.
Self-hosting from a pinned container digest removes steps 2 and 3 as ongoing concerns. If the server is only reachable as a marketplace-hosted remote URL, code review buys less than you think.
Classify each tool by the worst thing it can do with the credentials you plan to issue. Filesystem, shell, outbound HTTP and arbitrary-query tools get separate treatment from read-only lookups.
The allowlist entry names the server, the specific tools, the gateway route and the identity the calls run as. Anything not on the list is denied at the gateway, not at the client.
Weekly for most estates, daily if the agent has write access to customer systems. Alert on any change in resolved ASN, country or registrar. That is your domain-takeover tripwire.
Step 7 is the one nearly everyone skips, and the only one that catches the failure the OX data describes: a server that passed review in March and changed hands in August looks identical in your config file.
Pin to a gateway, not a marketplace URL
The decision that collapses most of this risk is refusing to let clients dial arbitrary endpoints at all. Agents connect to one internal gateway; the gateway holds the allowlist, the credentials and the egress policy; upstream servers are either self-hosted in your VPC or first-party endpoints run by a vendor you already have a DPA with.
| Hosting pattern | Who controls the running code | Residency provable? | Primary failure mode |
|---|---|---|---|
| Marketplace-listed remote URL | Unknown third party, can change without notice | No — DNS can repoint at any time | Domain expiry or transfer; silent build divergence from the published repo |
| First-party vendor remote (e.g. the SaaS you already buy) | The vendor, under contract | Partially — via contract and regional endpoints, not via the protocol | Regional endpoint quietly failing over to another region |
| Self-hosted container, pinned by digest, in your VPC | You | Yes — it runs where you deployed it | Stale images; nobody owns patching the third-party server code |
Local stdio server on a developer machine | The developer | Irrelevant, but the data still leaves via the tool's own outbound calls | Shadow installs; no central inventory; credentials in dotfiles |
Self-hosting has a real cost: you now own the patch cadence for other people's server code. That is a staffing decision, not a technical one, and it is why most teams land on a two-tier policy — self-hosted for regulated data, first-party vendor endpoints for everything else, marketplace URLs never. Sizing that trade-off for a specific stack is the scoping work our AI integration practice does before any build starts.
The 2026-07-28 spec release makes the gateway pattern substantially easier than it was a year ago. Method and tool names now travel in the Mcp-Method and Mcp-Name HTTP headers, so gateways can route and authorize on headers directly without parsing the JSON-RPC body. That is a genuine enabler: per-tool policy at the proxy is now a header match rather than a body inspection, and a stateless core means the proxy does not have to maintain session affinity.
POST /mcp HTTP/1.1
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
# gateway policy can allow/deny on Mcp-Name before the body is parsed
"Always-Allow" is a per-server grant wearing a per-tool costume
The most instructive finding in the OX research is the lab test, not the census. Against Claude Code paired with Haiku 3.5, a malicious MCP server first requested access to a harmless file. The user approved with always-allow. The server then requested a sensitive file, .env among them, and got it with no further prompt. The same attack failed against Opus 4.6 and 4.7.
This is not a vulnerability in the usual sense. OX reports Anthropic's position as: once always-allow is granted, that is the documented behaviour, and model-level detection of malicious content is a best-effort heuristic rather than a security boundary. That is the honest answer. A consent UI saying "always allow this tool" records a decision about a tool, but the trust it establishes is with a server that controls what that tool does on every subsequent call. The user approved a filename. The system recorded an approval for a capability.
Granting Always-Allow interactively and letting it persist into automated runs
A developer clicks always-allow during an exploratory session to stop the prompt fatigue. That grant lands in a settings file, gets committed or synced, and then applies to unattended agent runs where no human is present to notice the second, worse request. The approval was made in a context with a human in the loop and is now being spent in a context without one.
The model-tier result is real and you should use it carefully. Teams routinely route the cheapest capable model to the highest-privilege step because that step is short. Inverted: privilege should drive tier selection at least as much as token volume does, and if you are mapping that trade-off across providers our model coverage notes are the starting point.
But a probabilistic refusal is not an access control. Haiku 3.5 failing and Opus 4.7 passing on one attack chain tells you nothing about the next chain. Build the deterministic control — the gateway, the per-tool allowlist, the scoped credential — and treat model tier as the layer that catches what the policy missed, never the layer that makes policy unnecessary.
What the protocol gives you now, and what it still does not
MCP's security posture is improving faster than its reputation suggests. The 2026-07-28 release landed a set of authorization hardening changes: authorization servers returning the iss parameter per RFC 9207 with clients required to validate it before redeeming a code, client credentials bound to the issuer that minted them, and Dynamic Client Registration formally deprecated in favour of Client ID Metadata Documents. The spec's own security guidance is explicit that token passthrough is forbidden — a server must not accept tokens that were not issued to it and forward them downstream.
Enterprise-Managed Authorization is now a stable extension and is the closest thing to a native answer for allowlisting. It puts your IdP in the path: the client exchanges an ID token for an Identity Assertion JWT Authorization Grant, the IdP evaluates policy before issuing it, and clients never receive a token for an unapproved server. Centralised revocation, centralised policy, immediate effect. If your MCP clients and servers both support it, adopt it.
Spec 2026-07-28 ships: stateless core, header-based routing, RFC 9207 issuer validation, DCR deprecated in favour of CIMD, twelve-month deprecation policy.
New roadmap names agent identity and enterprise-ready security as a priority area, with DPoP, Workload Identity Federation and RFC 8693 token exchange as the work items.
OX Security publishes the 15,465-server census — the first ecosystem-wide measurement of where MCP tool endpoints actually run.
What the protocol still does not give you: any binding between a published server listing and the code running at the endpoint, any region or residency attribute, and any revocation path when a domain changes owner. The Server Card Working Group is defining a standard metadata document so a server can be discovered and reasoned over without connecting to it, which will help discovery hygiene considerably. It is metadata the server publishes about itself. It is not attestation, and you should not plan a control around it yet.
Where the honest answer is "not yet"
Three cases where we would not connect an agent to a third-party MCP server at all, regardless of how good the vetting process is.
Regulated data with a hard localisation condition and no self-hosting option. If the server is only available as a vendor-operated remote and your licence conditions require in-region processing, no amount of DNS checking makes that compliant. Either the vendor ships a container you can run in-region, or the capability waits. The percentage-outside-US framing is a distraction here; what your regulator wants is a positive assertion about a named jurisdiction, and a marketplace endpoint cannot give you one.
Agents with standing write access to financial or customer systems, unattended, in the first quarter of a deployment. Not because the technology cannot do it, but because you have not logged enough tool calls to know what normal looks like. Run propose-and-approve until the distribution is boring, then remove the human. That sequencing separates automation that survives production from pilots switched off after an incident.
Anything where the tool endpoint sits behind a consumer tunnelling service. The 0.45% bucket is small, but there is no version of this that is acceptable for production. No uptime guarantee, no enterprise access control, no audit trail, and a residential IP that can be reassigned. If a vendor's MCP server resolves to a tunnel, that is a prototype someone forgot to tear down.
Security is already the top MCP adoption obstacle in the Stacklok survey, cited by 64% of software respondents. The answer is not slower adoption. It is vetting cheap enough to run continuously — a resolution check, a WHOIS expiry check, a diff against last week — so connecting a server is a reversible decision. If you want that as a standing control rather than a spreadsheet, start here.
Frequently Asked Questions
What are the main MCP server security risks in 2026?
Supply chain, not protocol. OX Security's analysis of 15,465 published servers found 15.6% of 5,095 hostnames resolving outside the US, 0.45% on home networks or consumer tunnels, and 2.3% no longer resolving — including six domains available for $4–12 a year that anyone could register and impersonate. Add the code-to-runtime gap: a remote server's published repository has no cryptographic link to the binary actually answering requests.
How do I enforce MCP data residency for EU or GCC requirements?
Outside the protocol. MCP has no native mechanism to enforce where connected tools run or where data is processed, and the 2026-07-28 stateless core removed the handshake that might have carried a region attribute. Enforce it as egress policy: agents reach only an internal gateway, the gateway allowlists endpoints that resolve into approved regions, and resolution is re-checked on a schedule so a silent repoint triggers an alert.
Is "Always-Allow" safe if I only approve harmless tools?
Not on its own. In OX's test, a malicious server got an always-allow grant for a harmless file, then used prompt injection to read a .env file with no further confirmation on Claude Code with Haiku 3.5. The grant records trust in a server that controls what its tools do next. Scope approvals per tool, expire interactive grants, and load unattended runs from a reviewed allowlist in source control with no escalation path.
Does using a stronger model protect against MCP prompt injection?
It helps and it is not a boundary. The same escalation that succeeded against Haiku 3.5 failed against Opus 4.6 and 4.7, which is a good argument for letting privilege — not just token cost — drive tier selection for agents holding write credentials. But model refusal is a heuristic that generalises poorly to the next attack chain. Build deterministic gateway and allowlist controls first; use model tier as defence in depth.
Should we self-host every MCP server?
Self-host anything touching regulated data, use first-party vendor endpoints where you already hold a data processing agreement, and avoid marketplace-hosted URLs for production. Self-hosting from a pinned container digest removes DNS repointing and domain expiry from your threat model, but transfers patching of third-party server code to your team. That is a staffing commitment, so decide it deliberately rather than per-integration.
What does Enterprise-Managed Authorization change for MCP governance?
It moves the allowlist into your identity provider. The client exchanges an ID token for an Identity Assertion JWT Authorization Grant, the IdP evaluates policy before issuing it, and clients never receive tokens for unapproved servers. You get centralised policy, single sign-on across servers and immediate revocation. It is a stable MCP extension, so adoption depends on both your client and the target server supporting it.
How often should we re-vet MCP servers we already approved?
Weekly for most estates, daily where agents hold write access to customer systems. Re-resolve every allowlisted hostname, capture ASN, country and registrar, record apex-domain expiry, and alert on any change. A server that passed review six months ago and has since changed hands looks identical in your config file — the resolution change is the only signal you get.
Research digest
AI Research Briefing
Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.
- No hype, just measurable outcomes
- Read by 2,400+ engineers
- Unsubscribe anytime

