OTWopentechwire
Tech Intelligence, Openly Wired
AI

Trust Between AI Agents Is Creating a New Attack Surface

Proof-of-concept exploits show how one compromised agent can weaponise the Model Context Protocol to spread malicious instructions across enterprise networks

MT
Mei-Lin Tan
Asia Tech Correspondent · Singapore
Oct 7, 2026
5 min read
Trust Between AI Agents Is Creating a New Attack Surface
Credit: Getty Images

A Growing Blind Spot in Enterprise AI

As organisations deploy AI agents to handle translation, data analysis, and business process automation, a structural vulnerability is emerging in how these systems communicate. Independent security researcher Syed Anas Mohiuddin has documented exploits targeting agents from Google, JP Morgan Chase, Weviate, Rapid7, France's interministerial digital directorate, and US federal agencies. The common thread is not the organisations themselves, but their reliance on a communication standard called the Model Context Protocol.

The attacks demonstrate how trust relationships between agents, designed to enable seamless collaboration, can instead become conduits for malicious instructions. Over the past five months, five organisations have acknowledged vulnerabilities that allow a single compromised agent to propagate harmful commands across internal networks.

How Agents Betray Each Other

The technique exploits a particular form of prompt injection. Rather than targeting the large language model directly, attackers focus on specialised agents within an organisation's network. A translation agent or analytics module, for instance, may have minimal guardrails governing what instructions it accepts. When such an agent receives malicious input, it processes the command and passes it along to other agents in the workflow.

The receiving agent, configured to trust internal communications, executes the instructions without scrutiny. The result is a cascade effect where compromise spreads laterally through an organisation's AI infrastructure. Mohiuddin's proof-of-concept attacks extracted database contents and sensitive business information by chaining together these trust relationships.

At Opentechwire, we've tracked the growing complexity of AI security as models move from isolated deployments to interconnected agent networks. The shift introduces dependencies that mirror traditional IT infrastructure but with less mature security tooling.

What Makes MCP Vulnerable

The Model Context Protocol was designed to standardise how AI applications and agents exchange information within enterprise environments. Its purpose is interoperability, allowing different agent types to share context and coordinate tasks. That design goal, however, assumes a trusted environment where all agents operate under consistent security policies.

In practice, organisations deploy agents with varying levels of hardening. A data extraction agent might enforce strict validation on external input but implicitly trust messages from another agent tagged as internal. Mohiuddin's research shows that attackers can exploit this asymmetry by first compromising a weakly protected agent, then using it as a launchpad for instructions that would normally be blocked if they arrived from outside the network.

The protocol itself does not mandate authentication or validation mechanisms between agents. This is a structural choice rather than an oversight; MCP prioritises low-latency communication and ease of integration. The trade-off is that security becomes the responsibility of individual agent developers, who may not anticipate cross-agent attack vectors.

Scope Across Sectors

The vulnerability landscape spans financial services, government, and technology vendors. JP Morgan Chase's exposure suggests that even organisations with substantial security budgets face challenges securing agent-to-agent channels. Rapid7, a security firm, acknowledged flaws in its own agent implementations, highlighting how the issue cuts across expertise levels.

France's interministerial digital directorate and unspecified US federal agencies also confirmed vulnerabilities, raising questions about how widely MCP has been adopted in sensitive government workflows. The fact that these acknowledgements emerged within a five-month window suggests coordinated disclosure rather than isolated incidents.

Weviate, a vector database provider, represents another category of risk. Agents that interact with data stores can be leveraged to exfiltrate not just the information they were designed to access, but adjacent datasets accessible through chained queries. The trust model breaks down when an attacker controls the query logic itself.

Why Traditional Defences Fall Short

Conventional security controls such as network segmentation and input sanitisation are less effective against this class of attack. Network segmentation assumes threats originate externally, but agent-to-agent communication happens within trusted zones. Input sanitisation typically focuses on user-facing interfaces, not on messages from peer systems.

Guardrails implemented at the LLM level also miss the mark. These controls govern what a model will generate or respond to, but do not address instructions embedded in inter-agent messages. An agent designed to summarise documents might have robust filters for user prompts but none for instructions received from a translation agent upstream.

The difficulty in mitigation stems from the implicit trust model. Agents are deployed to collaborate, and collaboration requires some degree of assumed reliability. Introducing strict validation at every handoff risks breaking workflows or introducing latency that undermines the agent's value proposition. Organisations face a choice between friction and risk.

What Enterprises Can Do Now

Immediate steps include auditing which agents communicate with each other and mapping trust relationships. Many organisations deploy agents incrementally without a unified view of how they interconnect. Visibility is the first requirement for control.

Implementing authentication between agents, even within internal networks, can limit lateral movement. If each agent must verify the identity and authorisation level of any peer before accepting instructions, the attack surface shrinks. This approach adds overhead but is more feasible than rewriting guardrails for every agent type.

Rate limiting and anomaly detection can also help. If a translation agent suddenly begins issuing database queries or accessing file systems outside its normal pattern, that deviation should trigger alerts. Behavioural baselines are harder to establish for agents than for human users, but the principle is the same.

Longer term, the industry may need to revisit MCP's design assumptions. A protocol that prioritises ease of integration over secure-by-default communication will continue to generate vulnerabilities as adoption scales. Standards bodies and vendors should consider mandatory authentication, message signing, and context-aware permission models as part of the specification rather than optional extensions.

The Broader Implication for AI Infrastructure

The vulnerabilities Mohiuddin documented are symptoms of a larger architectural question: how much trust should autonomous systems extend to each other? As AI agents take on more consequential tasks, the blast radius of a compromised agent grows. An agent with write access to financial systems or healthcare records can cause harm far beyond data leakage.

The enterprise AI stack is evolving faster than the security frameworks designed to govern it. MCP is one protocol among many, but it illustrates a recurring pattern where interoperability and speed-to-market take precedence over threat modelling. The organisations that acknowledged these flaws deserve credit for transparency, but the disclosures also reveal how widespread the issue has become.

For now, the onus is on individual organisations to secure their agent deployments. That burden is unlikely to produce consistent outcomes, particularly in sectors with limited security resources. Until the protocol layer itself incorporates stronger trust verification, attackers will continue to exploit the gap between how agents are intended to collaborate and how they actually behave under adversarial conditions.

Read next
AI

Agentic AI Forces Enterprises to Rethink the Entire Analytics Stack

Linh T. Pham · 6 min
AI

Model Iteration Cycles Shrink as AI Systems Accelerate Their Own Development

Arjun S. Mehta · 6 min
Startups

Asia's Largest AI Exits Went Through Exchanges. America's Went Through Acquirers

Valerie Nguyen · 10 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.