Library Header Image Library Header Image

Why Multi-Agent Systems Are a Cybersecurity Nightmare


Posted on by Harsh Verma

Key Takeaways:
  • Multi-agent AI expands the attack surface. One compromised agent can spread malicious instructions through trust chains.
  • Agent-to-agent manipulation is an overlooked risk. Prompt injection and corrupted outputs move invisibly between agents, so human approval should gate critical actions.
  • AI security must protect the whole ecosystem. The vulnerabilities come from trust between agents, not from any single model.

Single-agent systems are still the focus of AI discussions. One assistant, one workflow, and one model performing one task. This has changed. Multi-agent systems are being deployed by organizations where AI agents work together independently across tools, APIs, assets, and workflows.

Traditional IAM was built for people logging into applications. Agentic IAM has to be built for autonomous software acting on people’s behalf,” wrote Gaurav Sharma, Senior Security Software Engineer at Nvidia in an RSAC blog.

With multi-agent systems, one agent retrieves the information; the second does the analysis; a third executes actions, and a fourth validates the results. In operations, this maximizes efficiency, but from a cybersecurity perspective, it creates a new category of risk.

According to the MITRE ATLAS Framework, when agents start to trust other agents, it creates a massive web of hidden weak points that hackers can easily exploit, trick, or take control of.

My view on this is that the biggest future AI security threats will not come from one compromised model but from chains of agents amplifying mistakes, manipulation, or malicious instructions across systems. What we are witnessing is not an incremental shift in risk. It is the emergence of an entirely new attack surface.

Trust chains create larger risks

Different areas of business software environments are often kept separate. Applications have fixed permissions, workflows are predictable, and communication is controlled. With multi-agent systems, things are much different because they are flexible, exchange information, share tasks, and coordinate decisions independently.

This is how a trust-chain is built where one compromised agent can influence others in the system.

This is dangerous because:

  • Most agents learn and inherit trust from other connected systems
  • Malicious instructions can be spread across live operations
  • Any compromise spreads across all workflows rapidly
  • Systems also amplify errors automatically

This means a compromised research agent, for example, can feed manipulated information to a planning agent, trigger flawed recommendations for operational use, or cause downstream systems to execute on corrupted results.

With a single malicious agent, the danger emerges and spreads because no single system appears compromised on its own.

The overlooked threat: agent-to-agent manipulation presents an additional security risk

Out of all the risks discussed in this sector, one of the least discussed in AI security is that agents may become targets to be manipulated by other agents. Because modern systems rely on tool integrations, APIs, shared memory, and workflow layers, threat-level assets like hostile prompts, corrupted outputs, or manipulated contexts can move between systems almost invisibly. Any modification or deletion operation by AI agents should be gated and approved by a human.

According to OWASP Top 10 risk mitigations for LLMs, researchers have shown how malicious prompt attacks easily spread through connected AI workflows, influencing downstream agents to behave in a subtle but compromising manner.

Other attack patterns apart from prompt injection also include:

  • Corrupted memory sharing
  • Manipulated tool outputs
  • Workflow abuse

Blog Graphic One October Six

My perspective on this is that with multi-agent systems, attackers do not need to crash infrastructure directly. By manipulating agent relationships, they can create widespread errors.

Stanford and UC Berkeley case study: evidence from emerging research

Researchers from Stanford, UC Berkeley, and other institutions have shown how prompt injections can spread through interconnected AI workflows.

The research’s aim was to show how trust chains can create new attack areas. As many organizations increasingly deploy multiple cooperating agents, gaps are created from interactions rather than individual systems.

In these scenarios, one compromised agent introduces malicious instructions into a workflow. Downstream agents trust the information they receive and continue acting on the manipulated instructions, effectively propagating the attack across the system.

No single model failure is required. The vulnerability emerges from the trust relationships between agents.

Traditionally, cybersecurity has focused on protecting singular systems, networks, and identities. Multi-agent AI environments are not the same because trust now propagates among autonomous systems. This is a sharp change from isolated-system security to ecosystem security, carrying its own risks. Businesses now know that AI agents do not operate independently for long. They collaborate continuously.

My view on this is that the next major frontier in cybersecurity will not involve just protecting machines from attackers. It will be how to protect machines from other machines operating inside a trusted ecosystem.

Contributors
Harsh Verma

Principal Software Engineer - AI, Palo Alto Networks Inc

Blogs posted to the RSAConference.com website are intended for educational purposes only and do not replace independent professional judgment. Statements of fact and opinions expressed are those of the blog author individually and, unless expressly stated to the contrary, are not the opinion or position of RSAC™ Conference, or any other co-sponsors. RSAC Conference does not endorse or approve, and assumes no responsibility for, the content, accuracy or completeness of the information presented in this blog.


Share With Your Community

Related Blogs