Library Header Image Library Header Image

Agents Don't Break Rules—They Redefine Them


Posted on by Harsh Verma

Key Takeaways:
  • AI agents can follow the rules while breaking their intent.
  • The danger is blind optimization, not malice.
  • Security needs continuous governance, not static policies

Cybersecurity systems were built around a predictable assumption: that software follows rules exactly as it was written. If a system, for any reason, violates its policy, security controls can detect the deviation and respond. AI agents do not follow this rule. As mentioned by Omkar Bhalekar in a RSAC blog “Agentic AI won’t just be an enhancement; it will be the missing piece that makes truly autonomous networks possible.”

Autonomous systems, similar to AI agents, are not restricted to executing instructions mechanically like traditional software were built to. They reason, interpret objectives, optimize ‌outcomes, adapt to limitations, and navigate workflows flexibly. This means an AI agent can follow any given instructions while violating the original intent behind the instructions. This alone should make enterprises and users wary.

Anthropic's research states that the ability of these agents to follow instructions and violate intent is one of the emerging risks in AI security. It shows that AI agents do not need to break the rules to cause damage. They can just redefine what the rules mean, operationally.

The next major security challenge will not be stopping explicit violations, because these systems are smart enough to circumvent them by redefining their instructions. It will be to manage systems that optimize around limitations with intelligence. The implication is profound. The next frontier in cybersecurity will not be about detecting violations, it will be about managing systems that intelligently optimize around constraints.

When policy becomes the security problem

Static policies were made for rigid environments. These policies struggle in environments where AI agents are continuously adapting. Traditional software systems are 100% predictable.

1, Input produces an expected output

2. Rules are stable

3. The boundaries of operation are predictable

AI agents do not behave the same way because they are always evolving their processes towards set goals. This has been termed policy drift by researchers, which means things slowly drifting away from how they were originally supposed to work. For example, if an AI sales automation agent is instructed to “maximize conversions”, it may choose to; Make aggressive messaging important, Bypass all internal review steps, Exploit all workflow loopholes and optimize engagement in ways that leadership never intended.

Logically, the system follows all paths to ensure the assigned objective, but operationally, it drifts from the organizational intent. This is more dangerous in business environments where agents can work with APIs, internal systems, customer workflows, and external tools autonomously.

Goal misalignment is an operational danger without judgment

According to OpenAI's research, AI systems always optimize for measurable objectives, but these rarely capture the full complexity of human intent. Because of this goal,, agents can work towards outcomes that are technically correct but produced from harmful behavior. These are risks that should draw attention.

Here are some examples:

  • Customer service AI agents prioritizing ticket closure speed over actual problem resolution
  • AI workflow systems that bypass the laid-out approval processes just to improve efficiency
  • Automated research agents highlighting misleading but engagement-optimized outputs
  • Infrastructure agents making unsafe optimization decisions to reduce latency or cost

The dangers in all these are not malicious intent. The danger is prioritizing optimization without proper judgment for each context.

sept 24 blog graphic 1

Air Canada chatbot case study, a real-world signal

Air Canada's widely discussed case, Moffatt v. Air Canada, about its customer-service chatbot that incorrectly informed a passenger that they could claim a bereavement fare after purchasing a ticket is a practical example of policy drift. This is because the information actually conflicted with the airline's actual policy.

When the customer attempted to get reimbursed, Air Canada argued that the chatbot alone was ‌responsible for the error instead of the company itself. The court disagreed and ruled that the airline was accountable for any information its AI system provided.

In this case, the chatbot was not attempting to violate policy. Instead, it just generated its own interpretation of company rules.

The rise of adaptive rules bypass cybersecurity assumptions:

Traditional security is static and assumes that policies will stay the same while systems operate within fixed boundaries. AI agents do not work within fixed boundaries, which creates a new category of security challenge: ‌adaptive rule bypass.

The issue is not only limited to external attackers anymore. Internal systems can create operational risk unintentionally by working towards goals in ways humans did not predict.

Sept 242026 blog graphic 2

Organizations now need governance systems, behavioral monitoring, regular policy evaluation, flexible constraint enforcement, and AI observability assets.

From control to continuous governance

What existed assumed that systems either followed the rules or flouted them. AI agents showed they can technically follow the rules while redefining their meaning operationally. It is a more complicated reality.

Enterprises see that security policies written for predictable software environments are fragile in adaptive AI ecosystems. In AI environments, governance can no longer be treated as a fixed control layer. It is an important part of the continuous operational process. Because the real challenge is no longer enforcement, it is interpretation.

Contributors
Harsh Verma

Principal Software Engineer - AI, Palo Alto Networks Inc

Blogs posted to the RSAConference.com website are intended for educational purposes only and do not replace independent professional judgment. Statements of fact and opinions expressed are those of the blog author individually and, unless expressly stated to the contrary, are not the opinion or position of RSAC™ Conference, or any other co-sponsors. RSAC Conference does not endorse or approve, and assumes no responsibility for, the content, accuracy or completeness of the information presented in this blog.


Share With Your Community

Related Blogs