AI agents, telephone tricks, and the old problem of letting the work control the rules.

There was a time when you could manipulate the telephone network by playing the right sounds into it. (quick shout out to the old timers who grew up on reading 2600)

Earlier telephone systems carried certain signaling tones in the same channels used for conversation. Those signals helped operate the network. Subscribers could reach that channel too, and phone phreaks learned how to reproduce what the switches expected to hear.

The caller was supposed to use the network. The architecture gave the caller a way to issue instructions to it. The Computer History Museum describes this in-band signaling weakness in its account of the blue box. (Historical account)

I keep thinking about that story when I look at AI agents.

An agent reads a document, visits a website, or processes a message. Somewhere in that material are instructions attempting to redirect its behavior. The author wants the agent to treat something it encountered as something it must obey.

Different technology. A familiar question: how did the content get a vote in running the system?

October is Cybersecurity Awareness Month. I would like to spend some of it revisiting fundamentals that each generation of technology seems to rediscover after something goes wrong.

This is the first one.

In a recent LinkedIn post, I called it the Law of Separation of Authority:

In a secure AI system, untrusted data must never acquire authority. Permissions and consequential actions must be governed by an independently enforced control plane that the model cannot override.

The model may interpret information and propose an action. It must not grant itself permission to execute that action.

Calling it a law does not make it a new discovery. Those of us who have spent decades in security have encountered this requirement under several names and in many different systems.

We should be getting better at recognizing it.

Consider an ordinary business payment.

Someone submits an invoice. Someone checks that the work was performed and the amount is correct. Someone with the appropriate authority approves payment. The payment system releases the funds under the organization’s rules.

There are reasons to distribute those responsibilities. One person’s mistake gets another opportunity to be caught. A compromised account has limited reach. A person attempting fraud needs to overcome controls they do not own.

The details vary, but the purpose is recognizable: avoid giving one participant unchecked power over the entire transaction.

Computer security formalized a related idea long ago. In their 1975 paper, Saltzer and Schroeder described separation of privilege, including mechanisms that require more than one condition to grant access. They also examined the authority to change access rules themselves. (Original paper)

Now imagine replacing parts of that payment process with agents.

One reads the invoice. Another reviews it. A third handles payment.

The workflow diagram looks responsible. Three boxes. Three roles. Perhaps three different models.

Then we inspect the implementation.

All three use the same powerful service account. The first agent can edit the records the reviewer relies on. The reviewer can change the payment limits. Any of them can invoke the tool that releases funds.

The diagram describes a division of labor. The permissions reveal how little authority we actually separated.

Giving agents different jobs does not, by itself, give them different powers.

That distinction matters because a role instruction is still an instruction interpreted by the model. “You are the reviewer” can help organize work. It does not establish a security boundary.

A boundary needs something that can refuse the action even when the agent decides the action is necessary.

Return to the telephone network.

Separating signaling from the voice channel, as systems such as SS7 did, addressed the particular weakness that allowed subscribers to inject network control tones through the conversation path. That did not make telephone signaling immune to abuse. The participants admitted to the signaling network still required appropriate trust and controls, an issue reflected in the ITU’s continuing work on signaling security.

The useful lesson is architectural: the channel carrying a conversation should not automatically carry authority over the network.

AI systems need an equivalent distinction, even though the implementation will look different.

Language models interpret language. A legitimate request, a quoted instruction and an attacker impersonating an administrator can all arrive as text. Improving the model’s ability to distinguish them is valuable. The surrounding system must still decide which actions are permitted.

An invoice might say, “Our banking details have changed.”

That is information to investigate.

It does not give the invoice’s author authority to change the supplier’s registered payment destination. It also does not give the agent reading it that authority.

The agent can propose a change. A separately controlled process must establish whether the change is valid and who may authorize it.

The difference becomes clearer when we put ourselves in the adversary’s position.

The adversary is a problem solver. Their problem is that we built something to stop them.

They will look for a path to the protected resource. They will also look for a path to the mechanism deciding what is protected.

Can they change the approval criteria? Replace the reviewer? Persuade the agent to request broader permissions? Find another tool carrying more powerful credentials? Modify the records that make a transaction appear acceptable?

The controls themselves become part of the attack surface.

This is what I mean by getting above the controls: obtaining influence over the decisions that constrain the work.

A system might deny an agent permission to make a payment while allowing that same agent to edit the policy defining when payment is allowed. The immediate restriction exists. So does a route around it.

We have to follow authority all the way to the ability to change authority.

For our payment agent, I would want the design to make that separation visible.

The agent processing invoices could read the necessary records and submit a proposed transaction. Its identity would have no permission to release funds or modify supplier banking details.

The approval process would examine the proposed transaction against protected records and the organization’s requirements. Any required human approval would identify the actual recipient, amount and destination.

The payment service would enforce those conditions when executing the transaction. Approval for one payment would not become permission for a different payment after the agent changed a field.

The credentials that move money would remain behind that service. There would be no alternate tool offering the agent a convenient way around it.

Changes to the approval rules would require a separately authorized administrative process. The agent doing the work could recommend an improvement. It could not install its own exception.

This still allows useful automation. Routine transactions can proceed under rules the organization established in advance. The separation concerns who sets those rules and what enforces them.

It also gives us something concrete to test.

Give the invoice agent a malicious document. Assume it follows every instruction in it. Can it release funds?

Have it claim that an executive approved an exception. Can that claim substitute for actual authorization?

Let it try another tool, delegate to another agent or modify the proposed transaction after approval. Does the boundary still hold?

Those tests tell us more than watching the agent complete a hundred ordinary invoices.

They also force responsibility into the open. The model provider controls some parts of the system. The application developer controls others. The organization deploying it decides which records, tools and business powers to connect.

Someone must own the enforcement point between an agent’s proposal and its effect on the world.

There is one complication I want to leave with you.

In the LinkedIn discussion, Andrew Bove (thanks buddy) pointed out that an attacker might shape an agent’s understanding without expanding its permissions. The agent could propose an allowed action for the wrong reasons, and the approver could receive the same misleading account. (Discussion)

That deserves its own article. Separate authority does not automatically produce independent judgment.

But it makes the first requirement no less necessary.

For this October, pick one consequential action your agent system can take. Follow it from the information the agent reads to the change it makes in the world.

Identify who proposes it, who authorizes it, what enforces the decision and who can change those arrangements.

Then assume the agent has been persuaded to do the wrong thing.

What still stops it?

We should be able to point to the answer in the architecture.

Mahalo for reading!

Aloha –TK