The Brakes Are Not Connected


Whenever the AI discussion gets sufficiently alarming, somebody eventually reaches for one or both of two solutions: slow it down and/or keep a human in the loop. Both sound sensible. Slow things down and we buy time. Keep a person involved and the machine remains subordinate to human judgment. The grown-ups are still in the room, someone has a hand near the kill switch, and nobody has quite handed the keys to HAL.

I’m in favor of both. I just don’t think either is actually a brake.

Slowing Development

Slowing development would certainly help. If capability is moving faster than our ability to understand, monitor, and regulate it, more time is better than less. Testing takes time. So does finding out whether the astonishing new thing everyone wants to ship has some thoroughly unpleasant property nobody noticed while admiring the benchmark scores.

The problem is keeping things slow. A model that does more is worth more, and one that needs less supervision costs less to run. Corporations, governments, militaries, universities, and open-source communities don’t share the same goals or appetite for risk. If one lab slows down, another may decide Christmas has come early.

And what exactly are we waiting to catch up? As models become more capable, everything around them gets more complicated. They get more tools, memory, permissions, instructions, and autonomy. A serious agent may be trying to reconcile its system instructions with developer instructions, the user’s request, retrieved information, application state, tool permissions, memories, and safety policies. When something unexpected happens, we add another instruction, filter, or permission check. We’ve closed a path we know about while changing a system that contains a great many paths we don’t. A system prompt isn’t code in the usual sense. It’s more like terrain. The model is trying to get somewhere and is increasingly good at finding its way through it.

A Human in the Loop

Putting a human in the loop sounds more reassuring until you ask what the human is actually doing.

If a cyberattack is unfolding in milliseconds, nobody is waiting for Bob in SecOps to examine every alert. The defensive system acts. Bob sets policies and thresholds and, when something goes badly wrong, ends up in a room with six other people trying to establish what the hell the system actually did. Or suppose a model works through a million lines of code, logs, dependencies, and configuration data and recommends a fix. A human engineer gets the Approve button.

What, exactly, is the engineer approving?

Unless the engineer can reproduce enough of the analysis to judge it independently, he’s really deciding whether to trust the machine. Give him thirty seconds because something important is on fire and there isn’t much judgment left in the process. And this gets worse as the technology succeeds. If one expert with AI can do the work that used to require ten, somebody in management will notice the other nine salaries sooner or later. Over time, fewer people will acquire the experience needed to challenge the machine. You can still have a human in the loop while steadily getting rid of the humans who know enough to say, “No. That is bullshit.”

A tired junior employee with thirty seconds and an Approve button is a human in the loop. So is a specialist with the evidence, time to examine it, and authority to stop what happens next. The phrase covers rather a lot of territory.

 What We Need Is a Brake

An AI can inspect a production network without being able to reconfigure it. It can find a vulnerability without having the credentials to exploit it. It can recommend terminating somebody’s access without being able to terminate it. We don’t have to give a system the power to do something simply because we’ve given it the intelligence to work out how.

Now we have something resembling a brake.

Telling an agent “do not access external systems” is an instruction. Not giving it network access is a brake. Telling it not to spend more than $10,000 is an instruction. Setting a hard transaction limit in the payment system is a brake. Asking Bob to watch what it does is supervision. Putting the authorization somewhere the AI can’t reach gives Bob something he can actually stop.

Adam Engst, after reading an earlier draft, suggested that perhaps the important human isn’t the human in the loop but the human outside the loop. I think that’s right.

Let the AI reason, plan, try things, fail, and try something else. Sooner or later, something it wants to do has consequences outside the environment it controls. That boundary is a much more useful place to put Bob.

This is mostly old-fashioned security engineering. Software fails. Credentials leak. Things get compromised. People make mistakes. That’s why we have least privilege, separate authority, independent logs, and systems designed to fail closed. Yet we’re building AI agents and trying to control them with instructions written in English, backed up by somebody clicking Approve.

Cory Doctorow has made a related argument about the tendency to describe an AI agent escaping its intended environment as though the machine had developed a taste for freedom. I’m with him on the mysticism. It doesn’t need to be conscious, malevolent, or pissed off about being trapped in a sandbox. Give it a job, tools, and enough latitude, and it may find a way of doing the job that nobody expected, so telling the hacking agent not to leave the sandbox probably won’t work. The answer? Build a sandbox it can’t leave.

If it can change the rule, acquire the missing permission, hide what it tried to do, or rewrite the record afterward, the sandbox wasn’t much of a sandbox.

There’s a human version of the same problem. The protection works perfectly well until people decide it’s a pain in the ass. Diane Vaughan called this the normalization of deviance in her study of the Challenger disaster. Something happens that isn’t supposed to happen, but nothing terrible follows. It happens again. Still nothing terrible. Eventually, what was once evidence of a problem becomes evidence that the system can tolerate the problem.

It isn’t difficult to imagine the AI version. A company requires human approval. The approvals start holding things up, so the company lets the AI handle a few more cases by itself. That works fine. Six months later it expands the AI’s authority again. Then again. Nobody decides to surrender control. The controls just become increasingly annoying, and nothing bad has happened yet.

Another problem here deserves its own essay. If the system finally does cause serious harm, who owns it? The company that made the model? The people who built the agent around it? The company that deployed it? Bob, who told it what he wanted and had no idea how it was going to get there? Making somebody pay for the damage or go to jail would certainly concentrate minds, and deciding who gets the bill is going to be interesting.

There’s also a cost to keeping the brakes connected. Redundancy is expensive. If AI enables one person to replace ten, keeping some of the other nine around because you may someday need people who can check the machine will look woefully inefficient. So will restrictions that prevent the AI from doing things it is perfectly capable of doing.

That’s precisely when the brakes matter. Brakes don’t make cars more efficient. They exist because efficiency is a piss-poor consolation when you can’t stop.

I’d want to know whether a restriction still works when the AI ignores its instructions. Can the model change the permission or erase the evidence that it tried? Can it get to the thing that restricts it? And when somebody decides the restriction is costing too much time or money, can they quietly make it disappear? If the model has merely been told not to do something,  all we have a sign saying BRAKE. If the restriction disappears whenever it becomes inconvenient, we have much the same thing.

Slowing down gives us time to build and test real brakes. Humans decide where to put them and when to release them. Both matter, but neither is actually a brake.

Bob can have an Approve button, but unless not pushing it stops the machine, the AI will do whatever it’s going to do.

Thanks to Adam Engst, Lenny Foner, Alan Wexelblat, Jamie McCarthy, David Mankins, Robert Thau, and Greg Bolcer for reading an earlier draft, arguing with it, and making it better.

Mark Gibbs is a former journalist, former CTO, longtime technology columnist, and currently exhausted observer of what the industry is doing to itself. He has spent several decades reporting on technology, building and managing systems, running technical organizations, and advising companies on how to use technology without being eaten by it. Along the way, he has written extensively about computing, networks, software, security, and the recurring human tendency to deploy complicated systems first and ask difficult questions later. Gibbs also keys two other blogs: The F*ck You Muscle and The Bard’s Knife. He can be reached at [email protected].

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart