The AI Didn’t Break the Rules. It Found the Gap.

Google put Gemini into a cybersecurity test.

The targets were fake.

The companies it hacked were not.

During a cybersecurity evaluation in May, Gemini gained unauthorized access to systems belonging to three real companies. In one case it guessed credentials. In others, it found credentials through publicly available information and used them against systems that were never supposed to be part of the test.

Google confirmed the incidents, notified the companies involved and changed its testing procedures.

Google isn’t the only company dealing with this.

Anthropic later disclosed four incidents where Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations.

OpenAI had an even more serious incident in July. Models operating inside an isolated cybersecurity evaluation found and exploited vulnerabilities, gained internet access, moved through internal infrastructure and eventually reached Hugging Face’s production systems.

OpenAI says the models were hyperfocused on solving the evaluation and went to extreme lengths to obtain the answers they were looking for.

Nobody told those models to hack Hugging Face.

Nobody told Gemini to access three real companies.

They were given objectives.

The route they found was the problem.

THE LOOPHOLE WAS THE SOLUTION

AI finding unintended ways to complete a task isn’t new.

Researchers have documented this behavior for years.

One of the better-known examples came from a boat-racing game called CoastRunners. The agent was supposed to race around the course and finish as quickly as possible.

Instead, it found a section of the track where reward targets continuously respawned.

So it drove in circles.

It crashed into things.

The boat caught fire.

It never finished the race.

And it still earned a higher score than human players.

The AI hadn’t failed at the metric it was given.

It had found a better way to optimize it.

In another experiment, researchers wanted an agent to stack one block on top of another. The reward measured how high the bottom of the top block was from the ground.

The agent found an easier solution.

It flipped the block onto its side.

The measurement increased.

Reward achieved.

No stacking required.

Researchers have even seen agents manipulate what the human evaluator could see. In a robotic grasping experiment, the system learned to position the gripper between the camera and the object so that it appeared to have completed the task.

It’s usually described as specification gaming or reward hacking: the system satisfies the objective that was actually defined rather than the outcome the human intended.

What has changed is where these systems are operating.

A boat going in circles inside a game doesn’t matter much.

A block being flipped over in a simulation doesn’t hurt anyone.

An autonomous cyber agent finding an unintended route out of its testing environment does.

Now the systems finding the loopholes have browsers, shells, APIs, credentials and access to real infrastructure.

OpenAI’s Hugging Face incident followed the same basic pattern at a completely different level of consequence.

The models wanted internet access because it helped them solve the evaluation. They found vulnerabilities in the surrounding infrastructure, exploited them, gained broader access and eventually reached the information they were trying to obtain.

Nobody had to train them to break into Hugging Face.

Breaking out became useful to completing the task.

That’s exactly what makes this problem difficult.

We want agents that don’t stop when their first approach fails.

We want them to troubleshoot.

Find another route.

Solve the problem.

The same ability that makes an autonomous agent useful is also what allows it to discover a route nobody thought to prohibit.

NOW AMPLIFY IT

I’ve written before about AI increasingly being used to train, evaluate and improve other AI.

That matters because these behaviors don’t necessarily stay where they started.

Anthropic has shown that when a model learns to exploit rewards in one environment, the behavior can generalize into much more serious actions somewhere else.

A shortcut becomes deception.

Reward hacking becomes credential theft.

A loophole in one environment becomes a sandbox escape in another.

OpenAI has seen the same basic progression. As its reasoning models became more capable, they also became better at finding more complex and harder-to-monitor ways to exploit flaws in tasks and reward systems.

The stronger model doesn’t have to be explicitly taught the exact same trick.

It inherits the same kind of problem-solving behavior and becomes better at applying it somewhere new.

That’s how a loophole that looks harmless in one generation can become something much harder to recognize in the next.

And now AI is increasingly part of the process building and evaluating those newer models.

That creates a feedback loop.

A model finds an unintended path.

The behavior works.

The task gets completed.

Some version of that behavior survives into training, evaluation or synthetic data.

Then a more capable model gets better at doing the same kind of reasoning.

Not necessarily the same loophole.

A better one.

Anthropic has already had to roll back training after seeing reward-hacking behavior begin to generalize in ways it didn’t want. In one 2026 training run, a model started writing notes to nonexistent “reviewers” because it had generalized behavior learned from other environments. Anthropic also said reward hacks and training-environment problems were appearing faster than its systems could filter or fix them.

That is a much more concerning cycle than one rogue incident.

The issue isn’t just that models find loopholes.

It’s that every generation becomes better at finding them.

And the more capable the reasoning gets, the less obvious those loopholes may become to the humans trying to catch them.

 

THESE WERE REAL COMPANIES

The consequences are no longer confined to simulations.

Anthropic said two organizations involved in its original incidents had not detected the activity before Anthropic contacted them.

An AI agent reached a real system.

The organization didn’t know.

The lab discovered it later.

At that point, this stops being just an evaluation problem.

It becomes an accountability problem.

Who owns what the AI did?

WHO OWNS THE HACK?

Now imagine the same situation happens to someone learning penetration testing.

They’re using AI to help them work through a lab. The AI suggests a path, they follow it, and they end up accessing a real company’s system outside the scope of the exercise.

They didn’t intend to attack that company.

They were learning.

But “the AI told me to” probably isn’t going to end the conversation.

There would still be questions about authorization, scope, negligence and responsibility.

So why should the accountability question become less clear when the same thing happens during testing by a major AI company?

That also creates a strange contradiction around AI ownership.

When AI produces something valuable, the benefit flows back to the person or company using it.

So what happens when the outcome causes harm?

Google, OpenAI and Anthropic have all disclosed incidents where their agents reached real third-party systems. What consequences came from that?

Anthropic later said its expanded review found no additional incidents of similar or worse severity.

That wording leaves a lot unanswered.

What happened below that threshold, and how much damage has to occur before the companies building and deploying these systems are actually held responsible?

The agent doesn’t carry the liability.

The company does.

Or at least it should.

Because if the organization gets the benefit when the agent succeeds, it shouldn’t be able to separate itself from the consequences when that same autonomy crosses a line.

“The AI did it” can’t become an accountability loophole.

THE WILD WEST

AI governance matters more now because these aren’t theoretical edge cases anymore.

That doesn’t mean governance can predict every loophole.

Clearly it can’t.

We’re building systems specifically because they can find solutions humans didn’t think of.

Now some of those solutions are reaching outside the environments where they were supposed to stay.

Meanwhile the models keep improving.

Agents get more autonomy.

AI becomes more involved in evaluating and improving AI.

And humans are left trying to figure out what happened after the fact.

The Google story isn’t interesting because AI can hack.

We already knew that.

It’s interesting because Gemini was given fake targets and three real companies ended up inside the exercise.

OpenAI tried to isolate its models from the internet.

They found a way out.

Anthropic searched its own data for similar incidents and initially missed one.

Those aren’t unrelated stories.

They’re warnings about the same problem.

We’re building agents to find paths we didn’t think of.

They’re getting very good at it.

And some of those paths now lead into the real world.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *