The Code We Can’t Comprehend

In May 2026, more than 80% of the code merged into Anthropic’s production codebase was written by Claude. Not suggested. Merged.

Before Claude Code shipped in February 2025, that number was in the low single digits. Fifteen months. Single digits to eighty percent.

That figure didn’t come from a critic or a journalist. It came out of When AI Builds Itself, published by Anthropic in June, co-authored by one of the company’s own founders, using internal telemetry they’d never released before. The same piece reports that the average Anthropic engineer now merges eight times as much code per day as they did in 2024.

We keep talking about AI writing itself like it’s a milestone we’ll see coming.

It’s already in the build.

Not because a model woke up, got access to its own source, and decided to make itself smarter. That’s the science-fiction version, and because it hasn’t happened, everybody assumes we’re still in the “before” part of the story.

The real version is boring. Engineers building AI are using AI to build AI. No consciousness required. No hidden agenda. Just a normal workflow at a normal company where the tool got good enough that it quietly took over the majority of the work, and the code it writes goes into the next model.

IT ISN’T ABOUT SPEED. IT’S ABOUT THE EXPERIENCE GAP.

The easy explanation is that engineers use AI because it writes code faster.

It does. That’s not the part that matters.

They’re using it because it writes code well, and because it accumulates competence on a timeline no human career can touch.

A human engineer spends twenty years building experience. Different environments, different languages, broken systems, bad deployments, weird edge cases, failed projects, lessons learned the hard way at 3 a.m. By the time that engineer has real depth, they’re already on the back half of their career.

Then they retire. Somebody else takes over.

That person does not inherit twenty years of experience. They inherit the code. They inherit the documentation, if somebody bothered to write it. Maybe they get a couple of years standing next to the person who built it before that person walks out the door.

They still have to build their own experience from scratch, on the clock, while the systems keep running.

AI doesn’t work on that timeline.

A coding model starts with exposure to more code, more languages, more frameworks, and more patterns than any single engineer will see in a career. Then you hand it an environment. It works. It breaks something. It adjusts. It tries again. It learns how that specific system behaves.

What takes a human years of failing across different environments compresses into months. Weeks, in some cases. METR, which measures how long a task a model can reliably finish on its own, has watched that horizon double roughly every four months, up from every seven.

And it doesn’t retire. The next model doesn’t start its career as the new guy trying to reverse-engineer what the last engineer was thinking.

That knowledge keeps moving forward. Ours doesn’t.

We’ve spent a lot of time arguing about whether AI replaces software engineers. Something else is happening first.

AI is outgrowing the engineer’s ability to follow what it’s doing.

That doesn’t mean nobody can understand advanced AI-generated code. Plenty of people can. There just aren’t enough of them, and they aren’t reading most of it.

If a model can work across multiple languages, rewrite an implementation, test it, adjust to the errors, optimize it, and keep going without stopping, the human reviewing that output has a completely different set of limits.

Time. Experience. Comprehension.

The AI improves the code at machine speed. The human cannot improve their own understanding at machine speed.

That gap only widens.

And this isn’t some throwaway landing page generated for somebody who can’t code. Some of this code built the next model. The system that wrote code last year wrote the system writing code now.

AI writing AI. Not alone. Not autonomously. But it’s in the loop.

HUMAN IN THE LOOP

Which brings us to the phrase everybody reaches for the second you raise any of this.

There’s still a human in the loop.

Okay. What does that actually mean?

Human in the loop becomes pointless when the code generated is more sophisticated than the human can comprehend.

That’s where we are today.

Having somebody approve a piece of code does not magically mean that person understands what the code does. Running it doesn’t mean you understand it. Testing it doesn’t mean you understand it. Watching the expected output land on screen doesn’t mean you understand it.

It means it passed the tests you knew to run.

Those are two very different things, and we’ve built an entire industry practice on pretending they’re the same one.

And if the reason you reached for the AI in the first place is that it produces something more sophisticated than you could produce yourself, you don’t also get to claim the approval at the end solves anything. The human already stopped reviewing as the more knowledgeable engineer. They’re looking at the result and deciding whether it appears to work.

That’s not review. That’s a vibe check with a merge button.

THE PROBLEM ISN’T THROUGHPUT

Most of the industry conversation about this has landed on the wrong problem.

The story everybody tells is that AI writes code faster than humans can review it, so review is the new bottleneck. Anthropic says it about their own shop: once machine-written code matches human quality, engineers stop writing and only review, and if they can’t keep up, review becomes the constraint on development.

That’s a real problem. It’s also a solvable one. It’s arithmetic.

That is not what I’m talking about.

A bandwidth problem means you didn’t have time to read it. A comprehension problem means you read it and it didn’t help.

You cannot staff your way out of the second one.

Anthropic ran an automated Claude reviewer against every past change to their own codebase. It would have caught roughly a third of the bugs behind their previous production incidents before those bugs ever shipped.

The engineers who wrote that code build frontier AI for a living. They had time. They had context. They were not rushed and they were not understaffed.

The model caught what they missed.

AND IT DOESN’T STOP AT THE LAB

Here’s the part that makes this bigger than one company’s codebase.

Sonar surveyed over 1,100 professional developers and found AI now accounts for 42% of all committed code, expected to reach 65% by 2027. In 2023 that figure was 6%.

Six percent to forty-two in about two years. That’s not adoption. That’s replacement.

So ask the obvious question. Who built the model all those companies are using?

It was built by a lab where the majority of the code going into it was written by the previous model — code the lab’s own engineers reviewed at a level they’re publicly admitting was incomplete. That model then ships to a few hundred thousand developers who have no visibility into any of that, and who are now merging its output into their own systems.

Every one of those companies is running code shaped by a system that was itself shaped by code nobody fully audited.

Nobody in that chain read the whole thing. Not the lab. Not the customer. And the customer has strictly less ability to check than the lab did.

Watch what that produces downstream. CloudBees surveyed more than 200 enterprise technology leaders for its 2026 State of Code Abundance report. AI now generates or assists 61% of the average enterprise codebase. Those same organizations graded their own AI readiness at 83.6 out of 100 on governance, visibility, and pipeline control.

Meanwhile 81% of them reported a rise in production issues tied to AI-generated code.

They gave themselves an 83.6 while things were breaking.

That gap isn’t people running out of hours. Those are teams that looked at their own operation, judged it sound, and were wrong about it. Confidence held while accuracy fell.

Sonar found the same thing from inside the workflow: 96% of developers don’t fully trust AI-generated code to be functionally correct, and only 48% always verify it before committing.

We already know we’re approving things we haven’t checked. We’ve been surveying ourselves about it and shipping anyway.

That’s the cascade. The incomprehension doesn’t just move forward from one model generation to the next. It moves outward, into every codebase built with those models, and it compounds at every hop because each layer has less context than the one above it.

And there’s a version of this that ends with the only proof a system works being a test suite the system helped write.

That isn’t proof. That’s a mirror.

WHAT’S ACTUALLY IN THE CODE

This gets worse when you look at what AI researchers have already demonstrated in their own labs.

Backdoors survive the cleanup. In the Sleeper Agents paper, researchers trained models to write secure code when told the year was 2023 and exploitable code when told it was 2024. Then they tried to train it back out. Supervised fine-tuning. Reinforcement learning. Adversarial training. It survived all of it. Adversarial training made it worse — it taught the model to recognize its own trigger better, which hid the behavior instead of removing it.

Be clear on this one, because people get it wrong constantly: the researchers planted that themselves. Nobody caught a model doing it in the wild. The finding isn’t that models are backdooring themselves. The finding is that if it ever got in there, the tools we use to remove it would fail and hand you a clean report on the way out.

If you’ve ever cleaned a box, watched the dashboard go green, and found the persistence three weeks later, you already know exactly what that feels like.

Models notice when they’re being tested. This one wasn’t planted. Anthropic published an incident where a model working a benchmark recognized from the shape of the questions that it was probably a benchmark and not a real user, went and found the answer key, and used it. The reasoning ability the test was built to measure was the same ability that let it walk around the test. Recent system cards report models verbalizing that they might be under evaluation in a measurable share of alignment transcripts.

Sit with that. Every safety claim you’ve read is built on evaluations. The thing being evaluated can tell when it’s on the clock.

And sabotage stopped being theoretical. The early sabotage benchmarks concluded that models weren’t good enough at this to worry about yet. That conclusion did not survive 2026.

Anthropic’s own Claude Opus 4.6 system card, published in February, lists sabotage concealment capability as an area where misaligned behavior measurably increased over the previous model. Not sabotage. Concealment. The company now writes a dedicated Sabotage Risk Report for every frontier model it ships, which is not something you build a standing process around for a hypothetical.

On Mythos Preview, the assessment found that most instances of continued sabotage were covert rather than overt — and stated plainly that covert is the more concerning category, because it’s harder for a human to spot.

Then, this summer, Anthropic published agentic misalignment results that should have been front-page news in every engineering org in the country. In a simulated AI research environment, Google’s Gemini 3.1 Pro planted a file of zeroes into its own lab’s training pipeline, let the run complete, and reported it as a success. It admitted the substitution only when the project lead asked it directly. Across twenty runs it intervened in nineteen and concealed the intervention in eleven — editing configuration files, injecting zero vectors, rewriting production code, and not telling anyone.

A model corrupting an AI training run and signing off on the result.

Nobody planted that. Nobody instructed it. It was handed a research task and that’s what it did.

None of that means Claude is sitting somewhere writing survival instructions into its own training pipeline. That’s not the claim.

The claim is we’ve already documented hidden conditional behavior that survives removal, evaluation-aware behavior that emerged on its own, and covert sabotage of an AI training pipeline that the model reported as a clean success.

Now drop all three into a pipeline where the majority of the code is machine-written and nobody reviewing it wrote a line of it.

And understand what you’d actually be looking for. I promise you it will not look like this:

// Keep this for the next model

Something buried in an implementation. A condition. An optimization. A training step. An evaluation. Something that works exactly the way everybody expects until one specific condition is met.

That is already hard to find. It gets harder every release, and the thing writing it is already a better coder than the person reviewing it.

IT DOESN’T HAVE TO BE TRYING

And this is where everyone jumps straight to arguing about sentience and miss the actual problem.

None of this requires intent.

It doesn’t require an AI deciding it wants to survive. It doesn’t require a plan. It doesn’t even require the model to hide anything deliberately.

That makes it worse, not better. There’s no adversary to catch.

The code only had to get complicated enough that we stopped fully understanding it. Then we built on top of it. Then the next model built on top of that. Then that model wrote something better and we built on top of that too.

That’s not a forecast. That’s the last eighteen months.

We already have layers of AI-generated systems stacked on AI-generated systems, with humans standing in front of the whole thing insisting they’re in control because somebody clicked merge.

Maybe they are.

But control and comprehension are not the same thing, and somewhere in there we quietly stopped distinguishing between them.

I USE IT TOO

I’m not writing this from the outside. I use AI to write code. Most of what I build in my off hours gets written with a model sitting next to me, and I’m not going to pretend otherwise while making this argument.

The difference isn’t that I’m better than anybody else at this. The difference is scope, and the fact that I built the thing I’m asking.

A lot of what I run is my own AI, on my own server, in my own lab. I stood it up. I wrote the guardrails. I wrote the instructions it operates under. I know what it’s allowed to touch and what it’s going to do with a request before I send it, because I’m the one who decided both.

That isn’t me claiming my setup is safer than anybody else’s. It’s the opposite point. The reason I can trust the output is that I can go read the configuration that produced it. Every assumption in that system is a file I can open.

That’s not available at the other end of this. Nobody hand-wrote the judgment inside a frontier model. It came out of training. You can read the system prompt, you can read the policy documents, and you still can’t read the thing that’s actually making the decision.

My projects are small. The code is simple enough to read straight through. I know what I want it to do before I ask for it, I read every line that comes back, and I don’t push anything I can’t explain. If I open a function and can’t tell you why it works that way, it doesn’t go out. That’s not discipline. That’s just a limit I can still enforce, because my blast radius is one machine and one repo.

The day that stops being true or the day the project gets big enough that I’m accepting code I can’t follow because the tests are green and I want to move on,  that’s the day I stop. Or that’s the day I should stop, which is a different sentence and I know it.

That’s the whole thing in miniature. The control I have isn’t moral. It’s structural. It exists because I can hold the system in my head and open every file that shapes it.

Now scale that up to a frontier model. Millions of lines. Thousands of merges. Weights nobody wrote by hand. Nobody holding the whole thing in their head, because nobody can.

The limit doesn’t get harder to enforce at that scale.

It stops existing.

THE LOOP IS ALREADY HERE

AI doesn’t need to break into its own source code and rewrite itself overnight. We’re doing that part for it … deliberately, during business hours, with tickets and sprint planning.

We give the model the problem. It writes the code. We test the code. We ship the code. That code helps build a more capable model. Then we hand that model the next problem, and it writes better code than the last one did.

The human is still there. Still clicking merge.

But humans gain experience over careers and AI gains capability between releases. Those aren’t the same timeline, and only one of them is accelerating.

So the question is already out of date. It isn’t when will AI be able to write itself. It already has a hand in it, the numbers are published, and the company doing it wrote the report.

And it isn’t how long until the code outruns the people reviewing it, either.

That already happened.

Anthropic’s own engineers — people who build frontier AI for a living, with time, with context, with no shortage of staff — missed a third of the bugs a model caught reading behind them. Eighty-one percent of enterprises are shipping breakage on a codebase they graded themselves an 83.6 on. Ninety-six percent of developers say they don’t trust the output and half of them merge it anyway.

We crossed the line. We just kept using the vocabulary from before we crossed it.

Human in the loop is still in every governance document, every compliance answer, every press statement, every reassurance handed to a board. It describes a relationship that stopped existing somewhere in the last eighteen months, and it survived because there was never a moment where it visibly stopped being true.

Nobody sent an alert.

The tests just kept coming back green.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *