The share of AI-generated code inside many engineering organizations has been climbing sharply over the past year. The distribution is uneven, but I now often hear engineers say that nearly all of the code they write is AI-assisted in some form, whether through prompt-and-paste workflows in chat windows, copilots embedded in editors, or autonomous agents operating inside CI pipelines.
The upside is obvious: higher velocity. As an engineering manager, how do you mitigate the downside? Will an AI-generated change break production? Worse: will it pass every unit test, integration test, and CI check you have and still introduce a bug that surfaces weeks later? And what happens to the architecture itself as the rate of code change climbs and the half-life of design assumptions shrinks?
I increasingly see AI-generated code fail in ways that differ from human-authored code: confident calls to APIs that do not exist, change sets much larger than a human would typically author for the same task, and tests that pass because they encode the same incorrect assumptions as the implementation itself. I have been working through a layered model of guardrails to mitigate some of these issues:
- Scope and authority: what the AI is allowed to touch.
- Structural: what the code must look like before it reaches review.
- Behavioral: what the code must prove it actually does.
- Human verification: what a person must check before merge.
(My deeper concern though is architectural drift: AI systems are good at producing locally reasonable changes, but they do not naturally preserve long-term consistency across an entire system. Over time, this creates what I think of as “architecture slop”: a codebase where each individual change makes sense in isolation while the overall structure gradually loses coherence. More on that in a follow-up post!)
Layer 1: Scope and authority
Define as engineering policy what AI can do without escalation, what requires review, and what should be largely human-authored.
- Green zones: This is where AI can generate code freely. POCs/prototypes, Internal tooling, test scaffolding, documentation, greenfield features behind a feature flag, build scripts, log parsers, one-off data migrations.
- Yellow zones: AI can suggest, but a human authors the final commit. User-facing flows, database schemas, public APIs, anything that crosses a service boundary.
- Red zones: AI can draft or assist with review, but never authors the final commit, and a senior reviewer signs off in addition. Terraform scripts, Authentication and session handling, payment processing, anything that touches PII or regulated data, security-sensitive code paths, cryptographic primitives.
A concrete example: SQL migrations. An AI agent can be good at drafting them, but they belong in a yellow zone, which means the agent can draft the migration but a human has to read it, run it locally, and sign off before it runs against staging. The agent does most of the rote work, and the human catches the one in twenty cases where the migration would have locked a high-traffic table during business hours or dropped an index that some downstream report was depending on. Putting migrations fully into the green zone would make those rare failures much more dangerous.
Across all three zones, the engineer who merges the change owns it, whether or not an AI wrote it.
The other dimension worth defining is the scope of any single change. Adding a new endpoint that lives in its own module is a bounded change with a contained blast radius. Refactoring the request pipeline across twenty files is not. The first is fine for AI to handle; the second is exactly where AI tends to make confident, sweeping mistakes that take hours to unpick.
As an example, engineering policies could limit AI-generated code to one new module per PR, and require authors to write a summary in their own words for any PR that crosses module boundaries or changes more than 300 lines.
Layer 2: Structural guardrails
These are automated checks that run before a human ever reviews the code, and they catch the bad habits AI-generated code tends to exhibit.
Dependency policy checks
AI-based coding tools are enthusiastic about pulling in libraries. They often import something that solves a problem you already have a solution for, or pick an abandoned package that has not seen a commit in years. A written justification for any new dependency, even one or two lines in the PR description, makes someone stop and ask whether the dependency is actually needed. Every new dependency should be a deliberate choice and not just an autocomplete.
Hallucinated API checks
These checks matter most in dynamically typed languages, where calls to non-existent methods only surface when the code actually runs.
An undefined-symbol linter (pyflakes, ESLint's no-undef) catches the easy cases: importing a module that does not exist, or calling a function that is not defined in the codebase or its dependencies.
A type checker in strict mode (mypy, pyright, the TypeScript compiler) catches the harder ones, including invented methods on real objects (user.getDisplayName() when the real method is user.displayName) and return values used as if they could not be None.
Take the last case as a worked example. The AI generates:
def get_user_email(user_id: str) -> str:
user = db.users.find_one({"id": user_id})
return user.email
mypy or pyright in strict mode rejects this immediately: find_one returns Optional[User], not User, so user.email will throw on the unhappy path. A human reviewer would probably miss it on first read because the code looks reasonable.
A type checker catches it on every run, and most editors flag it in real time as the engineer types.
Layer 3: Behavioral guardrails
A massive, almost entirely agent-coded refactor passed all unit and pre-merge tests but broke a critical feature.
It was only caught due to my own excessive paranoia making me run end-to-end tests before the prod deploy.
The traditional testing pyramid puts unit tests at the wide base, integration tests in the middle, and end-to-end tests at the narrow top. It emphasized unit tests because they are fast, relatively stable, and cheap to run repeatedly. End-to-end tests, by contrast, have historically been slower, more brittle, and more expensive to maintain.
AI systems are extremely good at generating unit tests that appear thorough while proving very little, because the tests often encode the same assumptions and blind spots as the implementation they are validating.
Large agent-coded refactors can pass full unit-test suites and pre-merge checks and still break critical end-to-end behavior. Unit tests will still have a role, but the key guardrail needs to shift to end-to-end and behavioral tests.
I see two major shifts happening in how we test.
The first is around test scope. There is a class of bugs in AI-generated systems that will routinely escape unit-test guardrails. For example, an endpoint may call the correct helpers, but in the wrong order. Every unit test against those helpers will still pass because each helper behaves correctly in isolation. An end-to-end test is often the only thing that catches that the system as a whole no longer behaves the way users expect. For every user-facing workflow that matters, there should be at least one end-to-end test that exercises the full stack.
Historically, end-to-end tests were avoided because they were slower, flakier, and more expensive to maintain. That tradeoff is beginning to shift now that AI systems themselves can generate and maintain large portions of the test suite.
The second shift applies across every level of testing. Tests should validate what the system does, not how the implementation happens to be organized internally. A behavioral test passes when the system produces the correct output for a given input. A structure-coupled test passes only when the internals follow a particular shape: a specific field was mutated, a helper was called, or a sequence of operations occurred in a specific order.
Consider a function that calculates shipping cost. Both tests below are unit tests; the difference is in what they assert.
def calculate_shipping(order: Order) -> Decimal:
weight = sum(item.weight for item in order.items)
zone = lookup_zone(order.destination)
base = WEIGHT_TABLE[zone][weight_bucket(weight)]
return base + tax_for(order.destination, base)
A structure-coupled test might look like this:
def test_calculate_shipping_calls_dependencies():
mock_zone = patch("shipping.lookup_zone", return_value="ZONE_A")
mock_tax = patch("shipping.tax_for", return_value=Decimal("1.50"))
with mock_zone, mock_tax:
result = calculate_shipping(sample_order)
assert mock_zone.called
assert mock_tax.called_with("US-CA", Decimal("10.00"))
This test passes only when the internals are arranged a specific way. The moment you regenerate the function and the AI inlines lookup_zone or combines the tax calculation into the base lookup, the test breaks even though the code is still correct. The test is punishing you for refactoring.
A behavioral test for the same function:
def test_calculate_shipping_returns_correct_total():
order = Order(
items=[Item(weight=Decimal("2.5"))],
destination="US-CA",
)
assert calculate_shipping(order) == Decimal("12.47")
This passes whenever the function produces the right answer, regardless of how it gets there. The implementation can be regenerated repeatedly, and as long as the behavior remains correct, the test continues to pass.
Structure-coupling matters more under AI-generated code than it did under human-authored code because of a higher rate of implementation churn. A human refactoring a module often has a mental model of how that code should be organized and code tends to preserve that shape over time. An AI regenerating the module will happily restructure things in ways that are functionally equivalent but structurally very different. Tests coupled to structure become noisy with every code regeneration, and will end up being ignored.
What about unit tests?
This does not mean unit testing should be abandoned. End-to-end tests do not tell you where the breaking change is, and the combinatorial logic paths in your code cannot all be exercised by end-to-end tests alone.
You need to understand which unit tests should be kept. A useful heuristic: if you regenerated this code tomorrow, would the test still make sense? If yes, keep it. If it would break just because the internal implementation shifted, then rewrite it against the public interface or delete it.
Practical behavioral test guardrails
- Require integration tests for new public interfaces. Every new endpoint, queue handler, or exported function gets at least one test that calls it the way a real caller would. One happy path and one realistic failure case are usually enough to catch the things mocks let through.
- Use property-based tests for combinatorial logic. Tools like Hypothesis in Python or fast-check in JavaScript generate inputs the AI did not think of, which is precisely the failure mode AI-generated code is prone to: empty inputs, unicode strings, negative numbers, off-by-one boundaries at the limits of an integer range.
- Require additional human review whenever AI-generated changes modify both implementation and behavioral expectations test simultaneously. One common failure mode is an AI system “fixing” a failing test by weakening the assertions rather than addressing the underlying issue.
- Measure coverage on the diff, not the whole codebase. Whole-codebase thresholds are easy to game: engineers add tests to easy code to keep the number up, and the hard code stays untested. Coverage on changed lines is honest. Setting the threshold at 80% on changed code routinely takes coverage of new code from the mid-40s to the high-70s within a quarter, with the only visible behavior change being that engineers stop shipping untested branches.
Layer 4: Human verification
The last layer is what a human must do before code merges. Generic code review is not enough, because reviewers skim AI-generated code the same way they skim human code, and AI code is plausible-looking in ways that defeat skimming. The classic failure is the reviewer who approves with “LGTM” on a 400-line AI-generated diff.
How deep human review needs to go depends on the zone (Layer 1). In green zones, an AI code reviewer in CI is enough on its own; the volume of internal tooling and test scaffolding being generated does not justify pulling a human onto every PR, and the blast radius if something is wrong is small. In yellow zones, the AI reviewer clears the obvious issues but a human reviewer is required on top, with the three practices below. In red zones, the AI reviewer is advisory at best, never gating: a senior reviewer with explicit authority over the code path signs off in addition to the three practices, and for the most sensitive code (payments, auth, crypto, regulated data) a second human reviewer is worth the friction.
Three practices may help reviewers avoid skim mode on yellow- and red-zone AI-generated PRs (especially red zone!):
package.json.Planning for a shrinking architectural half-life
Guardrails keep individual changes safe. They do not solve the longer-term problem, which is that all software architecture ages. Most codebases start showing their age around the two-year mark, when accumulated decisions and shifting business requirements begin to outgrow the architecture they were originally built around.
Several forces are compressing the effective lifespan of software architecture at the same time.
The first is sheer code volume. AI dramatically lowers the cost of producing code, which means systems accumulate complexity much faster than before.
The second is implementation churn. Decisions that teams previously deferred for months can now be regenerated in an afternoon. Each individual change may be locally reasonable, but over time the cumulative effect pushes the architecture in inconsistent directions.
The third is faster product cycles. AI lowers the cost of shipping software, which means competitors iterate faster, customer expectations shift faster, and requirements evolve more rapidly.
The fourth is what I think of as “local coherence with global drift”. Each individual change looks reasonable because the AI optimized it against the immediate context it was given. But no single mind consistently holds the broader architectural model in view, so consistency gradually erodes: patterns emerge that were never deliberately chosen, conventions drift, and abstractions appear that nobody consciously designed. Over time, the codebase begins to resemble a city where every block was designed independently, without a shared plan. That is what I would describe as “architecture slop”.
Taken together, these forces likely compress the traditional two- to three-year architectural refactor cycle significantly, especially in fast-moving products. Refactoring at a two-year mark may itself become too slow and too expensive to keep pace.
More on this in a follow-up post!
Closing thoughts and caveats
Most of us are still figuring out what effective guardrails for AI-generated code actually look like. The ideas above are not a finished framework, but rather an exercise in thinking in public.
What is clear though is that we are dealing with far larger volumes of code, much higher rates of change, and failure modes that differ from traditional human-authored software. The next couple of years of engineering management are going to be about figuring out which practices need to change, and how.