Evidence before acceptance

How to Detect Unintended Behavior Changes in AI-Generated Code

How to Detect Unintended Behavior Changes in AI-Generated Code

A one-line code change can alter three business workflows while every test stays green. Here's how to compare behavior against an accepted baseline and make each acceptance decision inspectable.

Lindsay Britt

8 min read

A one-line code change can alter three business workflows while every test stays green. When that happens, the engineering organization has approved more than it understood.

AI coding tools make this problem harder to manage because they can produce changes faster than teams can investigate their consequences. The constraint becomes the organization’s ability to establish what changed, what depends on it, and whether the difference is acceptable.

That requires evidence about behavior. A small diff and a passing build don’t provide enough of it.

What is an unintended behavior change? An unintended behavior change is a change in what software does that nobody authorized. It can appear outside the code a developer meant to change, in a workflow that shares the edited logic, which is why it can pass a code review and tests focused on the intended change.

Why a one-line AI change is hard to verify

Consider a Java billing service. A team asks an AI coding assistant to adjust an invoice calculation. During the work, the assistant changes a rounding rule in a shared helper:

// Before

public static BigDecimal normalizeAmount(BigDecimal amount) {

    return amount.setScale(2, RoundingMode.HALF_UP);

}

// After

public static BigDecimal normalizeAmount(BigDecimal amount) {

    return amount.setScale(2, RoundingMode.HALF_EVEN);

}

In this illustrative example, invoice calculation, month-end settlement, and fee calculation all use the helper. The existing tests pass because none exercises a value that distinguishes the two rounding rules. The helper is in an allowed file, so the change also passes the team’s file-scope check.

For an exact decimal input of 2.345, the old rule returns 2.35. The new rule returns 2.34. HALF_UP rounds ties away from zero; HALF_EVEN rounds them toward the even retained digit.

Neither rule is inherently wrong. The problem is that the request didn’t authorize changing the rounding behavior of settlement and fees.

When settlement passes the exact decimal 2.345 into the helper, the result changes from 2.35 to 2.34. Whether that difference is acceptable depends on the settlement requirement, not the passing invoice tests.

The edit is easy to see. Establishing its consequences requires tracing the affected workflows and checking their requirements. An unchanged caller or method signature provides no assurance that its behavior stayed the same.

Why passing tests can’t validate AI-generated code on their own

Tests provide evidence for the cases they exercise. A boundary test using 2.345 would expose this difference immediately. The requirement would establish which answer is correct.

What a passing suite cannot establish on its own is whether the team identified every relevant consequence of the change. Coverage can show that a helper executed without showing that the tests exercised the distinguishing input or asserted the required rounding rule. Coverage is a common gate for AI-generated code: Futurum Research found that 58.6% of software lifecycle engineering decision-makers mandate automated test coverage thresholds for AI-generated code reaching production.[1]

AI-generated tests don’t remove that limitation. When an assistant writes both the implementation and its tests from the same mistaken assumption, they can agree perfectly and still violate the requirement. Expected results need an independent basis in requirements, accepted behavior, or another appropriate reference.

The operational problem is making the investigation happen consistently across a growing volume of changes, without requiring the most experienced person on the team to reconstruct the system for every review. Left untraced, those effects accumulate into architectural drift.

Each common check establishes something specific, and leaves something open:

Check

What it establishes

What it leaves open

Diff review

What text changed

Which workflows depend on the changed code

Tests

Whether the exercised cases produce expected results

Cases no test exercises, such as 2.345 in this example

File-scope checks

Whether the edit stayed within allowed files

Effects on workflows outside the edited files

AI code review

A model’s assessment of risk or scope

A source-traced record of what changed and who accepted it

Behavioral comparison against a baseline

Changed rules and affected relationships, traced to source

Behavior outside the analyzed scope, which stays explicitly unverified

Compare behavior against an established baseline

Detecting unintended change requires two references: what the software currently does and what the proposed change is authorized to alter.

Within a defined analysis scope, source can be used to establish the rules, relationships, and behavioral semantics represented by the software. Authorization comes from requirements and accountable decisions. Existing software contains defects and obsolete rules, so preserving everything indiscriminately would be a poor governance policy.

A behavioral baseline makes the comparison explicit. For the rounding example, it records the existing operation: monetary values are normalized to two decimal places using HALF_UP, with identified relationships to the invoice, settlement, and fee workflows.

The proposed revision changes one part of that record. Scale remains two decimal places. The rounding rule becomes HALF_EVEN. The identified callers remain connected to the helper, which means their unchanged source does not exempt them from impact review.

This is the mechanism that matters: derive the behavior represented by each revision, compare the rules and relationships, and evaluate the differences against the authorized change.

CodeIntent maintains a governed baseline grounded in source-derived semantic evidence, so each proposed change can be evaluated against an established, accepted reference, with prior acceptance decisions available for review.

Caller search contributes to the investigation. By itself, it doesn’t establish which rule changed, whether each affected workflow permits the new behavior, or who accepted the difference. Those connections need to become part of the review evidence.

Make the acceptance decision inspectable

A useful record separates what analysis established from what a reviewer decided.

Illustrative review record, not captured CodeIntent output. The boundary input and test-gap assessment are supplied for this example.

Review item

Finding or decision

Authorized scope

Adjust the invoice calculation; preserve existing settlement and fee behavior

Behavioral difference

Shared rounding rule changed from HALF_UP to HALF_EVEN

Traced relationships

Invoice, settlement, and fee workflows reference the helper within the analyzed scope

Distinguishing input

Exact decimal 2.345 produces 2.35 before the change and 2.34 afterward

Test evidence

Existing tests in this example do not exercise an input that distinguishes the rules

Required action

Restore the shared rule, or explicitly authorize and validate the affected behavior changes

Reviewer’s merge decision

Hold while required findings remain unresolved

The record should identify the compared revisions and link findings to their source. Test results should reference the tests that produced them. Approvals should identify the decision-maker and apply to the revision actually reviewed.

This gives an engineering leader an inspectable account of the behavioral change and its authorization. It also makes review more focused. Instead of asking someone to “take another look,” the team can ask the settlement owner to resolve a specific change to a specific rule, supported by a concrete input and result.

Direct review toward the changes that need it

A governance process that sends every shared-helper edit into a manual queue will become a bottleneck. The evidence should determine where additional review is necessary.

If the comparison establishes that a formatting change leaves the represented rules and relationships unchanged, there is no behavioral exception to authorize. If the required checks pass and the analysis has no blocking gaps, policy can allow it to continue through the normal approval path.

The rounding change produces a different outcome. It alters a shared rule outside the requested scope. That calls for correction or targeted authorization by the owners accountable for the affected behavior.

  1. Declare the intended change and protected behavior. Establish what may change and who can authorize an exception.

  2. Compare the accepted and proposed revisions. Identify changed behavior within the governed scope.

  3. Trace the affected relationships. Include consequences in code that wasn’t edited.

  4. Validate the differences. Use requirements, distinguishing inputs, tests, and appropriate technical review.

  5. Apply the acceptance policy. Let sufficiently supported changes proceed. Hold changes with unresolved required findings unless an authorized exception applies.

  6. Retain the evidence and advance the baseline. Verify the accepted revision and preserve the decisions supporting it.

The benefit is a more precise use of engineering judgment. Reviewers spend their attention on identified differences and unresolved questions instead of repeatedly assembling the context needed to discover them.

What CodeIntent adds

CodeIntent gives reviewers changed rules and affected relationships with evidence traced to source. That reduces the context they have to reconstruct for each pull request and gives the team a record it can continue to use as the software evolves. See how CodeIntent works.

Applied to this example, a behavioral record would capture the normalization rule, its scale and rounding mode, and its relationships to invoice, settlement, and fee calculation. Comparing revisions would expose the changed rounding rule while retaining the relationships that make its impact visible.

The record persists as the code changes. The next review starts from an accepted reference, and the evidence behind previous decisions remains available. Engineering can evaluate successive changes against an explicit account of the behavior it has accepted, rather than treating every pull request as an isolated investigation.

With the same governed revisions, analysis version, and rules, the deterministic analysis result is repeatable. Model-assisted assessments can help explain or investigate findings, but they remain distinguishable from source-derived evidence.

The analysis boundary also stays explicit. If the rounding mode depends on deployment configuration unavailable to the analysis, the effective rule remains unverified. Additional evidence or an authorized exception may support acceptance; neither silently turns an unknown into a verified result.

For the next change to shared business logic, require the affected workflows, a demonstrated behavioral difference where one exists, and the authorization supporting acceptance. If required evidence is missing and no authorized exception applies, the governance result remains unresolved, even when every test is green.

Bring a pull request that touches shared business logic to a CodeIntent walkthrough. Examine the baseline comparison, the affected relationships, and what remains unresolved before acceptance.

Share this post

Stay in the loop

Get new writing on deterministic modernization, evidence, and governed AI.

1. Futurum Research, 2H 2026 Software Lifecycle Engineering Global Enterprise Decision Maker Survey Report, July 2026 (subscriber access): app.futurumgroup.com/share/OPSH6u8XSAxYTgB7

Supporting public source: Futurum, “Software Lifecycle Engineering Market to Reach $226 Billion by 2030,” announcement of the 2H 2026 research, July 13, 2026: futurumgroup.com/press-release/software-lifecycle-engineering-market-to-reach-226-billion-by-2030

HOLONIC

The deterministic evidence layer underneath legacy modernization and the agentic enterprise.

© 2026 Holonic Technologies, Inc. · Atlanta, GA · Tucson, AZ

CodeIntent® is a registered trademark of Holonic Technologies.