Evidence before acceptance
A one-line code change can alter three business workflows while every test stays green. Here's how to compare behavior against an accepted baseline and make each acceptance decision inspectable.
Lindsay Britt
8 min read
A one-line code change can alter three business workflows while every test stays green. When that happens, the engineering organization has approved more than it understood.
AI coding tools make this problem harder to manage because they can produce changes faster than teams can investigate their consequences. The constraint becomes the organization’s ability to establish what changed, what depends on it, and whether the difference is acceptable.
That requires evidence about behavior. A small diff and a passing build don’t provide enough of it.
What is an unintended behavior change? An unintended behavior change is a change in what software does that nobody authorized. It can appear outside the code a developer meant to change, in a workflow that shares the edited logic, which is why it can pass a code review and tests focused on the intended change.
Why a one-line AI change is hard to verify
Consider a Java billing service. A team asks an AI coding assistant to adjust an invoice calculation. During the work, the assistant changes a rounding rule in a shared helper:
// Before
public static BigDecimal normalizeAmount(BigDecimal amount) {
return amount.setScale(2, RoundingMode.HALF_UP);
}
// After
public static BigDecimal normalizeAmount(BigDecimal amount) {
return amount.setScale(2, RoundingMode.HALF_EVEN);
}
In this illustrative example, invoice calculation, month-end settlement, and fee calculation all use the helper. The existing tests pass because none exercises a value that distinguishes the two rounding rules. The helper is in an allowed file, so the change also passes the team’s file-scope check.
For an exact decimal input of 2.345, the old rule returns 2.35. The new rule returns 2.34. HALF_UP rounds ties away from zero; HALF_EVEN rounds them toward the even retained digit.
Neither rule is inherently wrong. The problem is that the request didn’t authorize changing the rounding behavior of settlement and fees.
When settlement passes the exact decimal 2.345 into the helper, the result changes from 2.35 to 2.34. Whether that difference is acceptable depends on the settlement requirement, not the passing invoice tests.
The edit is easy to see. Establishing its consequences requires tracing the affected workflows and checking their requirements. An unchanged caller or method signature provides no assurance that its behavior stayed the same.
Why passing tests can’t validate AI-generated code on their own
Tests provide evidence for the cases they exercise. A boundary test using 2.345 would expose this difference immediately. The requirement would establish which answer is correct.
What a passing suite cannot establish on its own is whether the team identified every relevant consequence of the change. Coverage can show that a helper executed without showing that the tests exercised the distinguishing input or asserted the required rounding rule. Coverage is a common gate for AI-generated code: Futurum Research found that 58.6% of software lifecycle engineering decision-makers mandate automated test coverage thresholds for AI-generated code reaching production.[1]
AI-generated tests don’t remove that limitation. When an assistant writes both the implementation and its tests from the same mistaken assumption, they can agree perfectly and still violate the requirement. Expected results need an independent basis in requirements, accepted behavior, or another appropriate reference.
The operational problem is making the investigation happen consistently across a growing volume of changes, without requiring the most experienced person on the team to reconstruct the system for every review. Left untraced, those effects accumulate into architectural drift.
Each common check establishes something specific, and leaves something open:
Check | What it establishes | What it leaves open |
Diff review | What text changed | Which workflows depend on the changed code |
Tests | Whether the exercised cases produce expected results | Cases no test exercises, such as 2.345 in this example |
File-scope checks | Whether the edit stayed within allowed files | Effects on workflows outside the edited files |
AI code review | A model’s assessment of risk or scope | A source-traced record of what changed and who accepted it |
Behavioral comparison against a baseline | Changed rules and affected relationships, traced to source | Behavior outside the analyzed scope, which stays explicitly unverified |
Compare behavior against an established baseline
Detecting unintended change requires two references: what the software currently does and what the proposed change is authorized to alter.
Within a defined analysis scope, source can be used to establish the rules, relationships, and behavioral semantics represented by the software. Authorization comes from requirements and accountable decisions. Existing software contains defects and obsolete rules, so preserving everything indiscriminately would be a poor governance policy.
A behavioral baseline makes the comparison explicit. For the rounding example, it records the existing operation: monetary values are normalized to two decimal places using HALF_UP, with identified relationships to the invoice, settlement, and fee workflows.
The proposed revision changes one part of that record. Scale remains two decimal places. The rounding rule becomes HALF_EVEN. The identified callers remain connected to the helper, which means their unchanged source does not exempt them from impact review.
This is the mechanism that matters: derive the behavior represented by each revision, compare the rules and relationships, and evaluate the differences against the authorized change.
CodeIntent maintains a governed baseline grounded in source-derived semantic evidence, so each proposed change can be evaluated against an established, accepted reference, with prior acceptance decisions available for review.
Caller search contributes to the investigation. By itself, it doesn’t establish which rule changed, whether each affected workflow permits the new behavior, or who accepted the difference. Those connections need to become part of the review evidence.
Make the acceptance decision inspectable
A useful record separates what analysis established from what a reviewer decided.
Illustrative review record, not captured CodeIntent output. The boundary input and test-gap assessment are supplied for this example.
Review item | Finding or decision |
Authorized scope | Adjust the invoice calculation; preserve existing settlement and fee behavior |
Behavioral difference | Shared rounding rule changed from HALF_UP to HALF_EVEN |
Traced relationships | Invoice, settlement, and fee workflows reference the helper within the analyzed scope |
Distinguishing input | Exact decimal 2.345 produces 2.35 before the change and 2.34 afterward |
Test evidence | Existing tests in this example do not exercise an input that distinguishes the rules |
Required action | Restore the shared rule, or explicitly authorize and validate the affected behavior changes |
Reviewer’s merge decision | Hold while required findings remain unresolved |
The record should identify the compared revisions and link findings to their source. Test results should reference the tests that produced them. Approvals should identify the decision-maker and apply to the revision actually reviewed.
This gives an engineering leader an inspectable account of the behavioral change and its authorization. It also makes review more focused. Instead of asking someone to “take another look,” the team can ask the settlement owner to resolve a specific change to a specific rule, supported by a concrete input and result.
Direct review toward the changes that need it
A governance process that sends every shared-helper edit into a manual queue will become a bottleneck. The evidence should determine where additional review is necessary.
If the comparison establishes that a formatting change leaves the represented rules and relationships unchanged, there is no behavioral exception to authorize. If the required checks pass and the analysis has no blocking gaps, policy can allow it to continue through the normal approval path.
The rounding change produces a different outcome. It alters a shared rule outside the requested scope. That calls for correction or targeted authorization by the owners accountable for the affected behavior.
Declare the intended change and protected behavior. Establish what may change and who can authorize an exception.
Compare the accepted and proposed revisions. Identify changed behavior within the governed scope.
Trace the affected relationships. Include consequences in code that wasn’t edited.
Validate the differences. Use requirements, distinguishing inputs, tests, and appropriate technical review.
Apply the acceptance policy. Let sufficiently supported changes proceed. Hold changes with unresolved required findings unless an authorized exception applies.
Retain the evidence and advance the baseline. Verify the accepted revision and preserve the decisions supporting it.
The benefit is a more precise use of engineering judgment. Reviewers spend their attention on identified differences and unresolved questions instead of repeatedly assembling the context needed to discover them.
What CodeIntent adds
CodeIntent gives reviewers changed rules and affected relationships with evidence traced to source. That reduces the context they have to reconstruct for each pull request and gives the team a record it can continue to use as the software evolves. See how CodeIntent works.
Applied to this example, a behavioral record would capture the normalization rule, its scale and rounding mode, and its relationships to invoice, settlement, and fee calculation. Comparing revisions would expose the changed rounding rule while retaining the relationships that make its impact visible.
The record persists as the code changes. The next review starts from an accepted reference, and the evidence behind previous decisions remains available. Engineering can evaluate successive changes against an explicit account of the behavior it has accepted, rather than treating every pull request as an isolated investigation.
With the same governed revisions, analysis version, and rules, the deterministic analysis result is repeatable. Model-assisted assessments can help explain or investigate findings, but they remain distinguishable from source-derived evidence.
The analysis boundary also stays explicit. If the rounding mode depends on deployment configuration unavailable to the analysis, the effective rule remains unverified. Additional evidence or an authorized exception may support acceptance; neither silently turns an unknown into a verified result.
For the next change to shared business logic, require the affected workflows, a demonstrated behavioral difference where one exists, and the authorization supporting acceptance. If required evidence is missing and no authorized exception applies, the governance result remains unresolved, even when every test is green.
Bring a pull request that touches shared business logic to a CodeIntent walkthrough. Examine the baseline comparison, the affected relationships, and what remains unresolved before acceptance.
Share this post
Stay in the loop
Get new writing on deterministic modernization, evidence, and governed AI.
Related articles

AI Is Shipping Your Code Faster. What Does It Take to Accept It?
AI can write the code faster than any team can fully account for it. Each accepted change looks fine on its own, but architectural drift, and the understanding debt underneath it, compounds quietly until it doesn't.

Why COBOL to Java Migrations Stall at Sign-Off, and How Bounded Verification Fixes It
AI can generate COBOL to Java migrations fast, but production sign-off still requires a human reviewer's certification. This piece explains why unbounded review breaks down at scale, and how a bounded, source-traced evidence trail makes sign-off possible again.

Evaluating AI Modernization Vendors? Ask What Replaces the Two-Year Parallel Run
Prompt-derived AI can translate legacy code fast, but it can't replace the multi-year parallel run enterprises rely on for proof. Here's what can.
1. Futurum Research, 2H 2026 Software Lifecycle Engineering Global Enterprise Decision Maker Survey Report, July 2026 (subscriber access): app.futurumgroup.com/share/OPSH6u8XSAxYTgB7
Supporting public source: Futurum, “Software Lifecycle Engineering Market to Reach $226 Billion by 2030,” announcement of the 2H 2026 research, July 13, 2026: futurumgroup.com/press-release/software-lifecycle-engineering-market-to-reach-226-billion-by-2030
HOLONIC
The deterministic evidence layer underneath legacy modernization and the agentic enterprise.
© 2026 Holonic Technologies, Inc. · Atlanta, GA · Tucson, AZ
CodeIntent® is a registered trademark of Holonic Technologies.