We Thought TDD Would Make AI Code Reliable. Our Tests Had Holes.

We can repair the safety net by fixing the contract first and testing where it could fail.

11 min read

We thought we had found a way to get consistently solid software from AI coding agents: make them use test-driven development.
If the agent writes the tests too, reaching 90 percent coverage or more can look almost routine.
That sounds like the safety net we have been waiting for.

Then I realized it was not that simple.
In many cases, an AI agent can write poor tests, get them from red to green, and still fail to prove the behavior we wanted.
The suite can be green and coverage can be high while the code is wrong at a boundary nobody described.
The safety net exists, but it has large holes.

This does not make testing useless.
It means we need to adapt the process when the same agent can interpret the requirement, write the implementation, write the oracle, and explain why everything is correct.
We need to supply stronger intent before the code, choose evidence that reaches the real risk, and challenge the completed lot against the contract afterward.

Office worker on trapeze platform above green net with two holes; caped figure spots gaps; objects fall through; 90% gauge.
The safety net exists, but it has large holes.

TDD still gives us a starting point

Martin Fowler describes TDD as a cycle in which a test expresses the next piece of functionality, the code makes it pass, and the developer refactors while maintaining an evolving list of cases.
That red, green, refactor loop remains useful because it turns an expectation into executable feedback before more code accumulates.
Fowler's description of TDD also makes the case list central, which matters because the quality of that list determines what the suite can detect.

Our earlier workflow used TDD-inspired guidance, but our audit does not establish that every test was written before the corresponding code.
The agent often delivered code and tests together in one reviewed lot.
Calling that strict TDD would invent a chronology we did not observe.

The deeper problem is independent of chronology.
When one reasoning process writes both the implementation and its tests, both can preserve the same mistaken assumption.
A test can execute every branch and still assert the wrong outcome.
Coverage tells us which code ran, not whether the requirements were correct, complete, or meaningfully asserted.

Consider an illustrative example rather than an event from our pilot.
A test may initially fail because its setup references a missing symbol or constructs invalid data.
The agent can repair that setup, turn the test green, and leave the intended behavior untested.
The color changed for a real technical reason, but the red did not prove that the product behavior was absent and the green did not prove that it was implemented.

Benchmarks should narrow the claim

Kun Chen reported an October 8 experiment with Sonnet 5.5 high on DeepSWE in which banning newly written tests produced a slightly higher success rate, although the difference was not statistically significant.
He reported statistically significant reductions in time and token use.
The baseline split was 65 percent unit tests and 35 percent integration tests.
On a 44-task subset where existing tests were disabled, he reported no measured difference in success.
Only 17 end-to-end tests appeared beside more than 3,000 unit and integration tests, so the experiment did not settle the end-to-end question.
His clarification concerned tests that agents add spontaneously, without explicit human guidance on what to test.
These are Chen's reported results from his archived statement, not an independent replication or a review of the raw protocol.

A separate 2026 study examined trajectories from six models and prompt interventions across four models.
It found no statistically significant outcome changes from those interventions in its setting, which does not establish equivalence or remove the long-term value of regression protection.
The study on agent-generated tests is useful evidence against sweeping promises, not evidence that tests never help.

A green suite missed two boundaries

One reviewed lot added response tracking to an internal dashboard.
The agent delivered the code with five new tests, and the targeted suite passed 47 tests.
Those tests were not decorative.
They checked that an abandoned draft was handled explicitly, an external send receipt remained idempotent, inconsistent identities were rejected, normalized message references could attach a reply, and ambiguous references were refused.

An independent diff review still found two blocking defects.
First, a thread identifier from one account could attach an incoming message to another account.
Second, a UTC ISO timestamp initialized a datetime-local field and was then interpreted as local time, shifting the recorded instant by one or two hours depending on the time zone.

Two additional targeted tests covered those boundaries after the fixes.
The suite then passed 49 tests, with seven new tests in total.
The original tests remained useful evidence for receipts, idempotence, identity checks, and ambiguity.
They had simply omitted two partitions that mattered to delivery.

The integration boundary also remained incomplete.
The thread tests injected already normalized In-Reply-To and References values directly into the database mutation.
They proved the mutation's behavior but bypassed the actual adapter-to-worker chain.
Adding more tests at the same mocked boundary would not by itself prove that the real chain preserved the headers correctly.

No finite suite is exhaustive, so “write more tests” cannot be the whole answer.
We should target the missing partitions, failure positions, representation changes, and integration boundaries that could reverse a delivery decision.
That means asking what happens across accounts, time zones, permissions, retries, partial failures, protocol adapters, and external effects.

Stronger techniques need a sound oracle

Black-box and end-to-end tests can exercise a public interface and observe actual effects.
They are valuable when internal mocks erase the boundary where data changes shape or a remote system behaves differently.
Our Safari testing postmortem shows why the environment behind that interface matters.
They are not a magical source of truth, because an end-to-end assertion can encode the same false expectation as a unit test.

We can also separate the context used to design tests from the context used to implement the change.
Both contexts should refer to a fixed human-approved contract, representative examples, and counterexamples.
Using a different model or a fresh context may reduce correlated implementation details, but it does not create independence if both receive the same ambiguous requirement.
We have not tested whether a stronger model changes this result.

Property tests can explore broad input spaces around invariants such as idempotence, ordering, isolation, or conservation.
Mutation testing can show whether a suite notices controlled changes to the implementation.
Both improve sensitivity, but sensitivity to change is different from agreement with intended behavior.
A false oracle can reject many mutations while still defending the wrong rule.

We need a contract-guided hybrid process

Before code, we should record the source of human intent in an accepted contract.
The contract should include representative cases, counterexamples, invariants, and the consequence of failure.
We should resolve ambiguity only when it can materially change the implementation or the decision to ship.

A useful counterexample names a failure position rather than saying “handle errors.”
For an email reader, the third-party client may connect successfully and then reject mailbox lock acquisition.
The required invariant can state that the connection must still close.
That case distinguishes an implementation that protects cleanup only after lock acquisition from one that protects every resource acquired after connection.

During development, the agent should choose unit and integration tests that match the risk.
It should not mock away the boundary under examination.
For a reproducible bug, the new test should be red for the expected reason and turn green because the same scenario now behaves correctly.
Existing gates stay active, even when changing an assertion would be the fastest route to green.

When a test fails, we should decide whether the code or the oracle is wrong by returning to the source of intent.
Green pressure is not a product decision.
If the source says that a limit must remain within 1 and 100, clamping may satisfy the contract even when a speculative test expected rejection.
That test should be corrected or discarded with the reason recorded.

After the lot, the supervising agent should map each material behavior to evidence and identify omitted risks.
The normal review of the diff remains the default.
An independent review should appear only when the risk can change the delivery decision, such as account isolation, authentication, data loss, concurrency, migrations, or paid external effects.
For a bounded lot, we cap this stage at one corrective pass.
If a blocking defect remains, we report it and stop rather than declaring success or starting another review loop.

Some tests may be written after the implementation when inspection reveals the discriminating boundary.
That is not strict TDD when the entire test is written afterward.
The process is a hybrid: human intent and decisive examples come first, useful red-green cycles guide development, and a final evidence review probes what the completed lot still fails to demonstrate.

AI coding agent workflow linking agreed behavior contract, test-driven development, and post-implementation review to close test coverage gaps
Green suite, 90% coverage, wrong boundary. The agent graded its own homework and gave itself an A.

Define the behavior before development, then challenge the evidence before you ship.

This practical workbook helps you answer that question on your next AI-assisted change.
You will learn to set the expected behavior before coding, challenge the tests after development, and make a release decision you can explain.

Put the workbook to work on your next change

Put the workbook to work on your next change

Our pilot found one cleanup defect

On October 9, we tried this process on a small email-reading library whose baseline was green with 57 tests and 151 assertions.
Before delegated inspection, we fixed six requirements covering activation at the current watermark, ordered cursor progress, UID validity changes, header preservation, bounded reads, and resource cleanup without marking messages as read.

The review considered nine scenarios through a mix of execution and source inspection.
They were not nine new tests.
Existing tests already supplied useful evidence for watermark behavior, UID order, UID validity resets, MIME parsing, and truncation.

Three new probes exercised the real public reader methods with a fake IMAP client.
The probe where fetch failed passed because the lock was released and logout was called.
Two probes failed when connection succeeded but mailbox lock acquisition was rejected, one through activation and one through incremental reading.
Both failures came from one cleanup defect repeated in two public methods: lock acquisition happened before the protected try/finally, so logout was never called after the rejection.

A temporary counterfactual correction moved lock acquisition inside the protected region and made lock release conditional.
All three probes then passed instead of one passing and two failing.
That result supports the local causal explanation, but the correction was not integrated or deployed.

The fake client still limited the evidence.
The probes used real reader methods, but they opened no real socket, reached no real server, changed no production mailbox, and performed no MIME parsing before the failure.
The supervising session verified that the real dependency can reject mailbox opening without disconnecting, but the pilot did not test the effect through a live server.
A known repeated-headers defect was excluded, and a proposed 0 or 101 limit rejection was discarded because clamping met the actual contract.

The pilot was guided by code inspection, so it was neither blind nor a causal comparison of workflows.
The delegated run took 3 minutes and 42 seconds, while the total study time, parent tokens, and monetary cost remain unknown.
We measured no savings and no general reliability improvement.

The two cases left different gaps:

  • The dashboard suite passed 47 targeted tests after 5 new tests. Independent review found 2 blocking defects at account and time boundaries. After correction, 49 tests passed with 7 new tests. The full adapter-to-worker chain remained untested.
  • The email library passed 57 tests and 151 assertions. Review covered 9 scenarios and added 3 cleanup probes. One defect caused 2 failures, and the temporary correction made all 3 probes pass. Those probes did not exercise a live server, production mailbox, or MIME parsing.

The faulty methods had 100% line coverage

We then measured coverage on the exact code revision used in the pilot, with Bun 1.4.0 and the same four existing test files.
The 57 tests still passed.
The two affected methods already had every line reported by the coverage tool marked as executed: 14 out of 14 for activation, and 60 out of 60 for incremental reading.
Both methods still contained the cleanup defect.

The measured scopes were different:

  • The 8 source files present in the baseline LCOV report totaled 1,000 covered lines out of 1,814 reported lines, or 55.13% when combined by line count.
  • The complete reader file had 255 covered lines out of 373, or 68.36%.
  • The activation method had 14 covered lines out of 14, or 100%.
  • The incremental-reading method had 60 covered lines out of 60, or 100%.

The 100 percent figure applies specifically to those two method ranges.
It does not describe the file or the entire project.
The eight-file aggregate excludes source files absent from the report, including the CLI.
Bun’s displayed overall figure was 74.18 percent, which matched the unweighted average of its file percentages.
The 55.13 percent above weights each reported line equally.
We keep the denominator visible so those figures cannot be interchanged.

Adding the three probes to unchanged code produced 58 passing tests and two failures.
Coverage stayed at 255 out of 373 lines for the reader and 1,000 out of 1,814 across those same eight files.
The probes exposed a missing behavior without adding a single covered source line.
The temporary correction then produced 60 passing tests, with 257 out of 375 reader lines covered.

The existing tests had executed connection, lock acquisition and cleanup in other circumstances.
They had not verified cleanup when lock acquisition itself rejected after connection.
Line coverage could not express that missing sequence.
The reporter supplied no branch-coverage data, and we make no claim about branch or path coverage.

Important tests need traceable evidence

One local case is enough to justify trying this process on the next comparable lot.
It is not enough to make it a global rule or claim that AI-generated code is now more reliable.

Our standard should be concrete.
Every delivery-critical test should trace to a stated behavior, fail for the right reason when it is introduced for a bug, exercise the boundary it claims to cover, and state what remains unknown.
Coverage, test counts, and green checks can support that evidence, but they cannot replace it.

We can retain useful existing tests while feeding the next pilot back into the process.
The goal remains more reliable code.
We will know that we are getting closer only when repeated, comparable lots show fewer serious acceptance defects without hiding the cost or the limits of the evidence.

Get the workbook, open the acceptance brief, and write the behavior you need before the next agent run.
Then use the after-batch review to make the agent's evidence visible before you accept its work.

Put the workbook to work on your next change

Put the workbook to work on your next change

These sources support the analysis


High test coverage from AI agents can hide real failures at boundaries nobody described. The demo-vs-product checklist in the welcome kit shows why tests alone are not enough, and what else must change before code ships.

→ Get the welcome kit

This post may contain affiliate links. If you click them, I might earn a small commission. It costs you nothing and helps me keep shipping quality articles for your reading pleasure.