Yeison Daza
4 min read

Tests need to earn their place in my codebase

Over the last few weeks, there has been a lot of discussion about tests. Now that AI can write them too, it is easy to end up with a test suite that looks healthy: lots of cases, good coverage, and green checks in CI.

But that does not necessarily mean it is protecting us from bugs.

A test can be nothing more than an echo of the implementation. It can check details of how the code is written without checking the behavior that should remain true. If the implementation changes, we change the test with it. And if we introduce a bug, everything might still pass.

I have seen two responses to this:

  1. Delete all the tests in the codebase.
  2. Stop writing unit tests and replace them with E2E tests.

Both have pros and cons. But I have come to a different conclusion:

Tests need to earn their place in my codebase.

A test should give you confidence

For me, a good test is one that gives me confidence when the code changes.

It should help me catch bugs when someone changes part of the system without having all the context. It does not matter whether that someone is a person, an AI, or myself six months from now.

Code changes. We refactor, add features, and fix bugs. Tests should prove that they can catch changes that break important behavior.

They should not exist just because they pass in CI, maintain coverage, or were written at some point in the past.

Challenging tests

That is why I implemented mutation testing in my project.

The idea is simple: introduce deliberate changes to the code and then run the tests. For example, change a condition, invert a comparison, change a returned value, or remove a call.

If no test fails after that change, the mutation survives.

A surviving mutation does not automatically mean there is a problem. It might be an aesthetic change or a part of the code with no real impact. But it is a useful signal: there is probably behavior that our tests are not verifying.

Instead of asking how many lines we cover, we can ask a more useful question:

If this code changed in an incorrect way, would our tests notice?

Making it work at scale

Mutation testing can be very expensive. Mutating every file and running every test after each mutation, on every PR, is not feasible. It would be too slow and would end up blocking the development workflow.

My approach is to use it as an additional process in CI:

  1. Read the changes in the PR.
  2. Map those changes to the related test cases.
  3. Mutate only the affected code.
  4. Run the relevant tests.

This makes the process take around five minutes in CI.

I also made an important decision: it does not block anyone. The PR can be merged and the team can keep working. When mutation testing finishes, it posts a report directly on the PR with the mutations that survived.

I do not want to turn another metric into a barrier to shipping code. I want a signal of where our test suite has blind spots.

Closing the gaps

Finding surviving mutations is only the first part. The useful part happens next.

I have a scheduled task that, every morning, reviews the reports generated the previous day. For each mutation, it decides whether it represents a real gap or an irrelevant change. When it is a real gap, it adjusts the tests to cover that behavior.

The result has been closing around 100 gaps in the test cases every day. And almost every day, the process finds real bugs in the codebase.

Not because the tests were useless, but because they were not yet protecting everything we thought they were.

A continuous process

The point is not for tests to be a passive artifact that runs in CI, or for them to exist just to increase a metric.

Their job is to increase our confidence: if someone changes the code we wrote, they should catch that change and prevent us from introducing bugs.

With mutation testing, tests stop being something we write and forget. They continuously face deliberate changes, reveal their blind spots, and get better over time.

I am not looking for a perfect mutation score, or to block the whole team with another slow check. I am looking for a process that challenges tests every time the code changes and, day by day, makes it harder for a bug to reach production.

Tests need to earn their place in my codebase.