Back to blog

Engineering Got Faster. How Does QA Keep Up?

More changes arrive. The release date stays. QA needs more than a bug count to decide what deserves attention and whether the release is ready.

Mastra product PRs and repair work, February-August 2026. August is a provisional classifier estimate, not a reviewed BFTR result.

Imagine the release is tomorrow. Some changes are verified. Others are waiting. Another batch lands while QA is still working through the previous one. Do we move the date, remove changes, or ship with some behavior unverified? Who accepts that risk?

For a QA leader or release manager, this is where faster engineering becomes a difficult release decision. Claude Code, Codex, and other coding agents help developers produce implementations. When that increases the pace of change, the work of understanding and verifying those changes does not disappear. It arrives faster too.

I want the same ambition applied to the people responsible for release confidence. Not another expectation that QA will somehow absorb the extra volume. Tools that help them understand what the release could affect, examine the evidence, and direct their attention.

Our retrospective study of Mastra follows product PRs from February through August 2026. February had 381 eligible product PRs, including 219 fix-related PRs. The provisional August figures are 893 product PRs, including 528 repair candidates. We count eligible product work, not all repository activity. This is one project’s history, not proof that AI caused the growth or that QA staffing stayed constant.

Mastra product PRs split into repair work and the remaining eligible PRs. February-July totals: 381, 520, 440, 571, 547, 733, including 219, 320, 265, 271, 303, 500 fix-related PRs. Provisional August: 893 proposed eligible PRs and 528 repair candidates, 59.1%, not BFTR-reviewed results.

Mastra product PRs split into repair work and the remaining eligible PRs. February-July totals: 381, 520, 440, 571, 547, 733, including 219, 320, 265, 271, 303, 500 fix-related PRs. Provisional August: 893 proposed eligible PRs and 528 repair candidates, 59.1%, not BFTR-reviewed results.

Source: Early's retrospective Mastra PR study, evidence cutoff September 6, 2026. February-July PRs were reviewed for eligible product work and classified as repair-only, mixed feature/repair, or other work. Repair shares combine repair-only and mixed PRs, divided by eligible product PRs. August uses provisional classifier estimates, not the completed review. These are PR counts, not unique bugs, regression rates, effort, or Early detection performance.

QA has tools. The gap is context.

QA teams already invest in automation, CI, test environments, and repeatable checks. The problem is not that they have ignored tooling. Running those checks is only part of the job.

Someone still has to establish what is in the candidate, understand which existing behaviors might be affected, and decide where deeper verification is necessary. That includes behavior in components nobody edited. As the candidate changes, someone must also decide whether yesterday’s results still apply.

Generating more tests can help. But it does not automatically establish that the tests challenge the right assumptions. A new test can agree perfectly with an implementation that misunderstands how another component uses its output.

Before deciding what to test, someone has to understand what this release could change. That is work we should help QA do, not leave invisible behind the test count.

A bug list is not a release assessment

Repair work accounted for 57.5% of eligible product PRs in February, 68.2% in July, and a provisional 59.1% in August. The share fluctuates even as the volume grows. Delivery means managing repairs alongside new changes.

Suppose an analysis finds three issues and the team fixes them. What does that establish about the rest of the release? Do the results still apply after those fixes?

If it finds nothing, I need the same context: what was examined, which dependencies were included, and what remains unresolved.

Release confidence needs evidence tied to the candidate, not just a finding count. An empty report and a completed bug list both leave that question open.

A fix can change more than intended

The team fixes a problem. The next release includes the fix. For the person responsible for signing off, one question remains: what else changed along with it?

Our Mastra Case File shows why that question matters. A fix intended to preserve saved messages also caused other expected handling to be skipped when a model call failed. One problem was addressed, but behavior outside that problem changed too.

That is the burden on QA and release leaders. A ticket can be closed while the release still carries an unanswered question. They need help seeing beyond the fix, not just confirmation that it is there.

The Mastra Case File walks through the evidence in an independent historical replay.

What I would want before saying ship

I would want clear answers to four questions, not another report that requires an afternoon to interpret.

  1. What did this release touch, and what stayed unchanged?
  2. What changed unintentionally?
  3. What needs fixing, and what will it take?
  4. Do we understand this release well enough to ship with confidence?

Passing tests, security checks, and trustworthy build provenance are necessary controls. They do not, by themselves, prove that dependent behavior remained intact. Each contributes evidence about a different part of the release.

No single tool supplies every answer.

Give QA a basis for prioritization

Evidence should direct the next action: a focused test, a dependency owner’s review, or an explicit decision about unverified behavior. If a critical question stays open, reduce scope, move the date, or have the responsible owner accept the risk. QA should not silently inherit that decision because the deadline arrived.

Early’s Regression Guard contributes by comparing a candidate with an approved baseline and identifying affected and changed behaviors. It can surface downstream regressions across configured component relationships, not automatically discovered ones. That supports the decision without replacing testing, security review, runtime observation, or human approval.

A release candidate passes through a scan, then issue and no-issue-detected paths converge on confidence. Confidence depends on what was examined, not only what was found.

A release candidate passes through a scan, then issue and no-issue-detected paths converge on confidence. Confidence depends on what was examined, not only what was found.

Confidence depends on what was examined, not only what was found.

Engineering has tools to produce changes faster. QA and release leaders need help understanding those changes and fixing the issues they introduce at the same pace. Whether we find issues or not, show me what was examined and what remains uncertain. That is a basis for a release decision, not a promise of safety.

On this page

Read next

The Code Is Ready. The Team Is Still Figuring It Out.
The Code Is Ready. The Team Is Still Figuring It Out.
Coding agents can finish the implementation before the team has answered the questions that determine whether it is ready to release.
AI Code Review Is Not Release Verification
AI Code Review Is Not Release Verification
A clean pull request is evidence about the change. It is not evidence about every behavior the release could affect.
New in Early: Home Dashboard, Outcome Report, and Usage Report
New in Early: Home Dashboard, Outcome Report, and Usage Report
A daily view of release risk, analysis freshness, finding outcomes, and verification activity.

Understand the behavior behind your next release