Beyond Code Review: What Makes a Release Ready?
Over 50% of estimated implementation effort went to repairs. What does that mean for release confidence?

Over 50% of estimated implementation effort went to repairs. Across four open-source projects, 60.6% of eligible product pull requests contained fixes.
A pull request can be approved while the release still needs work. For a QA or release leader, those are different decisions. One concerns a change. The other concerns the product people are about to use. Release verification examines the merged candidate before it ships.
That distinction matters when engineering can produce changes faster than teams can verify them. A growing list of approved pull requests does not tell a release owner which behaviors were examined, which issues need fixing, or whether the candidate is ready to ship.
We wanted to understand how much of the recorded change stream was corrective work, what teams were repairing, and what that tells us about release confidence. The findings make a case for looking beyond individual code reviews to the release itself.
Delivery includes substantial repair work
Our study examined Pydantic AI, Mastra, Zed and Astro, four selected, company-backed open-source projects. Across February-September 2026, 7,293 of 12,038 eligible product pull requests (60.6%) contained repairs. Documentation-only and other out-of-scope changes were excluded.
In February, the combined sample contained 673 repair pull requests. In September, it contained 1,357. That is approximately twice the monthly volume. The repair share also varied: 52.4% in February, 70.6% in August and 65.4% in September.
The distinction matters. A team can face more corrective work even when repairs take a smaller share of a larger delivery stream. The release workload is not captured by a percentage alone.
Four-project monthly product pull requests, February-September 2026. Repairs are included in eligible totals. The line shows repair share, not a regression rate.
These counts do not establish that AI caused the growth, or that code review failed. They show something more direct: a substantial part of the recorded product-change stream was work to correct defects, not simply to add capabilities.
Pull request counts do not capture the effort
A one-line correction and a difficult integration repair each count as one pull request. To look at the work behind the counts, we estimated implementation effort in story points and allocated it across repairs, features and other work.
Repairs accounted for 3,759.5 of 6,732 estimated story points (55.8%), compared with 30.8% for features and enhancements.
Allocation of 6,732 estimated story points across the covered August-September cohorts. Percentages use total estimated effort, not repair effort alone.
The effort analysis covers August-September cohorts, not all eight months. It combines retrospective, project-calibrated model estimates: full cohorts for Pydantic AI, Zed and Astro, and the first seven days of each month for Mastra. Story points describe relative implementation effort, not measured hours or financial cost.
For engineering leaders, the significance is capacity. Repair work can compete with roadmap work for investigation, implementation and validation. Reducing avoidable rework could give some of that capacity back, improve delivery speed and reduce customer disruption. This study does not measure how much of the repair effort could have been prevented or saved.
Where repair effort goes
We also classified 2,508 repair-containing pull requests from August-September by the principal type of problem corrected. Four categories accounted for 69.3%:
- Behavior and logic: 33.6%. The code’s behavior did not match what it needed to do.
- Integration and compatibility: 14.2%. Components, providers or environments did not work together as intended.
- Reliability and error handling: 13.0%. Failures, retries or exceptional conditions required correction.
- API and schema contracts: 8.5%. Interfaces or data expectations needed to be repaired.
August-September repair categories. Pull request shares use 2,508 repairs. Effort shares use 6,732 total estimated story points. One principal category per repair pull request.
Performance, security and other categories made up the remainder. The concentration in behavior and interactions is especially relevant to QA leaders. Reviewing the edited lines is not the same as examining the paths those lines affect, including dependencies and failure handling.
For release leaders, the categories help make the next question concrete: which of these paths were actually verified in the candidate we intend to release?
A fix is not necessarily a regression
A repair corrects a defect. A regression breaks previously working behavior.
In the detailed sample, 86 repair pull requests (3.4%) had evidence supporting regression attribution, and 101 (4.0%) had possible regression evidence. Together, that is 7.5% of the repair-containing pull requests. The possible cases remain unconfirmed.
For 91.4%, the available records did not establish the defect’s origin. That does not mean those defects were never regressions. It means the retained evidence did not settle the question.
The release problem is broader than regressions alone. A newly introduced bug, a longstanding defect and a regression can each need attention before shipping. Knowing the origin is useful. Understanding the behavior and deciding what to do about it is essential.
Code review is not release verification
Code review remains an important checkpoint. It lets a reviewer examine a proposed implementation, its intent and its design. Release verification asks a different question: what is the state of the assembled candidate now?
For QA leaders, that means visibility into the behaviors and interactions examined, not just a growing queue of changes waiting for validation.
For release leaders, it means a clear account of findings, unresolved risks and the actions required before release. An approval needs to be tied to the candidate, with an accountable decision owner and a practical monitoring and rollback plan.
For engineering leaders, it means finding and addressing issues while turning avoidable repair work into capacity for delivery. The historical counts show the scale of corrective activity. They cannot tell us whether a particular release is safe.
Security checks, provenance controls and passing tests are necessary. They do not, on their own, prove that dependent behavior stayed intact. Nor does a report with no findings establish that nothing was missed. Confidence depends on knowing what was examined and what remains outside that examination.
Four questions before release
The useful outcome is not another report that takes an afternoon to interpret. It is an inspectable basis for action:
- What are we releasing, and what did we verify? Identify the candidate, intended changes, affected behaviors and verification scope.
- What bugs, regressions, or risks need attention? Show the findings and the evidence behind them, including uncertainty.
- What needs fixing before we release? Decide which issues block shipping, verify proposed fixes and make deferred risk explicit.
- Do we have enough evidence to release with confidence? Make an accountable release decision, with monitoring and a workable rollback or mitigation plan.
That is the shift beyond code review: from approving a change to understanding what we are releasing, addressing the issues that matter and making the decision visible.
Explore the research
Project comparisons use matching pull request and effort coverage within each project. Purple denotes total. Amber denotes repairs. Percentages below repair bars show their share of each total.
The study combines an eight-month pull request analysis with a detailed August-September repair classification and a separate effort analysis. The populations and coverage differ. The detailed results are not additional records to add to the eight-month totals. Classifications used AI-assisted analysis of preserved repository evidence, with project-specific review and quality controls.
Get the research white paper by email. For full methodology, detailed results and supporting data, email contactus@startearly.ai and request the full report.
About Early
Early’s Release Guard analyzes changes after merge to identify new bugs, existing defects and regressions missed by earlier checks. It automatically opens pull requests with proposed fixes, helping QA, release and engineering leaders act before release and make informed release decisions.
This research was sponsored by Early.

