Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    DevOps
    CI/CD
    Testing
    QA
    Pipeline Optimization

    Fix a Slow CI Pipeline: Cutting 58 Minutes to 14 by Fixing QA

    January 20, 2026
    9 min read

    If your CI pipeline takes close to an hour, people will tell you the same three things every time: add more runners, throw better hardware at it, and cache everything that doesn't move. Sometimes that works. A lot of the time, it just makes an already messy system slightly faster at being frustrating, because the problem usually lives somewhere less exciting, in QA.

    That was the case for one team whose pipeline averaged 58 minutes end to end. On paper, nothing looked especially broken. Builds were fine, and so were the deploy steps. The bottleneck sat squarely in automated tests, which were eating 42 minutes of every run.

    They deployed roughly ten times a day, so the math on what that kind of latency does to a team isn't hard. There were context switches everywhere and PRs stacking up, and people pushed changes and walked away because "CI will take a while anyway."

    A month later, the same pipeline averages 14 minutes. Developers wait for it again, and they trust the results. Faster machines and exotic tooling had nothing to do with the biggest gains. Those came from being honest about what tests are for and when you actually need them.

    Long pipelines change behavior as well as slowing code

    Slow CI carries a psychological tax that rarely shows up in dashboards. When feedback takes an hour, developers stop caring about it in real time. They open a PR, kick off the pipeline, and move on to something else. When it fails, the failure feels disconnected from the change that caused it. That's how broken tests get retried without investigation, or worse, ignored.

    Over time, teams adapt in unhealthy ways. People batch changes to "make the wait worth it." They merge late in the day and hope nothing explodes overnight. They treat CI as a formality instead of a safety net.

    The irony is that these behaviors make pipelines slower and flakier, which reinforces the cycle. QA turns into something that happens to the code, when it should be something that helps it.

    Not all tests deserve equal treatment

    The first big change this team made was also the most controversial internally. They stopped pretending that every test needed to run on every pull request and split their test suite into two clear tiers.

    The critical suite runs on every PR. It covers authentication, payments, and the core user flows that would be catastrophic to break. It takes about seven minutes when run efficiently, and if this suite fails, nothing ships.

    The full suite still exists, but it's no longer in the hot path. It runs nightly and on release branches, where catching obscure edge cases is worth the extra time.

    The goal was to match feedback speed to risk without lowering quality. Most changes don't touch the deepest corners of the system, and forcing every developer to wait for exhaustive coverage on every commit doesn't make software safer. It just makes teams slower and more cynical. Once that mental shift clicked, everything else became easier to justify.

    Parallelization is obvious until you try it

    Everyone knows parallel tests are faster. Actually making them work is another story.

    The critical suite was originally running serially, which meant every test inherited the sins of every test before it, and one slow setup step could ripple through the entire run. By splitting the suite across six runners, the team cut execution time almost in half immediately. That was the biggest single win on paper.

    Parallelism comes with sharp edges, though. Tests that share data, rely on global state, or make assumptions about execution order tend to fall apart when run side by side. Fixing that meant changing the test design itself, and the CI configuration was the easy part.

    Each test had to be responsible for its own setup and cleanup, and shared resources needed isolation. A few tests that were "fine" in serial turned out to be brittle all along. This was real work, and it paid off beyond speed, since tests that can run in parallel are usually better tests.

    Flaky tests are worse than slow ones

    If slow pipelines teach developers to ignore CI, flaky pipelines teach them not to believe it.

    Before the overhaul, about 18 percent of failures were false positives: random timing issues, UI tests breaking because a button rendered a few milliseconds late, and failures that vanished on rerun. That number is brutal. It means developers are conditioned to assume CI is lying to them almost one out of every five times.

    The team attacked this directly by replacing their worst offenders. Some legacy browser tests were rewritten using more stable tools. Others were rethought entirely, shifting away from brittle UI checks toward approaches that validated behavior without pixel-level precision.

    The result wasn't perfect, but it was dramatic. False failures dropped to roughly three percent, and that change alone probably saved more developer time than the raw speed improvements.

    Auto-retry is fine if you treat it honestly

    One of the more pragmatic choices was adding automatic retries for single-test failures. If a test fails once and passes on retry, the pipeline doesn't block the PR, but it does flag the test, so nothing gets swept under the rug.

    This approach acknowledges reality. Some failures really are random, infrastructure hiccups happen, and timing issues sneak through even well-designed tests.

    Accountability is what makes it work. Retried tests are tracked and reviewed, and patterns emerge quickly. Tests that need frequent retries become candidates for fixes or removal, so the retry mechanism works as a diagnostic tool instead of a crutch.

    Without that follow-up, retries can absolutely hide real regressions. With it, they reduce noise while still pushing the system toward stability.

    Speed changed trust, and trust changed everything else

    Once the pipeline reliably finished in around 14 minutes, developers started waiting for CI again. They'd open a PR and actually watch the checks complete. When something failed, it felt immediate and actionable, and fixing a broken test no longer meant losing an hour of momentum.

    That trust had knock-on effects: smaller PRs, more frequent deploys, and less temptation to push risky changes late in the day. The pipeline stopped being an obstacle and started feeling like part of the workflow.

    It also made conversations with management easier. A fast, reliable pipeline is easier to defend than an expensive one that still frustrates everyone.

    Focus cost more than compute

    This entire effort took about a month. The ideas weren't complex, but they required sustained attention across tooling, tests, and team habits.

    That's the part many organizations underestimate. Buying faster runners is easy. Carving out time to fix test design, reduce flakiness, and rethink QA strategy is harder, because it competes with feature work and doesn't ship anything visible to customers.

    And yet, the return on investment is massive. Faster feedback loops save mental energy and reduce friction on top of the minutes, and they make engineering feel less like waiting and more like building.

    Fix QA first, then worry about everything else

    If your CI pipeline is slow, it's tempting to start with infrastructure: scale up, optimize caches, shave seconds off builds. But if QA dominates your runtime, that's where the leverage is. Ask which tests actually need to run, which ones are flaky and why, and whether your test suite reflects how your team ships software today or how it shipped three years ago.

    Most teams have a trust problem more than a tooling problem. Fixing QA will make your pipeline faster, and it will also make it worth paying attention to again. That's when CI starts doing the job it was always supposed to do.