This is a category where the thing being bought is not software. It is people, increasingly augmented by machines, and the distinction explains more about the market than any feature list will.
Forrester's own opening to The Forrester Wave: Continuous Automation And Testing Services, Q2 2024, authored by Diego Lo Giudice, makes the shift explicit. In the previous edition, in 2021, what clients wanted from these providers was intimacy, collaboration, and early AI experimentation. Three years later, one word changed the conversation: generative AI.
That is not a cosmetic update. It is the difference between a vendor promising to work closely with you and a vendor promising to make its own testers dramatically cheaper, faster, and more productive. The second promise is measurable. The first was not.
What these firms actually sell
Continuous automation and testing services providers do the unglamorous, permanent work of making sure software works before it ships and after it runs. They supply the testers, the automation frameworks, the tooling, and the delivery models that let an enterprise move from testing as a phase at the end of a project to testing as a continuous, embedded activity.
The category is called continuous for a reason. In agile and DevOps shops, the idea of a testing phase that happens after development is over has already collapsed. Testing has to happen in parallel, at every stage, and increasingly in production itself, which is a genuinely different technical proposition.
Forrester's inclusion criteria say a lot about what scale this operates at. To be evaluated at all, a provider must clear three hundred million dollars in annual revenue from these services alone, operate in at least two regions with meaningful share in each, and carry significant analyst or peer recognition. This is a market of large, global services organisations, not boutiques.
The report's framing of what it scored them on is equally direct. Three things: how fast they move clients from manual to automated testing, with reference clients reporting seventy to eighty percent automation; how well they deliver for modern application development; and how pragmatically they infuse generative AI into their own testing work and into testing AI-infused applications.
Inside The Forrester Wave: Continuous Automation And Testing Services, Q2 2024
The evaluation, published in June 2024, scored a field of global providers against twenty seven criteria across current offering, strategy, and market presence.
Three firms were named Leaders.
IBM placed as a Leader, credited for the strength of its AI research feeding a testing practice, its IGNITE accelerator, and its use of generative AI and Watson in application modernisation testing, particularly large-scale legacy migration from COBOL to modern stacks.
Infosys placed as a Leader, singled out for an AI talent strategy aimed at an AI-ready workforce, for folding AI into its testing accelerators, and for flexible outcome-based contracts with transparent pricing.
Qualitest placed as a Leader, the only pure-play testing specialist at the top, credited for a vision of zero-touch quality orchestration, strength in testing in production, and strong capabilities in testing AI-infused applications.
Among the Strong Performers, Accenture and Capgemini placed, with HCLTech and LTIMindtree among the Contenders.
The detail that says the most about this market
One line in the evaluation deserves more attention than the tier placements around it.
Accenture declined to participate in the full Forrester Wave evaluation process, and was scored under Forrester's vendor participation policy, marked as a nonparticipating vendor on the graphic.
Pause on that. The world's largest professional services firm, with a testing practice that dwarfs most of the Leaders in scale, chose not to open its books for this evaluation.
The reason is not knowable from outside, but the incentives are. Full participation means submitting reference customers, opening the practice to analyst scrutiny, and accepting a score one does not control. For a firm that sells testing as one line inside a much larger consulting relationship, the upside of a good score is modest and the downside of an awkward one is real.
For a buyer it is a signal, not a verdict. A vendor that declines to participate is not necessarily hiding weakness; it is declining to be ranked. The practical consequence is that the report tells you less about that vendor than about the ones that showed up, and any shortlist should treat that asymmetry as a fact rather than a moral judgment.
The genAI promise and where it strains
The strategic heart of this edition is generative AI, and the honest reading of it is more complicated than the enthusiasm suggests.
Providers in this category were, as Forrester notes, among the first anywhere to experiment with AI and machine learning, as far back as 2019. So the industry has not been caught flat-footed. What changed is that generative AI moved the goalposts from assistive tooling to a direct question about headcount productivity.
The promise vendors are making is that a tester augmented by a model is smarter, faster, and more efficient. The promise has a specific economic shape: if a model lets one tester do the work of two, the provider's margin improves or its price falls, and the client captures some share of that.
The strain appears in the reference feedback. Even IBM, a Leader with the deepest AI research bench in the field, drew the observation that scale challenges prevent consistent delivery of its depth to all clients, and that reference customers do not consistently receive the benefit of its full capabilities. That is the services reality in one sentence: a provider's best capability and what a given engagement team actually delivers are two different things, and they drift further apart the larger the firm.
It is also why the pure-play Leader is worth noticing. Qualitest does nothing but testing, so its entire reputation rides on this one category in a way IBM's or Accenture's does not. For a buyer who wants testing to be the vendor's whole attention rather than one line item, that concentration matters.
The inclusion of AI-infused applications
The criteria include something that barely existed in the 2021 edition: testing AI-infused applications.
This is the quietest and most consequential addition. A conventional application produces deterministic outputs, and a tester verifies them against expected results. An AI-infused application produces probabilistic outputs, and the entire notion of an expected result changes.
Testing it means building new techniques: evaluating whether a model's answers are acceptable rather than correct, whether a retrieval pipeline returns the right sources, whether a generated response stays inside a policy boundary, and whether the system behaves the same way under drift.
The vendors scoring highest on this criterion are effectively claiming to have solved a problem the rest of the industry has not yet agreed on how to define. That is either the most forward-looking part of the report or the most speculative, depending on how much one trusts the claim.
Where this leaves a buyer
This is a services purchase, so the questions that decide it are different from a software purchase.
Who will actually be on your engagement, not who leads the firm's testing practice. The report can tell you IBM has depth; it cannot tell you whether your account gets the senior testers or the bench. Reference calls should ask that question specifically.
What is the demonstrated automation rate on accounts that look like yours, with the client there to confirm it. The seventy to eighty percent figure is an aggregate; your industry, your legacy estate, and your release cadence will move it.
And how is generative AI priced, because this is where services contracts get quietly rewritten. A provider that trains models on your code and data creates an asset question, a data question, and a pricing question all at once, and none of them are settled in this market.
The category has not so much changed as been redefined around one capability. What was once a relationship business, sold on intimacy and collaboration, is now being sold on the claim that machines make the people measurably better. The firms that can prove that claim on your specific work, not just in their marketing, are the ones worth a shortlist slot.
The throughput this automation frees up still has to be measured somewhere; see Software Engineering Insights Solutions, the renamed successor to value stream management, which scores that exact capability and flags the risk of turning the measurement against the people it is meant to help.
Analyst Source
Forrester Research
Category definition, vendor inclusion, and evaluation findings in this article draw on Forrester's coverage of continuous automation and testing services. The Q2 2024 Wave, authored by Diego Lo Giudice with Paul McKay, Min Say, and Kara Hartig, scored a field of global providers against 27 criteria across current offering, strategy, and market presence, following an earlier evaluation of this market in 2021. Forrester's inclusion criteria for the category require a minimum of $300 million in annual continuous automation and testing services revenue.
Source research
Forrester does not endorse any vendor named here, and tier placement should not be read as a recommendation to buy.
A companion piece published since covers this from the adjacent category: DevOps Platforms, where forrester renamed this category after a keyword survey found zero of forty one vendor websites used its old name, and the new scorecard made built-in security a requirement just to be evaluated at all.