Skip to content
←Back to blog
buying guideautomated testingaxe-coreCIEAA

Build vs buy accessibility testing: what your team needs

The axe-core engine is free; the crawling, history and reporting around it are not. How to decide whether to build or buy accessibility testing.

P

Pavel Charkasau

Build vs buy accessibility testing is mostly a question about everything around the rules engine, because the engine itself is already free. axe-core, the open-source library most commercial and in-house scanners run underneath, is published under the Mozilla Public License 2.0 and covers WCAG 2.0, 2.1 and 2.2 at levels A, AA and AAA (axe-core). So nobody is really choosing between writing rules and paying for them. The choice is whether your team builds and maintains the parts that make those rules useful: reaching pages behind a login, running on a schedule, keeping a dated history, deduplicating findings across templates, and turning results into a report someone outside engineering can read. Build when your needs stop at your own pull requests and you have a developer who will own the setup. Buy when you need coverage of the whole site, evidence for an EAA or EN 301 549 statement, or reporting for people who will never open a CI log. Plenty of teams do both. Either way, automation finds issues and a person still has to review the rest.

What are you actually building when you build accessibility testing?

Less than you might fear for a first version, and more than you'd expect by the second year.

The first version is an afternoon. You add @axe-core/playwright to an existing test suite, point it at a few routes, and fail the build on violations tagged wcag2aa or wcag22aa. We walk through that setup in accessibility testing in CI. For a team guarding its own components, that may be all you need.

What comes after the first version is where the estimate usually goes wrong. Here's the list of parts people end up building once the check has been running for a few months:

  • A crawler, or a maintained list of URLs, so the check covers more than the routes someone remembered to add.
  • Authentication that survives SSO changes, MFA prompts and expiring test accounts.
  • A baseline, so the build fails on new violations and not on the 300 old ones nobody has fixed yet.
  • Deduplication. One broken header component produces the same finding on every page that uses it.
  • Somewhere to store results over time, with dates, so you can answer "when did this start failing?"
  • A way to record what a human checked and decided, because the engine leaves some results open on purpose.
  • A report format for whoever asks: a client, a procurement team, a regulator.

Each is easy alone. Together they are a small internal product with no product manager.

How much of the work does the open-source engine already do?

The hard part, which is why building is a reasonable option at all.

axe-core's README states two design facts worth knowing before you decide. It aims to return "zero false positives (bugs notwithstanding)", and with it "you can find on average 57% of WCAG issues automatically" (axe-core). The 57% figure comes from Deque's 2021 analysis of more than 2,000 audits covering about 13,000 pages and nearly 300,000 issues, and it counts issues by volume (Deque). Counted by WCAG success criteria a tool can evaluate, the same study cites the widely held estimate of 20 to 30%. That's why we quote roughly 30 to 57% and not the top of the range.

The rules also get maintained for you. Between 9 October 2025 and 5 October 2026, axe-core shipped nine stable releases, from 4.11.0 to 4.14.0 (npm registry). If you build, you inherit those improvements by bumping a dependency. You also inherit re-baselining when new rules fire, and explaining why last week's green build is red with no code change.

There's a standards angle too. The W3C's Accessibility Conformance Testing (ACT) Rules exist because different tools interpreted WCAG differently, and publishing shared test rules "reduces confusion caused by different interpretations of accessibility guidelines" (W3C WAI). A home-grown rule you write yourself sits outside that shared interpretation, so keep custom rules for your own design-system conventions and leave the WCAG ones to the engine.

What does a bought platform add on top of the engine?

Mostly reach and history, plus output that people outside engineering can read. That answer is narrower than most vendor pages make it sound.

The W3C's guide to selecting evaluation tools lists the differences that matter. "Some tools check a single page, while others check entire groups of related pages." "Some can also access password-restricted content." Tools can be integrated "into your web browser, content management system (C-M-S), and your development and deployment tools" (W3C WAI).

Build (engine in your own pipeline)Buy (platform around the engine)
Rulesaxe-core or similar, freeUsually the same engine, sometimes extra rules
Pages coveredThe routes your tests visitA crawl of the site plus configured journeys
Logged-in areasReuses your test fixturesStored credentials or recorded sign-in
HistoryWhatever you store yourselfDated scan history per page
Who reads the outputDevelopers, in CI logsDevelopers plus compliance, clients, procurement
Ongoing costEngineering timeSubscription plus some setup time

Notice what is missing from the "buy" column: a different answer to the 30–57% question. A platform built on the same engine catches the same machine-detectable issues. Anyone who tells you their tool covers far more of WCAG automatically should be asked to show how, and which criteria. Our scanner buyer's checklist has the questions.

When does building accessibility testing in-house make sense?

When the thing you want to protect is your own code, and the person who reads the result is the person who wrote it.

Building fits well when:

  • Your product is one application with a test suite people actually maintain.
  • The goal is catching regressions before merge, not reporting on the whole site.
  • Someone on the team will own the setup, including the dull work of updating baselines.
  • Nobody outside engineering needs to see the results.

In that setup, a component-level check with jest-axe plus a Playwright check on the five or six routes that matter most will catch a missing label or an unnamed button the day it's written. I'd recommend every frontend team do this regardless of what else they buy, because the pull request is where a fix costs two lines.

When is buying the better call?

When the question changes from "did this merge break something?" to "where does the whole site stand, and can we prove it?"

The signals are usually one of these:

  • Content changes outside your codebase. Marketing publishes through a CMS, or a third-party widget injects markup at runtime. Your test suite never sees it.
  • You need to cover journeys behind a login across many pages, and nobody wants to maintain those fixtures. We covered why the homepage is the wrong place to look in why homepage scans miss your real risk.
  • You are in scope of the EAA. Article 13(3) of Directive (EU) 2019/882 requires service providers to ensure "that procedures are in place so that the provision of services remains in conformity with the applicable accessibility requirements", taking account of changes to the service and to the standards (EUR-Lex). A dated record of what was checked and when is the easiest way to show that procedure exists.
  • Someone asks for a report. A client, an auditor or a public-sector buyer reading against EN 301 549 wants findings mapped to clauses, with the open items marked as open. Writing that formatter yourself is real work.
  • You manage many sites. Agencies hit the maintenance cost of a home-grown setup fastest, because every client site needs its own crawl, credentials and report.

What does neither option give you?

Conformance. Building and buying both get you the automated half, and the other half is a person.

The W3C says it directly: "Tools cannot check all accessibility aspects automatically. Human judgement is required" (W3C WAI). axe-core itself returns some results as "incomplete", meaning it "could not be certain, and manual review is needed" (axe-core). Whether your alt text describes the image, whether focus lands somewhere sensible when a dialog closes, whether an error message makes sense when a screen reader reads it out: no engine decides those, home-built or bought.

So budget for human review in both scenarios. The difference is what the automated half hands that reviewer. A good setup tells the reviewer which pages were scanned, what passed, and which items still need a person, so they spend their time on judgement instead of re-checking contrast on page 40. We wrote about how to split that work in accessibility monitoring vs audit.

How do you decide without a six-week evaluation?

Run the cheap experiment first, then price the gap.

  1. Add axe-core to your existing tests on your three most important routes. One day of work, at most.
  2. Write down which pages that check can't reach, and who asked for something it can't produce.
  3. Estimate what closing that gap would cost to build and maintain for a year, in engineering days.
  4. Compare that against a platform, trialled on the same pages. Our guide to running a scanner proof of concept covers how to do the trial on your real authenticated flows.

My own opinion: the in-house CI check is worth having no matter what you decide, and the platform is worth paying for only when someone other than a developer needs the answer. If every reader of the results writes code, build. If a compliance lead, a client or a regulator is in the room, the reporting and history are what you're buying, and they take a lot longer to build than they look.

We built wcagc around the same engine and the same limits. It scans crawled pages and signed-in journeys, keeps the history, and marks the clauses it can't decide as open instead of quietly passing them. It won't tell you your site conforms. Nothing automated honestly can.

Frequently asked questions

Is it cheaper to build accessibility testing in-house?

It's cheaper to start, because the axe-core engine is free under the Mozilla Public License 2.0. The cost shows up later as engineering time spent on crawling, authentication, baselines, history and reporting. If you only need pull-request checks on your own code, building usually stays cheaper.

Do paid accessibility scanners find more issues than axe-core?

Usually not by much, because many of them use axe-core as their rules engine. Automated testing covers roughly 30 to 57% of issues depending on how you count, and a platform mostly adds reach, history and reporting rather than a different detection rate.

Can we use both an in-house check and a bought platform?

Yes, and many teams should. The in-house check blocks regressions at the pull request; the platform covers the whole site, content from outside your codebase, and the reporting other people need.

Does either option make our site WCAG or EAA compliant?

No. Automated testing finds issues but cannot confirm conformance; the W3C states that human judgement is required. Both options need manual review with a keyboard and a screen reader, and an accessibility statement documents your efforts and known gaps, not a guarantee.

See what the gap looks like on your site

Before you price anything, look at what an automated check finds on a page that matters. Run a free scan on a real journey page, then go through the items it leaves open with the WCAG checklist. If the results are something only your developers need, build. If other people need them too, you'll have a clearer idea of what you'd be buying.


Pavel Charkasau, founder, wcagc.com. Last updated 11 October 2026.

Sources

  • axe-core README, Deque Labs — distributed under the Mozilla Public License, version 2.0; rules for WCAG 2.0, 2.1 and 2.2 at levels A, AA and AAA; "It returns zero false positives (bugs notwithstanding)"; "you can find on average 57% of WCAG issues automatically"; "incomplete" results where "axe-core could not be certain, and manual review is needed." Accessed 11 October 2026.
  • axe-core package metadata, npm registry — stable releases 4.11.0 (9 October 2025), 4.11.1, 4.11.2, 4.11.3, 4.11.4, 4.12.0, 4.12.1, 4.13.0 and 4.14.0 (5 October 2026). Accessed 11 October 2026.
  • Automated testing study identifies 57% of digital accessibility issues, Deque — 10 March 2021; over 2,000 audits, over 13,000 pages, nearly 300,000 issues; 57% coverage measured by issue volume, against the widely held 20–30% estimate based on success criteria. Accessed 11 October 2026.
  • Selecting Web Accessibility Evaluation Tools, W3C WAI — "Tools cannot check all accessibility aspects automatically. Human judgement is required"; "Some tools check a single page, while others check entire groups of related pages"; "Some can also access password-restricted content"; integration into browsers, CMSs and development and deployment tools. Accessed 11 October 2026.
  • Accessibility Conformance Testing (ACT) Overview, W3C WAI — ACT Rules make testing "more transparent, and thus reduces confusion caused by different interpretations of accessibility guidelines." Accessed 11 October 2026.
  • Directive (EU) 2019/882 (European Accessibility Act), EUR-Lex — Article 13(3): service providers "shall ensure that procedures are in place so that the provision of services remains in conformity with the applicable accessibility requirements", taking account of changes to the service, the requirements and the harmonised standards. Accessed 11 October 2026.