Skip to main content
Matías Martínez Boylston

AI Direction

A Full WCAG 2.1 AA Audit, Run on a Multi-Agent Pipeline

The promise at stake: an aging, high-value customer who can't finish a task blocked by an accessibility defect.

Role
Head of UX — program owner and method designer
Context
Financial services; a design system and the product surfaces built on it
Timeframe
Phase 1 (design source) closed May 2026; Phase 2 (shipping code) completed June 2026 — accessibility is a 7-year throughline

Outcome recap

  • 270 findings, 19 components audited across design source and shipping code
  • Verification caught real gaps: of 18 re-checked audits, 15 were materially corrected
  • A live, self-accessible report — zero violations on its own automated scan

The challenge

Somewhere in the customer base is an older, higher-value customer who opens the app to do something ordinary — check a balance, make a payment — and can't. Not because the product is broken. Because a button doesn't announce itself to a screen reader, or a focus ring vanishes, or a modal traps the keyboard. That customer doesn't file a bug. They just fail, quietly, and the business never sees it happen.

A design system is the highest-leverage place in a product organization to fix — or silently multiply — defects like that one. Every barrier baked into a shared button or modal propagates into every screen that uses it. This system had never been audited for accessibility. The default culture treated it as a final-stage checkbox, not a discipline owned up front.

The case for fixing it isn't abstract, and in this market it isn't primarily a legal one — accessibility law here binds only the public sector. So the case has to be commercial: the customer base skews older, and the older, higher-value customer is exactly the person whom low-vision, motor, and cognitive barriers hit hardest. A control you can't operate is a sale you don't make. At scale, it's brand and reputational risk. Underneath all of it, the same fact keeps returning: behind every finding is a real person who can't finish a task.

What I did

I took charge of a problem nobody owned. I stood up a full WCAG 2.1 AA program with a deliberate phase gate: Phase 1 audits the design source (Figma) to establish ground truth free of implementation noise; Phase 2 audits the real shipping code; then the two are reconciled. No mixed-surface shortcuts. Scope: all 19 components plus the underlying token foundations, audited on both surfaces.

The differentiator was the method. I designed the audit as a multi-agent pipeline: each component ran through an automated pass, which catches only a fraction of WCAG issues, plus a manual technique pass for everything automation misses — keyboard operability, focus order, screen-reader labeling, and everything else automation can't see. Then I directed an independent, adversarial verifier to re-check every report, prompted specifically to find what the first pass missed.

Findings become evidence teams can act on through adversarial verification, root-cause synthesis, and attention to human impact.
Fig. 01 — Findings → Trust. Make barriers visible, challenge the first pass, find shared causes, and see the person behind the issue: a path toward evidence the organization can trust. View the full-size infographic.

That verification is where the method was tested, not just demonstrated. Of 18 re-checked audits, 15 were materially corrected — not rubber-stamped. That's a real number to sit with: on a first look, most of what we thought we knew about our own compliance was incomplete. I'd built the pipeline on the premise that a single pass, even an automated one, isn't enough to trust — and the verification proved the premise right in a way I didn't fully expect until I saw the correction rate. It meant standing by a process that was, in effect, grading itself down in front of me, and using that as the reason to trust the final report more, not less.

"I'd built the pipeline on the premise that a single pass, even an automated one, isn't enough to trust."

I synthesized rather than just listing. A reconciliation matrix compared design against code and grouped recurring barriers by shared causes and ownership, turning individual findings into decisions teams could act on. I shipped a deliverable a committee can actually use: executive summary, a human-impact layer, a design route and an engineering route, per-component detail grouped by who owns the fix, a full findings registry, and methodology.

WCAG 2.1 AA

Audit report

v1.0 · Baseline

Executive summary · Two-surface baseline

Accessibility across design and shipping code.

A versioned baseline designed to make progress measurable—not merely compliant.

Total findings
270
Critical
14
Serious
131
Components
19
AA conformant
0 / 19

Severity profile

Where the risk concentrates

Both surfaces

Component pressure

Highest finding counts

Top five

Public-safe aggregate view Design source ↔ Shipping code
Live WCAG Report. A sanitized view of the live report, rebuilt from aggregate data to preserve confidentiality while showing the decision-ready artifact.

I made the report practice what it audits — self-accessible by design, passing its own automated scan with zero violations across every page, fully keyboard-navigable with visible focus and AA contrast. It's a live artifact I can demo on the spot, and it's built to be re-run: each pass produces a versioned report with computed deltas, so accessibility becomes something the organization tracks like a health metric, not something it commissions once and forgets.

When a real contrast defect reached production, I used it to write a durable team standard — "in UX, we are accessibility's line of defense" — made objective and measurable rather than debated. And I made the findings human: composite personas, age and ability archetypes, joined to the actual findings that would block them, translating 270 technical line-items into who they affect and why it matters.

Outcome & impact

A previously-unaudited design system now has a complete, two-surface WCAG 2.1 AA baseline: 270 findings, each attributed to a WCAG success criterion and to an owner. Underneath the number is the insight that actually moves the needle — a small set of root fixes resolves a disproportionate share of the severe findings, which turns a 270-line backlog into a short, fundable list.

Design source and shipping code are audited independently, then their evidence is reconciled to identify shared causes and clear ownership.
Fig. 03 — Design ↔ Code. Compare what is specified with what is implemented. Reconcile the evidence from two independent audits into shared causes, clear ownership, and one actionable baseline. View the full-size reconciliation.

More durable than the count is what the method produces on every re-run: a repeatable way to know, at any point, exactly where the system stands on accessibility — evidence a leadership team can act on instead of a one-time report that ages out of relevance. And there's a culture shift underneath the artifact: accessibility reframed as the team's own line of defense, backed by a standard anyone can apply without being an expert.

Team accessibility standard: in UX, we are accessibility’s line of defense. Review contrast, keyboard access and visible focus, names, roles and states, and evidence with a named owner before handoff.
Fig. 04 — Team Accessibility Standard. A public-safe portfolio reconstruction of the operating principle: verify text and control contrast; check keyboard access and visible focus; specify names, roles, and states; record the check and name the owner. View the full-size standard.

The report itself is the closing argument: a demoable, self-accessible artifact that earns executive attention because the medium reinforces the message. And it's genuine conviction, not performance — a 7-year throughline that goes back to launching a previous company's first accessible products.

What this taught me

The correction rate — 15 of 18 — was the real lesson, not a footnote. It taught me not to trust a single pass of anything, including a first read of my own team's work, and to build the check for that distrust into the process itself rather than into a review meeting after the fact. The program's proof is the report; the credit belongs to the design and engineering owners who took 270 findings and worked the fixes without getting defensive about them. My job was to build the method rigorous enough that their work would hold up in front of a committee.

Skills demonstrated

  • Accessibility (WCAG 2.1 AA)
  • Design-system governance
  • Multi-agent AI orchestration
  • QA method design (adversarial verification)
  • Audit synthesis and root-cause analysis
  • Human-centered framing (personas and prevalence data)
  • Commercial framing of UX
  • Building team standards and culture
  • Technical reporting

Proof / artifacts

Automated and manual audit passes converge on an adversarial verifier; 15 of 18 reports were materially corrected.
Fig. 02 — The Multi-Agent Audit Method. An automated pass covers part of the audit; a manual technique pass examines what automation misses. Both feed an independent, adversarial verifier that re-checks every report; 15 of 18 were materially corrected. View the full-size audit method.