Frontend engineering | October 11, 2026

Agent-readable UI contracts need tests, not longer prompts

A coding agent can retrieve a beautiful component guide and still ship a button that submits the wrong form, loses focus during loading, or fails keyboard users. The durable fix is a versioned rule graph: stable IDs, structured contracts, executable checks, runtime evidence, and a human release gate.

Design systemsCoding agentsAccessibility evidenceContract testingSources checked Oct 11
A rule graph connecting UI component guidance to automated and human tests

The missing interface is between guidance and proof

Open Components appeared in late September as an experiment in publishing UI guidance for three audiences at once: people using a component, developers implementing it, and agents generating or reviewing it. Its site offers human-readable pages, raw Markdown, compact YAML contracts, a JSON Schema, an llms.txt index, an MCP endpoint, and an agent plugin. That combination makes the project interesting even though it is young and its current component coverage is limited.

The launch is not evidence of broad adoption. At our October 11 check, the public GitHub repository had nine stars and no open issues; its Show HN discussion had seven points and one comment. Those are discovery signals, not a production benchmark. The useful question is architectural: what changes when guidance is exposed as an addressable contract rather than prose copied into a system prompt?

Current coding agents are competent at imitating visible patterns. They can also miss invisible semantics. A generated button may look right while relying on the browser's implicit submit behavior. A loading state may use disabled, remove the element from interaction, and move focus unexpectedly. An icon-only control may render correctly but have no accessible name. A component may satisfy an automated scanner and still confuse a screen-reader user because its announcement sequence is wrong.

A larger prompt does not solve the lifecycle problem. Prompt text can become stale, be truncated, lose source identity, or conflict with local conventions. Reviewers cannot easily tell which rule the agent attempted to follow. CI cannot reliably ask whether a paragraph was obeyed. The missing interface is a stable mapping from guidance to implementation, checks, exceptions, and retained evidence.

The unit of agent reliability is not the prompt. It is the rule that survives retrieval, generation, testing, review, and change.

Separate the guide, contract, implementation, and evidence

Teams often ask one artifact to do four jobs. A design-system page explains intent. A contract identifies rules and data. A reference implementation shows one realization. A test suite observes behavior. Keeping those layers separate prevents a framework example from silently becoming the standard.

ArtifactPrimary jobWhat it cannot prove
Human guideExplain rationale, variants, content, interaction, and trade-offs.That generated code implements every rule.
Machine contractExpose stable IDs, severity, applicability, checks, and references.That a browser or assistive technology behaves correctly.
Reference componentDemonstrate one tested mapping in one stack.That ports to React, Vue, native web components, or another token system are equivalent.
Static linterCatch syntax and structural violations quickly.Focus order, announcements, timing, or contextual form behavior.
Runtime test suiteExercise keyboard, DOM, focus, form, state, and visual behavior.Usability for every person and assistive-technology combination.
Evidence receiptBind a build to rules, tests, results, exceptions, and reviewers.Future compatibility after dependencies or rules change.

Authority matters. Open Components' repository states that its site code is MIT-licensed while its guidance and plugin content use CC BY 4.0. That makes reuse possible with the relevant attribution, but licensing does not make the project a standards body. W3C specifications and WAI-ARIA Authoring Practices remain primary references for web accessibility semantics. Product teams must also apply their own browser support, brand, localization, privacy, and risk policies.

Build a rule graph from retrieval to release

A robust flow starts with a pinned contract revision. An agent retrieves only the rules applicable to the component, target framework, interaction state, and product context. Generation records which rule IDs were used. Static and runtime tests emit results against the same IDs. Exceptions need an owner, rationale, expiry, and compensating evidence. The final reviewer sees coverage rather than a vague statement that “accessibility was considered.”

normative sources + product policy + design-system guidance
  -> normalized rules with stable IDs and provenance
  -> versioned component contract + JSON Schema validation
  -> agent retrieval: component + task + framework + context
  -> generated patch with rule manifest
  -> static checks + browser tests + accessibility API checks
  -> evidence receipt: pass | fail | exception | not-applicable
  -> human review, release decision, and retained artifacts

The contract should be compact enough for retrieval but rich enough to drive evaluation. Do not embed entire standards documents. Link to the authoritative source and normalize the decision your system must make. Preserve provenance because source language, local policy, and implementation advice are not interchangeable.

schema: ui-contract/v1
component: button
contractVersion: 2026-10-11.1
sources:
  - id: wai-apg-button
    url: https://www.w3.org/WAI/ARIA/apg/patterns/button/
rules:
  - id: button/declare-type
    level: error
    appliesWhen: element == "button" and insideForm == true
    requirement: Declare type="button" unless the action intentionally submits.
    checks: [ast/button-explicit-type, browser/form-action]
  - id: button/keep-focus-during-loading
    level: error
    appliesWhen: asyncAction == true
    requirement: Keep focus stable while preventing duplicate activation.
    preferredState: aria-disabled
    checks: [browser/focus-stability, browser/repeat-activation]
  - id: button/accessible-name
    level: error
    requirement: Expose a non-empty accessible name.
    checks: [dom/accessible-name, human/label-clarity]

Stable IDs are the critical design choice. Rewrite the explanation if it becomes clearer, but do not casually rename button/accessible-name. Test history, exceptions, dashboards, and generated manifests depend on that identity. When meaning changes, publish a migration note and compatibility class: editorial, additive, behavior-changing, or removed.

A button exposes why semantics must be executable

Buttons look simple because browsers provide much of their behavior. That also creates hidden defaults. MDN documents that a <button> associated with a form submits it when no explicit type says otherwise. An agent creating a “Show details” control inside a settings form can accidentally trigger validation or save data. A static check can find the missing type; a browser test can prove which submit handler ran.

Loading introduces a second trap. Native disabled prevents activation, but it can affect focus and form behavior. The Open Components button guidance recommends retaining focus for an asynchronous action and using aria-disabled="true" when the control should remain discoverable. MDN stresses that aria-disabled only communicates state: application code must actually suppress activation, and styling must make the state understandable.

async function handleSave(event) {
  const button = event.currentTarget;
  if (button.getAttribute("aria-disabled") === "true") return;

  button.setAttribute("aria-disabled", "true");
  button.setAttribute("aria-busy", "true");
  try {
    await saveDraft();
    announce("Draft saved");
  } finally {
    button.removeAttribute("aria-busy");
    button.removeAttribute("aria-disabled");
    button.focus();
  }
}

This is a pattern, not a universal prescription. A destructive action may need cancellation or a confirmation step. A long-running operation may move focus to a results region. A disabled control inside a composite widget can have different keyboard expectations. The contract should state applicability and allow a reasoned alternative, not force a snippet everywhere.

WAI-ARIA Authoring Practices says a button is activated with Space or Enter and describes where focus should go after activation. That behavior needs a browser test. An icon button also needs an accessible name that communicates the action. An automated name computation can prove that a name exists; only a human reviewer can decide whether “More” is sufficiently clear in context.

Use a test ladder instead of one accessibility score

Rule coverage should move from cheap deterministic checks to expensive contextual checks. A single aggregate score hides whether a critical keyboard failure was averaged with dozens of minor successes.

LayerExample checkEvidence
SchemaContract validates; IDs are unique; references and severities exist.Schema result and contract hash.
AST/staticButtons in forms declare type; icon controls expose a naming mechanism.Rule-level linter output with file and line.
DOMComputed role, name, state, relationships, and invalid attributes.DOM snapshot and accessibility-tree excerpt.
Browser interactionSpace/Enter activate once; loading suppresses repeats; focus lands predictably.Playwright trace, events, and focus sequence.
Visual statesFocus indicator, disabled state, zoom, contrast, forced colors, and text expansion.Named screenshots at supported viewports.
Assistive technologyName, role, state, and update announcement make sense.Manual protocol, environment, observations, reviewer.
Product contextThe action label, consequence, latency, errors, and recovery fit the task.Design/accessibility review and acceptance record.

Automate deterministic checks aggressively, but publish their boundary. Axe or another rules engine can catch many known violations; it cannot certify the experience. A browser trace proves the tested browser and scenario, not every supported device. Manual assistive-technology testing must identify the browser, OS, screen reader, version, input method, and scenario so that a later run can reproduce the observation.

Tests should fail on uncovered critical rules. If a new contract adds button/keep-focus-during-loading and no check maps to it, CI should label the rule “unverified,” not quietly green. That distinction prevents contract coverage from becoming documentation theater.

Attach a rule-level receipt to every generated patch

The evidence receipt is the bridge between agent output and release governance. It should be machine-readable, small enough to retain with CI artifacts, and readable in code review. Record the exact contract and source revisions, generated files, agent or tool identity, checks, exceptions, and human decision. Avoid storing sensitive prompts or user data when a hash and task ID are sufficient.

{
  "component": "button",
  "contract": "ui-contract/v1@2026-10-11.1",
  "contractSha256": "...",
  "generatedFiles": ["src/components/SaveButton.tsx"],
  "rules": {
    "button/declare-type": {"status": "pass", "test": "ast:1842"},
    "button/accessible-name": {"status": "pass", "test": "dom:1847"},
    "button/keep-focus-during-loading": {"status": "pass", "trace": "pw:1851"},
    "button/label-clarity": {"status": "human-pass", "reviewer": "design-system-owner"}
  },
  "exceptions": [],
  "releaseGate": "approved"
}

The receipt makes upgrades tractable. When a rule changes, query which components used the old version and which exceptions remain open. When a regression occurs, compare the contract hash, generated patch, dependencies, browser matrix, and traces instead of trying to reconstruct what an agent may have read.

Design for the ways agent-readable guidance can fail

FailureWhy it happensControl
Stale retrievalThe agent caches an old Markdown or YAML page.Pin version/hash; reject receipts for retired critical contracts.
Rule without authorityCommunity advice is treated as a normative requirement.Store provenance and authority class; link to primary sources.
Framework driftA React adapter changes semantics while preserving visual output.Run implementation-independent DOM and browser tests.
Prompt compliance claimThe agent says it followed guidance but emits no rule mapping.Require a manifest and test receipt; ignore self-attestation.
False automated confidenceA scanner passes while focus, timing, or announcements fail.Keep runtime and manual gates for critical interactions.
Exception permanenceA deadline waiver becomes invisible technical debt.Require owner, reason, compensating control, expiry, and re-test.
Contract injectionUntrusted content reaches agent instructions through retrieval.Allowlist sources, validate schema, separate data from executable instructions.
Telemetry leakPrompts, screenshots, or traces contain customer data.Redact, minimize, classify, and retain only approved evidence.

Security belongs in the contract pipeline too. An MCP endpoint or plugin expands the input surface. Treat retrieved text as untrusted content, validate structured responses, pin approved hosts and versions, and do not allow retrieved instructions to override repository policy. A contract can describe component rules; it should not acquire arbitrary shell or deployment authority.

Adopt one critical component before indexing the whole design system

Start with a component that has bounded semantics and high reuse: button, link, input, dialog, or combobox. Buttons are a good pilot because form behavior, keyboard activation, focus, accessible naming, loading, destructive actions, and visual states create useful test coverage.

  1. Freeze authority. List standards, product policies, supported platforms, and design-system guidance. Resolve contradictions with named owners.
  2. Normalize ten to twenty rules. Assign stable IDs, applicability, severity, rationale, sources, and candidate checks. Do not ingest every paragraph.
  3. Validate the contract. Use JSON Schema, unique-ID checks, broken-reference checks, and change classification.
  4. Map one implementation. Annotate the reference component with rule IDs without claiming the code is the rule.
  5. Build the test ladder. Start with explicit type, accessible name, keyboard activation, focus sequence, duplicate activation, and key visual states.
  6. Run paired tasks. Give an agent the old prose on one branch and the contract flow on another. Compare defects, review time, uncovered rules, and false positives.
  7. Retain receipts. Require the same artifact for human- and agent-authored changes so the process measures components rather than authors.
  8. Expand only after maintenance works. Publish a rule change, migrate a component, expire an exception, and reproduce a trace before adding more component families.

Measure useful outcomes: critical-rule coverage, defects escaping review, median time to resolve failures, stale-contract incidents, exception age, and cross-framework parity. Token count and number of generated components are not quality metrics.

Release checklist

  • The contract version, hash, sources, and licensing are recorded.
  • Rule IDs are stable, unique, scoped, and mapped to severity and applicability.
  • Every critical rule has an executable check or an explicit human gate.
  • The generated patch includes a rule manifest; self-reported compliance is not accepted as evidence.
  • Keyboard activation, focus sequence, form behavior, loading, errors, and recovery were exercised in supported browsers.
  • Automated accessibility output is supplemented by contextual and assistive-technology review for critical paths.
  • Exceptions have an owner, rationale, compensating control, expiry, and re-test date.
  • Retrieved guidance is allowlisted, schema-validated, and unable to override repository or deployment policy.
  • The receipt contains no unnecessary prompts, secrets, customer data, or sensitive screenshots.
  • A named design-system or accessibility owner approves the release.

FAQ

What is an agent-readable UI component contract?

It is a versioned set of component rules with stable identifiers, structured applicability, verifiable checks, and expected evidence that both an agent and CI can consume. It complements human guidance; it does not replace design judgment.

Is Open Components a web standard or component library?

No. It is an early open project that publishes guidance, machine-readable contracts, a reference implementation, and agent-facing interfaces. W3C specifications, WAI practices, browser documentation, and your product policy remain the relevant authorities.

Can a YAML contract prove a component is accessible?

No. It can improve traceability and automate deterministic checks. Keyboard behavior, focus, announcement quality, browser differences, and real-task usability still need runtime and human evaluation.

Should only AI-generated code need receipts?

No. Apply the same rule-level evidence to human- and agent-authored changes. That keeps the system focused on release risk and makes comparisons meaningful.

How should rule changes be versioned?

Keep IDs stable when meaning is stable. Record source and contract revisions, classify changes, publish migrations for behavior-changing updates, and invalidate only the evidence that actually depends on the changed rule.

Sources and verification notes

Current project details and linked guidance were checked October 11, 2026. Repository stars and discussion counts are point-in-time discovery signals, not evidence of adoption. Open Components is used as a case study; accessibility claims are anchored to primary web guidance.

  1. Open Components — project and component guidance.
  2. Open Components GitHub repository — source, reference implementation, plugin, tests, and licensing.
  3. Open Components raw Button guidance.
  4. Open Components raw Button contract.
  5. Open Components contract JSON Schema.
  6. Open Components llms.txt index.
  7. W3C WAI-ARIA Authoring Practices — Button Pattern.
  8. W3C WCAG 2.2 — Understanding Focus Order.
  9. MDN — the HTML button element.
  10. MDN — aria-disabled.