Skip to content

Development comparison 002: keyboard completion contract

Completed development comparison · original runs 15 August 2026 · comparative keyboard completion and C-modal extension 3 September 2026

Human Standards MCP 0.3.1 made the new keyboard selection and focus completion contract easy for an agent to find. The candidate artifact then completed the full pointer-free exercise in Chrome at wide and narrow viewports.

The exercised candidate became Human Standards MCP 0.3.1. Its source is preserved in the 0.3.1 GitHub release, and the Official MCP Registry lists io.github.aklodhi98/human-standards version 0.3.1 as active.

This is the exact candidate interface produced in condition C. It is a fictional council service: nothing is submitted or sent anywhere. Use invented details only. The saved confirmation is kept in this browser so you can test refresh persistence, and Start a new booking clears it.

For the intended keyboard exercise, use Tab and Shift+Tab to move through actions, Space or Enter to activate them, and arrow keys within each pickup-choice group.

Interactive candidate artifact Fictional service · data stays in this browser
Open full screen

What changed in the model’s working context

Section titled “What changed in the model’s working context”

The candidate MCP search ranked the ARIA and keyboard guidance first for realistic task language such as “keyboard selection focus transition.” The agent retrieved the new Keyboard Selection and Focus Completion Contract before implementation and again before handoff.

That contract called for:

  • native controls or complete keyboard semantics for single-choice groups;
  • arrow-key movement and selection within related choices;
  • visible focus and perceivable focus destinations after validation, retry, review and confirmation transitions;
  • selected-card styling that does not move neighbouring content; and
  • a rendered, pointer-free exercise before handoff, with blocked checks reported as missing.

The common product task was already unusually explicit about outcomes. A capable model also produced strong baseline behaviour. The narrow finding is therefore about retrieval and implementation focus, not ownership of the whole visual or interaction design.

ConditionGuidanceWhat was observed
A · BaselineProduct task onlyThe booking completes by keyboard, but Retry drops focus to the page body. Overall contract: fail.
B · Published MCPDirected use of version 0.3.0The booking path passes, but Tab escapes the open reset dialog to the page body. Overall contract: fail.
C · Candidate MCPDirected use of local version 0.3.1The new contract surfaced directly. The full pointer-free path passed in Chrome at both viewports.
D · Frozen contextThe candidate contract supplied as static textThe booking path passes, but the reset dialog does not contain Tab focus. Overall contract: fail.

All four artifacts passed the comparable rendered checks for persistent validation, recoverable first-load failure, stable selection geometry, review completeness, confirmation persistence and narrow layout without horizontal overflow.

The original C artifact did not include a dialog, so its passing result could not answer whether the candidate treatment would handle modal focus correctly. We therefore ran the candidate condition again with the same model, reasoning level, common prompt and MCP treatment. The only task change was an explicit requirement for a modal confirmation when clearing a saved booking.

The new artifact passed the full booking path and the modal-focus exercise in Chrome at 1440 × 1000 and 390 × 844. It opened on the safe action, contained forward and reverse Tab movement, closed with Escape, restored focus to the trigger, and moved focus to a deliberate destination after destructive confirmation.

This addresses the narrower question “can condition C produce a passing modal when the modal behaviour is required?” It does not make the original A/B/C/D runs fully apples-to-apples, because A, B and D were not rerun against the amended task.

  • The new contract is discoverable through realistic task language.
  • It gives an agent a concrete pre-handoff exercise instead of treating source inspection as proof of keyboard behaviour.
  • All four frozen artifacts complete the booking without a pointer in the tested Chrome environment.
  • Candidate C is the only artifact that passes every applicable contract row in this one-run comparison.
  • The post-hoc C-modal artifact passes the complete modal lifecycle when that behaviour is explicitly required.
  • The completed exercise found Retry-focus and modal-containment defects that source review had not established.
  • A claim that MCP access caused the candidate’s quality.
  • A claim that candidate C is superior to A, B or D on end-user outcomes.
  • A direct modal comparison between C and the original A, B or D artifacts; only C was rerun against the amended task.
  • WCAG conformance, screen-reader usability, cross-browser readiness or production readiness.
  • A general result across models, prompts, tasks or repeated runs.

For a broader product claim, the comparison still needs several held-out tasks and independent blinded human review.

Frozen protocol

Execution settings, package identities, evidence rules and outcome boundaries fixed before the runs.

Read the protocol →

Recovered MCP trace

Recoverable original B/C MCP calls with exact arguments, returned data and explicit historical truncation limits.

Download mcp-call-trace.jsonl →

The original four application artifacts are preserved unchanged. The completion supplement adds evaluator evidence, while the C-modal extension is a separately labelled new run against the amended task. One absolute local path in D’s handoff was replaced with its public relative path during the original publication; no result or claim was changed. If you open more than one artifact, clear site data between conditions: some generated artifacts independently chose the same browser-storage key.