Skip to content

Studies

Human Standards runs small, controlled studies to examine whether access to human-factors guidance changes what AI coding agents build. The purpose is to produce inspectable evidence—not promotional demonstrations.

Study 001: Community clinic appointment flow

Status: Exploratory pilot concluded

Nine interfaces were generated from the same specification under three controlled guidance conditions. Directed MCP access improved the automated results, but did not reliably remove obvious usability and accessibility problems.

Read the case study →

Development comparison 002: Keyboard completion

Status: One-run development comparison complete

Four held-out artifacts tested whether MCP 0.3.1 made a new keyboard and focus completion contract easier for an agent to find and exercise. All four complete the booking by keyboard; A fails Retry focus, B and D fail modal containment, and C passes every applicable contract row. A separate C-modal run also passes, but only C was rerun, so it is supporting evidence rather than a new four-way comparison.

Try the candidate and inspect the evidence →

The planned independent-review stage was not completed. The protocol deviation, missing human outcome, coordinator observations, condition mapping and automated evidence are published with the result.

Every published Human Standards study or pilot will:

  • define its research question and outcomes before analysis;
  • preserve all recorded runs rather than selecting favourable examples;
  • distinguish MCP availability from actual MCP use;
  • separate automated findings from human judgement;
  • report mixed, null, and negative findings;
  • include limitations and enough raw evidence to audit the claims; and
  • disclose any protocol deviation or missing outcome rather than silently substituting a weaker method.

The study question, method, condition mapping and final evidence are public. Applicant contact details and private recruitment correspondence are not published. Study 001 coordinator observations remain clearly separated from the independent human-centred quality score that was planned but never collected.

Questions about the method can be raised in GitHub Discussions.