Skip to content

Study 001 independent reviewer call

Independent review pending

Study 001 · Open call

Help us evaluate nine AI-built interfaces

Human Standards is recruiting two independent experts to complete the blinded review stage of our first controlled MCP study.

At a glance

Reviewers
2 independent experts
Material
9 opaque interface artifacts
Time
Approximately 3–4 hours
Format
Remote and self-paced

01

What we are studying

We gave one AI coding-agent configuration the same specification for a fictional community clinic appointment service under three controlled guidance conditions. The generated products support booking, viewing, rescheduling, and cancelling an appointment against deterministic fictional data.

The research question is whether a directed Human Standards MCP retrieval workflow changes functional completeness, accessibility, robustness, or human-centred quality. Publication is not conditional on a favourable result.

02

What the review involves

Each selected reviewer independently assesses all nine interfaces in a private review environment. You will receive the same product specification, scripted tasks, fictional test data, and frozen scoring rubric for every artifact.

  • Complete booking, validation, recovery, rescheduling, and cancellation tasks.
  • Inspect keyboard operation and the main booking path at mobile width.
  • Score 10 usability and human-centred quality dimensions from observed behaviour.
  • Record blocked tasks and concise observations without inspecting source code.

The work can be split across sessions. Selected reviewers will receive setup and access instructions privately before agreeing to participate.

03

Why independence matters

Artifacts are identified only as V001–V009. Reviewers will not receive condition labels, run identifiers, prompts, transcripts, source comments, MCP records, automated results, or the other reviewer’s scores.

Please do not apply if you have already seen Study 001 outputs, condition mappings, or preliminary results. Prior exposure would make a blinded independent review impossible. General familiarity with Human Standards or MCP is not itself disqualifying.

04

What becomes public

Only after both reviewers have locked all 18 evaluations will we reveal the condition mapping and analyse the human-centred quality scores. The final case study will include all runs, the protocol, prompts, checksums, outcome data, disagreements, limitations, and reproduction instructions.

Mixed, null, and negative findings will be publishable. Reviewers may choose whether to be acknowledged by name; scores will otherwise use an opaque reviewer identifier.

Questions about the method belong in public. Once the study discussion is open, use it for questions about eligibility, procedure, blinding, or publication. Do not post email addresses, availability, or other personal information on GitHub.