Research contract v1

Expert AI benchmark study protocol

This protocol packages a blinded reviewer study without treating automation as expert evidence. The public research submission contains only aggregated pseudonymous reviews; raw responses, identity keys, consent records and transport metadata stay private.

Preregister and freeze

  1. Create a draft submission with stable AI- position IDs, decided pilot/tuning/holdout splits and empty expertReviews arrays. Choose a neutral studyId: it remains visible to reviewers and must not reveal a split, model choice or expected result.
  2. Store the dated protocol and corpusHash in a third-party timestamped system or signed study record before review. Repository tooling cannot prove when a hash was created.
  3. Keep the draft outside this repository and create a new private kit outside this repository:
node scripts/prepare-mahjong-consultant-kit.mjs PRIVATE_DRAFT.json NEW_PRIVATE_KIT_DIRECTORY

The generator refuses private inputs or output inside the publishable repository and refuses to overwrite an earlier freeze. It creates:

Send only a separate copy of SEND-TO-REVIEWER. Never send or publish the draft, manifest, predictions, operator records or raw responses. The packet uses opaque ITEM- IDs and omits source IDs, split assignments, evidence IDs, existing reviews, measured latency and model predictions.

Collect privately

Give each reviewer separate packet and response-template copies. Keep the real-person-to-REV- key, consent, recruitment record, messages and raw returns outside the public repository. Reviewers should work independently from only the acting hand, public state and legal actions.

node scripts/prepare-mahjong-ai-review.mjs check-response PRIVATE_MANIFEST.json REVIEWER_PACKET.json RESPONSE.json

This rejects missing or duplicate items, illegal actions, contradictory labels and unsupported fields, but cannot prove who completed a response or what they saw.

Assemble the public aggregate

Merge only pseudonymous reviews and explicitly choose draft, pilot or final:

node scripts/prepare-mahjong-ai-review.mjs assemble DRAFT.json PRIVATE_MANIFEST.json REVIEWER_PACKET.json PUBLIC_SUBMISSION.json pilot RESPONSE-1.json RESPONSE-2.json RESPONSE-3.json
node scripts/validate-mahjong-research.mjs PUBLIC_SUBMISSION.json

Score predictions only after frozen labels are assembled:

node scripts/score-mahjong-ai-benchmark.mjs PUBLIC_SUBMISSION.json PREDICTIONS.json

What automation cannot prove

Hashes do not prove preregistration timing or the absence of another corpus. Opaque IDs do not prove blinding. Anonymous IDs do not prove reviewers are distinct, independent, human, qualified or consenting. File checks do not prove the model avoided holdout labels, reviewers used only visible information, rationales are truthful or recruitment was unbiased. These claims require a documented human process and retained private evidence.

Back to Singapore Mahjong methodology