How we know
Every number we publish comes from a script over exported data, and the plan for that script was written before anyone took the test.
The plan
Analysis plan tag: prereg-v1 (not yet tagged). The plan names the outcome, the test, the exclusions and the sample size before the first participant.
What a participant does
- Reads the consent text and ticks two boxes.
- Picks the healthier of two creeks (a warm-up, not scored).
- Is assigned by the server to see the lesson first or the test first.
- Answers Yes, No or Can't tell on sixteen photos, four per feature.
- Sees their score per feature, and may keep it with a random token.
How numbers reach the README
Scripts in evals/ write files in results/. A check in CI compares every number in the README with those files. Nobody types a number by hand.
AI on the same test
Vision models take the same sixteen photos. A model may only raise a question about a feature it passed, and the volunteer always answers first.