**Pre-specified analysis plan, part 2: does the checker's question help?**

Not yet tagged. It becomes binding when tagged `prereg-v2`, before any part 2 session exists. It adds a second, optional block after the two-minute test. It changes nothing in `docs/analysis_plan.md` (tag `prereg-v1`): part 1's flow, items, outcome and analysis stay as they were, and the one change to part 1's end screen, the line that offers part 2, is logged in `docs/deviations.md`. The code enforces this text: `evals/assist_analysis.py` refuses to run on real data before the lock and refuses to run at all if the tag is missing or this file differs from the tagged version.

This is a usability test of our own tool, run to improve it and to report honestly how well it works. Participation is anonymous and voluntary. We collect no names, emails, IP addresses or precise locations. It is not designed or presented as generalizable human-subjects research.

1. Question. Track 3 asks teams to use AI responsibly to support stream assessment without replacing human judgment. Our AI never decides: its output becomes a flag through the gate, and a flag can only make one question appear. This part measures whether that one question makes a person more accurate. The AI's help is measured, not assumed.
2. Design. Who: everyone who finishes part 1, through the public link or the panel, in either part 1 arm. After the part 1 score screen they are offered part 2 with one line: "Eight more photos, two minutes, and this time a checker may ask you to look again." They can decline. A decline is recorded and analysed as not started. Randomization: 1 to 1 into `assisted` and `unassisted`, by the server, in permuted blocks of 4 from a stored seed, stratified by part 1 arm (one block sequence per part 1 arm). Only a person who starts part 2 is randomized. The arm is stored with the part 2 session. One part 2 session per part 1 session. Sessions on the demo route are never stored.
3. Materials. 8 items, never shown in the lesson, a practice card or the test: two per feature, one present and one absent. Photo role `part2` in `photos/manifest.csv`. Gold labels are the ones chosen at picking, from the written definitions and the evidence each source gives, recorded in the manifest, as in part 1. The item ids, photos and gold labels, frozen here: a01 artificial_bank ph-p2-01 present; a02 artificial_bank ph-p2-02 absent; a03 dug_out_channel ph-p2-08 present; a04 dug_out_channel ph-p2-03 absent; a05 invasive_plant ph-p2-04 present; a06 invasive_plant ph-p2-05 absent; a07 pipe_running ph-p2-06 present; a08 pipe_running ph-p2-07 absent. The same file is `content/part2_items.yaml`. The question wording is part 1's per feature, word for word: artificial_bank "Are the banks artificial, such as concrete or stones set in concrete?"; dug_out_channel "Has this channel been straightened or dug out?"; invasive_plant "Do you see any non-native or invasive plant species?"; pipe_running "Can you see a pipe or drain outlet that empties into this creek?". Answers: Yes, No, Can't tell. Order randomized per participant.
4. The checker's flags, fixed before anyone takes part. No model is called while a person takes part 2. Every flag is computed once and committed in `results/assist_flags.json` by `evals/assist_flags.py`, from the models' stored answers on the 8 items (`results/assist_answers_*.json`, three runs each) and the committed pass table `results/model_pass_table.json`. The checker is one model, `claude-opus-5-5`, the one that passed the most features. An item gets a flag when that model gave the same Yes or No in at least 2 of its 3 runs and the gate `core/gate.py` kept it, which it does only for a feature the model passed. A Can't tell majority, no majority, or a feature not passed gives no flag. So the assistance is identical for every participant. At the tag, `results/assist_flags.json` holds a flag on six items, a01, a02, a03, a04, a07 and a08, and each points the way of that item's gold label; a05 has none because no model passed plants, and a06 none because the checker's answer was Can't tell. So a question can appear only after a first answer that is wrong or Can't tell, and no right first answer is ever questioned. The secondary counts in item 6 are reported with that in mind.
5. The two flows, per item. Unassisted: the person answers and moves on. Assisted: the person answers first. If the item has a flag and it disagrees with the person's answer, one question appears: "The checker noticed something here. Look again?" with Keep and Change. A Can't tell answer disagrees with any flag. Change opens the three answers again and the person chooses. If the flag agrees, or the item has no flag, nothing appears. The flag never shows which way it points and never sets an answer. The person's final answer is the one scored. Stored per item: first answer, final answer, whether a question was shown, the choice made (keep or change), and the timings.
6. Outcomes. Primary: per-participant accuracy on the 8 items, the share answered correctly with the final answer, Can't tell counted as incorrect. Estimate: mean accuracy of the assisted arm minus mean accuracy of the unassisted arm. Secondary, descriptive only, no significance tests: among assisted participants, how often a question was shown, how often the answer changed, and the accuracy of changed answers against kept ones; accuracy by feature; accuracy by part 1 arm; time per item; the share of Can't tell; declines by part 1 arm.
7. Analysis. Test: two-sided permutation test on the difference in means, 10,000 permutations, alpha 0.05. Interval: percentile bootstrap, 10,000 resamples, stratified by arm, 95 percent. Seed 20260926. One confirmatory test. The code is `evals/assist_analysis.py`, written and tested on synthetic data before the tag in three scenarios: the question helps, the question does nothing, and people change to the wrong answer when asked. Sample size: no target is promised. Planning note from our own simulation, `evals/power_v2.py`, results in `results/power_v2.json`, with 8 items and unassisted accuracy near 60 percent: about 77 percent power for a 12 point difference at 40 per arm, and about 45 percent at 20 per arm. The confirmatory claim needs at least 20 finished part 2 sessions per arm. Below that we report the results as a description and say so.
8. Exclusions, decided now. Sessions that did not finish all 8 items. Sessions that finished the 8 in under 20 seconds. QA sessions, and any part 2 session whose part 1 session was a QA session. Sessions after the lock. Repeat part 2 sessions from the same browser token (only the first finished one counts). Declines are counted, not excluded, and reported as not started. We report how many sessions each rule removed, and how many were offered, declined, randomized, started and finished in each arm.
9. Data lock. The same instant as `prereg-v1`: 2026-09-28T01:00:00Z, which is Sunday Sep 27 at 18:00 PDT. Nothing is read before it. Before lock we look at counts per arm only.
10. What we publish whatever happens. All results, including no effect or a negative effect, and the case where people change right answers to wrong ones when asked. The anonymous part 2 response table at data lock. Exclusion counts and the deviations log.
11. Ethics and privacy. As in `docs/analysis_plan.md` item 10. Part 2 stores, besides what part 1 stores: the part 2 arm, the part 1 session it follows, the per-item answers and timings in item 5, and whether the offer was declined. Nothing else.
