**Pre-specified analysis plan: Second Look usability test**

Not yet tagged. It becomes binding when tagged `prereg-v1`, before the first real participant. Amendments from Update 02 section 2 are folded in below. The code enforces this text: `evals/usability_analysis.py` refuses to run on real data before the lock and refuses to run at all if the tag is missing or this file differs from the tagged version.

This is a usability test of our own tool, run to improve it and to report honestly how well it works. Participation is anonymous and voluntary. We collect no names, emails, IP addresses or precise locations. It is not designed or presented as generalizable human-subjects research.

1. Question. Does a two-minute photo lesson help untrained adults recognise four stream features that volunteer assessors commonly miss: artificial banks, a dug-out channel, invasive plants, and pipes or sewage signs? Source for the choice of features: the OneAquaHealth project lead, hackathon workshop 1.
2. Design. Two arms, randomized 1 to 1 in permuted blocks of 4, assigned by the server from a stored seed. Trained: lesson, then test. Untrained: test, then the lesson is offered afterwards. Participants are adults who open the public link. No screening beyond the consent and age checkboxes. Sessions on the /demo route are never stored.
3. Materials. 16 test items. Each pairs one photo with one feature and one question. Answers: Yes, No, Can't tell. 4 items per feature, 2 present and 2 absent. Order randomized per participant. Gold labels are set by a team member, blind to any model output, from written definitions: the first person to label a photo sets its gold label. A second labeller is welcome and optional. If every test photo carries two independent labels by the tag we report Cohen's kappa per feature; if it does not, we say plainly in the results and in the README that one person set the key, and it is listed under Known weaknesses. Where the source itself supports a label we record that evidence in the photo manifest. Photos any labeller calls ambiguous are removed before launch. One labeller: at the tag every gold label was set by one person, Alex Velazquez, from the written definitions and from the evidence each source gives for its own photo, which is recorded in photos/manifest.csv. No second labeller was available before launch, so no Cohen's kappa is reported and the README says so under Known weaknesses. No photo, and no photo from the same spot on the same day, appears in both the lesson and the test. For the pipe feature the label means a pipe or outfall is visible in the photo. It does not mean the pipe is running, because flow and pollution cannot be judged from a photo. The lesson still teaches the three dry days rule, which is for the field, not for a photo. Exact question wording, frozen on 2026-09-21 and approved by Alex Velazquez: artificial_bank "Are the banks artificial, such as concrete or stones set in concrete?"; dug_out_channel "Has this channel been straightened or dug out?"; invasive_plant "Do you see any non-native or invasive plant species?"; pipe_running "Can you see a pipe or drain outlet that empties into this creek?". The invasive plant wording is the official app's, word for word. The built bank wording is derived from the app's own option text "Artificial (concrete or stones with concrete)", because the app offers options there and not a question. The other two are ours. The creek check form keeps the app's own pipe wording, which is a different question asked in a different place.
4. Outcomes. Primary: per-participant accuracy, the share of the 16 items answered correctly, with Can't tell counted as incorrect. Descriptive only, no significance tests: accuracy per feature, hit rate and false-alarm rate per feature, share of Can't tell answers, median test time, median lesson time, accuracy by self-reported prior experience, completed sessions by source, and the warm-up item (share choosing each photo). The warm-up choice is made on the landing page and stored only after consent. An answer can be changed until the person presses Next. The confirmed answer is scored. The first choice, the number of changes and both timings are stored and reported as description only. The warm-up answer is revealed on the end screen, after the test, and never before it: shown on the landing page it would be a small lesson given to both arms, which would shrink the difference this study measures. The reveal says one creek is in a more natural state and makes no claim about health.
5. Exclusions, decided now. Sessions that did not finish all 16 items. Tests completed in under 40 seconds. Repeat visits from the same browser token (only the first completed session counts). Sessions flagged as QA. Sessions that filled the hidden form field that people cannot see. Dry-run sessions before launch (the table is wiped at launch and the wipe is logged). Sessions where the browser reports answers that never reached the server count as incomplete. They are excluded from the primary analysis and counted per arm. They are also left out of the sensitivity check that scores unanswered items as incorrect, because the person did answer and we do not know what they said. We report how many sessions each rule removed, and how many were randomized, started and completed in each arm.
6. Analysis. Estimate: mean accuracy of the trained arm minus mean accuracy of the untrained arm. Interval: percentile bootstrap, 10,000 resamples, stratified by arm, seed 20260920, 95 percent. Test: two-sided permutation test on the difference in means, 10,000 permutations, alpha 0.05. One confirmatory test. Also reported: Hedges g and the full distribution of scores per arm as a plot. Sensitivity checks: partial completers with unanswered items scored incorrect; excluding people who report prior stream assessment. The code is `evals/usability_analysis.py`, written and tested on synthetic data before launch, including a scenario where people say Yes more often without being more accurate, which must show no gain.
7. Sample size and stopping. Target: 80 completed sessions, 40 per arm. Planning note from our own simulation with 16 items and untrained accuracy near 60 percent: about 80 percent power for a 10 point difference at 40 per arm, and for about 12 points at 30 per arm. We make no promise of reaching the target. Data lock is 2026-09-28T01:00:00Z, which is Sunday Sep 27 at 18:00 PDT, whatever the count. Before lock we look at counts per arm only. The confirmatory claim needs at least 20 completed sessions per arm. Below that we report the results as a description and say so. Recruitment is not planned. Sessions that arrive through the public link are reported as a description with their count, and nothing in this project depends on them.
8. People and models on the same test. Each configured vision model answers the same 16 items with the same question wording, three repeat runs, settings recorded. We report each model's accuracy beside the two human arms with a Wilson interval, and say plainly that 16 items is a small set. People see the photo at phone size. Models receive it resized to 1092 px on the long side. Pass rule for rung 1, fixed now: a model passes a feature only if it answers all 4 items for that feature correctly in at least 2 of its 3 runs. Only a passed feature may ever produce a flag. The larger benchmark in rung 3 (target 150 photos) replaces this rule with one written into this plan as a dated deviation before that benchmark is run.
9. What we publish whatever happens. All results, including no effect or a negative effect, and any model that beats trained people. The anonymous response table at data lock. Exclusion counts, the deviations log and the cost of the model runs. The README shows the three test items with the largest gap between trained-arm accuracy and the best model's accuracy, in either direction, chosen by script.
10. Ethics and privacy. Consent screen as described in the repo. Stored: random session id, arm, answers, timings, a hashed random browser token, coarse device class, consent version, and a coarse source label taken from the link (poster, chat, friends, creek_group, other). Nothing else. Photos contain no faces, house numbers or licence plates. Automatic backups of the raw database are taken daily and kept private. Nobody computes outcomes from them. The only code that computes outcomes is the analysis script, which refuses to run before data lock.
11. Exploratory, run after lock, reported as description only, and kept off the README's first screen and out of the video: does leaving out low scorers change what a group gets right? The 16 items are split at random into two halves with two items of each feature in each half (seed 20260920, 200 splits). A person passes a half with 6 or more of its 8 items correct. For random groups of 5 within an arm (2,000 draws), the group answers each item of the other half by plain majority, once using everyone and once using only the people who passed the first half. When nobody passed, or the passers tie, the group falls back to everyone. We report the share of items each version gets right. Our own simulation before launch says four photos per feature is too coarse to weight votes by feature, and that this filter helps only when a fair share of people answer at random. We publish what we find.
