**Pre-specified analysis plan, second wave: the same study in a second window**

Not yet tagged. It becomes binding when tagged `prereg-v3`, before any sitting of the second wave. It changes nothing in `docs/analysis_plan.md` (tag `prereg-v1`) or in `docs/analysis_plan_v2.md` (tag `prereg-v2`): both stay as tagged, and the first wave's result stays as reported. This plan says what is the same by pointing at those two plans, and spells out only what differs. The code enforces this text: `evals/wave2_analysis.py` refuses to run on real data before the second lock, and refuses to run at all if the tag is missing, if this file differs from the tagged version, or if one of the two earlier plans or the two analysis scripts differs from its copy at the tag.

This is a usability test of our own tool, run to improve it and to report honestly how well it works. Participation is anonymous and voluntary. We collect no names, emails, IP addresses or precise locations. It is not designed or presented as generalizable human-subjects research.

1. Why a second wave. No real person took part before the first lock, 2026-09-28T01:00:00Z. The lock job ran both analyses on 2026-09-29 and each said "not computed: an arm is empty", with 0 sittings kept in each arm: `results/usability_20260929.json` and `results/assist_20260929.json`. The sittings in that export were our own checks and one automated walk of ours (`docs/deviations.md`, the entries dated 2026-09-25T17:48:55Z and 2026-09-29). After that the hackathon moved its deadline from Sep 30 to 2026-10-05T04:00:00Z, which is Sunday Oct 4 at 21:00 PDT. So the same study runs a second time, in a new window, with the same design and the same analysis.
2. The window. A part 1 sitting belongs to the second wave when the server stored its start at or after 2026-09-30T04:00:00Z, which is Tuesday Sep 29 at 21:00 PDT, and before the second lock, 2026-10-03T04:00:00Z, which is Friday Oct 2 at 21:00 PDT, 48 hours before the deadline. A part 2 sitting belongs to it when the part 1 sitting it follows does, and when it was itself offered or started before the second lock.
3. What is the same. Everything in `prereg-v1` and `prereg-v2` but the dates. The question of each part. The two arms of each part and their randomization by the server, 1 to 1 in permuted blocks of 4: the pre-registered sequences are not made again, they carry on from the slot where they stand. The materials: the 16 test items and the 8 part 2 items, their photos, gold labels, lessons and question wording, and the checker's flags in `results/assist_flags.json`. The primary outcome of each part. The one confirmatory test per part: a two-sided permutation test on the difference in means, 10,000 permutations, alpha 0.05, with a percentile bootstrap interval of 10,000 resamples, seed 20260920 for part 1 and seed 20260926 for part 2. The descriptive outcomes and the two sensitivity checks of part 1. The exclusions of each plan, in each plan's order. The rule for small numbers: an arm with fewer than 20 kept sittings makes that part's result a description, not a test, and we say so. The code is the same code: `evals/usability_analysis.py` and `evals/assist_analysis.py`, not edited, called by `evals/wave2_analysis.py`. Two items of `prereg-v1` are not run again: item 8, the models on the same test, whose results stand as reported, and item 11, the exploratory question, which is not part of the second wave's one run.
4. How the second wave reads the export. The server marks every sitting made after the first lock as `post_lock`, so that mark is true for the whole second wave and cannot say which sittings count. The second wave ignores the stored mark. The start time decides. A sitting that started at or after the second lock is dropped under each plan's own rule for sittings after the lock. A part 1 sitting that started before the wave opened is left out as a dry run is, under the rule "dry run before launch" of `prereg-v1` item 5, with the wave's opening as the launch. A sitting whose start time cannot be read is dropped with those after the lock. The rules run in the order of `prereg-v1`, so the rule for repeat visits comes before the rule for dry runs: a browser whose first finished sitting was before the wave opened is left out of the second wave, because that person has seen the photos. A part 2 sitting whose part 1 sitting started outside the window is left out before part 2's exclusions and counted. As in `prereg-v1`, the counts of randomized, started and completed sittings are taken before the exclusions, so they include sittings from before the wave opened; each results file also counts the sittings by start time, in its `window` block.
5. What a reader must hold against this wave. Judge mode shows the answers to the same photos the study uses. It was open from the first lock, 2026-09-28T01:00:00Z, until it was shut again on Tuesday Sep 29, Pacific time, before the wave opened, and it stays shut until the second lock. A person who used it while it was open may know answers, and this plan cannot tell who did. The second wave was decided after the first analysis had run. That run showed no outcome: no score and no difference, only that an arm was empty. Pages outside the test changed between the waves. The test's own screens, items, photos, lessons and consent did not, and neither did part 2's; the study's routes and the study's functions in the Worker are the same. Recruitment differs from what `prereg-v1` item 7 planned: a paid research panel shows the study to its members and pays them, and the public link stays open. Sittings are reported by source, as description only. What `prereg-v1` says about the gold labels holds here too: one person set the key.
6. Sample size and stopping. The target of `prereg-v1` item 7 stands: 80 completed sessions, 40 per arm, with no promise of reaching it. The second lock is 2026-10-03T04:00:00Z whatever the count. Before the lock we look at counts per arm only.
7. The one run. The analysis of the second wave runs once, by the lock job at 2026-10-03T04:10:00Z, ten minutes after the second lock: `scripts/lock_analysis.py --wave 2` runs `evals/wave2_analysis.py` once, from a backup taken after the lock. It writes `results/usability_w2_<date>.json` and `results/assist_w2_<date>.json`, which have the shape of the first wave's files plus the window. The README reports the second wave in rows of its own. The first wave's rows and files stay as they are, and the two waves are not pooled.
8. What we publish whatever happens. As `prereg-v1` item 9 and `prereg-v2` item 10, for the second wave: all results, including no effect, a negative effect, too few people, or nobody at all. The anonymous response tables at the second lock. Exclusion counts and the deviations log.
9. Ethics and privacy. As `prereg-v1` item 10 and `prereg-v2` item 11. The second wave stores nothing the first did not.
