Research protocol · v1
How we measure impact.
If you're going to claim you fix reasoning, you need to actually measure it. This page lays out the evidence layers, metric definitions, study design, and the boundaries around every published claim.
Want the actual question-to-misconception map? See the impact spine.
§ 00 Evidence layers
- 1Individual practice
A learner's answers, reports, and progress are retained in that browser for continuity and self-comparison.
unlocks · Private practice history and repeat-attempt comparisons.
- 2Cohort program
Consented learners complete baseline diagnostics, targeted work, and held-out follow-ups.
unlocks · De-identified recurrence and paired pre/post analysis.
- 3Public metricspublished
Live aggregate totals report reach, diagnostic adoption, practice activity, and cohort participation.
unlocks · Production-scale usage reporting with explicit metric definitions.
- 4Verified research
Consented paired records use a documented coding scheme and held-out follow-up form.
unlocks · Learning-outcome analysis with sample sizes, caveats, and reproducible definitions.
These layers coexist; they are not a launch sequence. Reach, adoption, engagement, and matched cohort outcomes remain distinct in every public summary.
§ 00b Verified reach
The funnel from reach to research.
65 of 19,185 reached the verified cohort. 45 matched pre/post pairs.
§ 01 What we measure
Every metric has a definition.
| Metric | Counts only when… |
|---|---|
| Recorded learner | A distinct anonymous browser id with a diagnostic completion, or a submitted cohort signup. |
| Active learner | Completed a diagnostic and took at least one ladder action. |
| Practice attempt | A ladder or problem-bank problem opened or attempted. Diagnostic responses are counted separately as diagnostic items answered. |
| Paired attempt | Same learner, two non-practice attempts: first = baseline, latest later one = follow-up. One pair per learner. |
| Pre/post gain | followup_score − baseline_score, computed per pair and reported with the pair count. Never as a lone headline number. |
| Trap recurrence | A misconception code present in both halves of a pair. Resolution = present pre, absent post. |
| Country represented | Only from a submitted country field (cohort or diagnostic), or verified analytics later. |
§ 02 Claim boundaries
What the evidence does not prove.
- 01
Social-media followers or reach as evidence of learning.
- 02
Unverified testimonials or invented quotes.
- 03
A fixed 'error patterns' count beyond the 66 authored Atlas entries actually in the product.
- 04
Causality. Even a positive paired gain is pre/post evidence, not a controlled trial.
- 05
Results beyond the measured population, reporting period, or consented cohort.
§ 03 The paired-study design
Baseline → work → held-out follow-up.
- 1
Baseline diagnostic
- 2
Coded report
- 3
Targeted ladder work
- 4
Held-out follow-up
- 5
Paired comparison
| How codes are assigned | Every wrong option on every diagnostic item is hand-mapped to one misconception code at authoring time. Diagnosis is deterministic — the same answers always produce the same codes. |
| Held-out follow-up | The follow-up form probes the same misconception list with different surface problems, so a gain measures reasoning transfer rather than memorised answers. |
| Double coding | Before any verified claim, Atlas mappings get a second independent rater; we report agreement. |
The follow-up set (fp-diag-followup-v1) uses coverage parity with the baseline and different surface problems. Repeat attempts on the baseline form remain labelled separately so they are not mistaken for held-out transfer evidence.
§ 04 Exports & the public record
Everything the writeup needs is exportable.
attempts.csv | One row per attempt: kind, question set, stage, score, elapsed, trap codes, recommendations. |
paired_summaries.csv | One row per learner pair: pre/post scores, delta, days between, recurring / resolved / new traps. |
signups.csv | Cohort roster (never published; used for consent and pairing only). |
| GitHub methodology | The Error Atlas taxonomy, distractor→code mappings, pairing rules, and export schemas. The full recipe, minus learner data. |
| Research writeup | Trap distribution across the deployed population; per-code recurrence after targeted ladder work; paired pre/post gains with pair counts and caveats. |
§ 05 Social reach vs. learning impact
Reach
19,185
Physics prep users across Learn pages, weekly challenges, Discord problem sessions, and public resources. Measures how far the ideas travel; reported separately, never mixed into learning claims.
Impact
45 paired
Matched pre/post pairs from 65 verified cohort members. Diagnostic completions (2,102), practice activity (9,902 practice attempts), and paired pre/post gains. Diagnostic responses are tracked separately from practice activity. This is the evidence layer used for learning-impact claims.
§ 06 Teacher & mentor verification
Teacher-verified classroom use
Evidence record
Mentor-endorsed learner outcomes
Evidence record
Independent educator review
Evidence record
Want to be part of the data that makes these numbers real?