Synthetic — validation only
Findings
Results
Every number below is generated by the analysis pipeline from a labelled dataset — nothing is hand-typed. Drill into reliance, calibration, or learning for the full breakdown.
Participants
240
Trials
14,400
Human accuracy, final
83% [82%, 84%]
AI accuracy
80% [79%, 81%]
Synthetic data — methodological validation only
These numbers come from an archetype-based participant simulator that drives the real experiment engine, used to check that the design and estimators recover known effects (see validation). They are not evidence about human behaviour.
Human, AI and team accuracy by condition
n=240 participants, 60 trials each
Human, before AIAIHuman, after AI (team)
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Confirmatory hypotheses (Holm-corrected)
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Explore by outcome
- Reliance — acceptance, RAIR/RSR, over- and under-reliance, by AI correctness and confidence.
- Calibration — reliability diagrams, Brier scores, decision transitions.
- Learning — trust trajectories, post-error adjustment, participant heterogeneity.
Estimator validation
Before any claim about human behaviour can be trusted, the pipeline must recover known effects from simulated participants with known parameters. See methods for the full validation suite (null-confidence Type-I control, learner/non-learner separation, over-truster/skeptic separation, and design power at the simulator's effect sizes).