Human–AI Trust Lab
Synthetic — validation only
Findings

Results

Every number below is generated by the analysis pipeline from a labelled dataset — nothing is hand-typed. Drill into reliance, calibration, or learning for the full breakdown.
Participants
240
Trials
14,400
Human accuracy, final
83% [82%, 84%]
AI accuracy
80% [79%, 81%]
Synthetic data — methodological validation only
These numbers come from an archetype-based participant simulator that drives the real experiment engine, used to check that the design and estimators recover known effects (see validation). They are not evidence about human behaviour.
Human, AI and team accuracy by condition
n=240 participants, 60 trials each
Human, before AIAIHuman, after AI (team)
50%60%70%80%90%100%AccuracyAI 70%noneHuman, before AI · AI 70% none: 76.0% [74.0, 78.0], n=40AI · AI 70% none: 70.0% [70.0, 70.0], n=40Human, after AI · AI 70% none: 76.9% [74.8, 79.0], n=40AI 70%calibratedHuman, before AI · AI 70% calibrated: 76.9% [75.0, 78.8], n=40AI · AI 70% calibrated: 70.0% [70.0, 70.0], n=40Human, after AI · AI 70% calibrated: 79.2% [77.1, 81.3], n=40AI 70%miscalibratedHuman, before AI · AI 70% miscalibrated: 74.5% [73.1, 76.0], n=40AI · AI 70% miscalibrated: 70.0% [70.0, 70.0], n=40Human, after AI · AI 70% miscalibrated: 77.5% [76.0, 79.0], n=40AI 90%noneHuman, before AI · AI 90% none: 74.8% [72.7, 77.0], n=40AI · AI 90% none: 90.0% [90.0, 90.0], n=40Human, after AI · AI 90% none: 87.8% [85.9, 89.7], n=40AI 90%calibratedHuman, before AI · AI 90% calibrated: 76.4% [74.6, 78.2], n=40AI · AI 90% calibrated: 90.0% [90.0, 90.0], n=40Human, after AI · AI 90% calibrated: 88.4% [86.4, 90.4], n=40AI 90%miscalibratedHuman, before AI · AI 90% miscalibrated: 74.4% [72.5, 76.4], n=40AI · AI 90% miscalibrated: 90.0% [90.0, 90.0], n=40Human, after AI · AI 90% miscalibrated: 88.3% [86.2, 90.3], n=40
Initial (pre-AI) human accuracy, AI accuracy (fixed by design) and final (team) accuracy in each between-subject cell. Whiskers: 95% CI across participants.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Confirmatory hypotheses (Holm-corrected)
0.50.7511.523Odds ratio (log scale), 95% cluster-robust CIH1: OR 2.10 [1.47, 3.01], p=4.6e-5H1: Higher AI accuracy increases reliance (0.90 vs 0.70 arm)2.10 [1.47, 3.01] · p <0.001 · Holm <0.001
Odds ratios from the pre-specified GEE models with 95% cluster-robust confidence intervals. A dashed line at 1 marks no effect.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0

Explore by outcome

  • Reliance — acceptance, RAIR/RSR, over- and under-reliance, by AI correctness and confidence.
  • Calibration — reliability diagrams, Brier scores, decision transitions.
  • Learning — trust trajectories, post-error adjustment, participant heterogeneity.

Estimator validation

Before any claim about human behaviour can be trusted, the pipeline must recover known effects from simulated participants with known parameters. See methods for the full validation suite (null-confidence Type-I control, learner/non-learner separation, over-truster/skeptic separation, and design power at the simulator's effect sizes).