Human–AI Trust Lab
Synthetic — validation only
Results

Learning & dynamics

How reliance changes over the course of a session, and how consistent that pattern is across participants.
Trust trajectory over the session
Trust trajectory over the session
Trust trajectories: acceptance of the AI recommendation and appropriate reliance in windows of ten trials by AI-accuracy arm. Bands: 95% CI across participants.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Post-error trust adjustment
Post-error trust adjustment
Post-error trust adjustment: change in the probability of following the AI on the trial after an observed AI error vs after an observed success, overall, in early/late windows of each block, and by accuracy arm. 95% CI across participants.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Post-error adjustment, overall (disagreement trials, immediate feedback)

2.1% [-1.6%, 5.8%] · one-sample t = 1.13, p = 0.262

Participant heterogeneity
Participant heterogeneity
Distributions of participant-level learning slopes and post-error trust adjustments. Values left of zero indicate declining appropriate reliance / trust withdrawal after an observed AI error.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0
Recovered reliance by simulated archetype (synthetic only)
Recovered reliance by simulated archetype (synthetic only)
Simulator validation: relative AI reliance (RAIR) vs relative self-reliance (RSR) recovered for each simulated archetype. Ground-truth labels exist only for synthetic data.
Synthetic — validation onlydataset sim-v1-001 · sha 1507f368 · code 57cb0777+ · research/scripts/01_analyze.py · v0.1.0

What "learning" means here

The learning slope is the within-participant trend in appropriate reliance (following a correct AI, or overriding a wrong one) across a session's disagreement trials — not raw accuracy, which is dominated by task difficulty. A slope near zero over 60 trials is not surprising; it is the reason this design pairs a short session with feedback conditions in the second experiment (v2) rather than expecting large within-session learning in v1. See methods.