Research Brief · Beyond Performance
Does better AI-assisted output mean the person actually improved?
Micheal Charles Preble · Independent Researcher
The Question
What evidence is required before a claim that AI-mediated performance improved can support the stronger claim that a human being, or a population, advanced?
The Problem
Claims that AI has improved human performance increasingly appear in education, professional work, and public administration. But improvement in an AI-mediated system does not by itself establish that the human participant has advanced. A person may produce stronger outputs while losing capacities necessary for meaningful agency, while formal oversight persists after practical authority has migrated elsewhere. The reverse mistake matters too: a decline in unaided performance does not automatically establish human regression, because reliable external systems can expand genuine human capability.
The Contribution
A diagnostic method built on three configurations of how capability can be arranged — autonomous capability, governed extended capability, and ungoverned dependency — and a three-pattern diagnostic that classifies observed change as coupled improvement, performance-only acceleration, or regressive decoupling. Before evaluation, a claimant must declare six elements: the human functioning at issue, the population and role, the scope of the claim, the required contrast, the durability window, and the failure signs that would defeat the claim. The method then examines five human-side conditions: governance-critical capacity, authority in practice, governability and continuity, distribution, and time.
How the Framework Works
The method returns one of three dispositions — succeed, unproven, or fail — for a declared claim against its required evidence. It does not produce a composite score; the disposition follows from the relationship between the claim, the required evidence, and the observed failure signs.
The paper applies the method to two published third-party studies as worked examples, not as Preble's own data collection. Bastani et al. (2025, PNAS) ran a preregistered field experiment with nearly 1,000 high-school math students: standard GPT-4 access raised assisted-practice performance 48% while unassisted exam scores fell 17% below a no-AI control; a guarded tutor removed the exam penalty but produced no positive unassisted gain. Under the method, “assisted performance improved” succeeds for the standard condition; “human learning improved” fails for it and remains unproven for the guarded tutor. Strömberg, Lei & Wu (2026, CEPR working paper) followed 26,811 Chinese secondary students for 30 months: homework scores rose 18%, homework time fell 30%, and closed-book exam scores fell 20% within six months, concentrated among the roughly 80% of users whose usage pattern was consistent with greater outsourcing. The broader capability claim fails for that majority and remains unproven for the minority, where no subgroup-level transfer measure exists.
Why It Matters
Evaluation of AI systems routinely stops at output. This method insists that a claim about the human participant requires evidence about the human participant — not a proxy drawn from system performance. It rejects both a blanket warning against AI assistance and a blanket assumption that better output means a better-off person.
Evidence Status
Evidence status: A pre-validation decision procedure, not a validated instrument. It is applied to two published third-party empirical studies as worked examples; it does not report new data collection of its own.
What This Does Not Claim
- Does not propose a universal index or a validated instrument.
- Is not equivalent to a general warning that AI assistance causes deskilling.
- Does not claim capacity, control, distribution, or time are newly discovered dimensions — the contribution is procedural, not ontological.
- Does not exercise its own residual “human conversion gain” category.
- Does not settle long-run effects on educational pathways or switching costs; it inherits the limits of the studies it evaluates.
Open Questions
Whether independent evaluators applying the same declared claim and failure criteria reach similar dispositions across cases is untested (predeclaration robustness). Cross-domain validation — public administration, professional work — remains future work. The thresholds separating material from immaterial deterioration are domain-specific and not yet fixed.
Read / Cite the Research
Read the canonical paper record → · Full paper PDF · View on SSRN
Preble, Micheal Charles. “Beyond Performance: A Diagnostic Method for Evaluating Human-Advancement Claims in AI-Mediated Systems” Manuscript v2.3. Available at SSRN 7416901, 2026.