Improvement in an AI-mediated system does not by itself establish that the human participant advanced.
Research Poster · Beyond Performance
Beyond Performance: A Diagnostic Method for Evaluating Human-Advancement Claims in AI-Mediated Systems
Micheal Charles Preble · Independent Researcher · SSRN 7416901
Problem · Background · Research Question
Builds on human-AI teaming and meaningful-control literatures (Vaccaro et al. 2024; Santoni de Sio & van den Hoven 2018) without claiming to replace them.
What evidence is required before AI-mediated performance improvement can support a stronger human-advancement claim?
Framework
Three configurations describe how capability is arranged: autonomous capability, governed extended capability, ungoverned dependency. Observed change is classified into three patterns: coupled improvement, performance-only acceleration, regressive decoupling.
Central diagram: the decision table
Beyond Performance is a diagnostic method, not a linear narrative, so its central representation is the decision table that maps claim type to required evidence — not a flowchart.
| Claim type | Required contrast | Fails if | Unproven if |
|---|---|---|---|
| Assisted performance | Treated vs. control, or tool-on vs. tool-off | Declared output measure does not improve | No valid performance contrast |
| Human learning or skill | Delayed unaided, transfer, or equivalent domain-valid measure | Assisted performance rises while declared human capability materially declines | No delayed or transfer measure |
| Meaningful control | Observed override, escalation, correction, or decision-right evidence | Responsibility stays human while practical influence becomes ineffective | Only formal/documentary rights measured |
| Governed extension | Assisted function plus continuity/recovery contrast | Functioning depends on a system the person cannot evaluate, correct, substitute, or recover from | Continuity and governability not tested |
| Broad social progress | Population- and role-disaggregated outcomes over a declared window | Mean improvement coexists with material deterioration in a relevant subgroup | Only aggregate or short-term evidence available |
Propositions / Observable Implications
- Predeclaration: the six declared elements (functioning, population, scope, contrast, durability, failure signs) should be fixed before evidence is interpreted.
- Disposition consistency: independent evaluators applying the same declared claim should reach similar succeed/unproven/fail dispositions.
- Directional deterioration: assisted performance rising while a predeclared transfer measure falls is sufficient to trigger failure of a learning claim, without needing a universal effect-size threshold.
Evidence Required
Bastani et al. (2025, PNAS): preregistered field experiment, ~1,000 HS math students — standard GPT-4 condition: +48% assisted practice, −17% unassisted exam vs. control; guarded tutor removed the penalty without a positive gain. Strömberg, Lei & Wu (2026, CEPR): 26,811 Chinese secondary students over 30 months — +18% homework score, −30% homework time, −20% closed-book exam within 6 months; losses concentrated among ~80% of users with an outsourcing-consistent usage pattern. These are third-party findings the method is applied to, not Preble's own data.
Limitations
Not a validated instrument. Material-deterioration thresholds vary by domain. Depends on predeclaration; vulnerable to post hoc rationalization if claims are redefined after seeing data. Inherits the limits of the evidence it evaluates.
Falsifiers / Open Questions
The method's reliability would be undermined if independent evaluators applying the same declared claim and failure criteria reached materially different succeed/unproven/fail dispositions. Cross-domain application (public administration, professional work) is untested.
Citation / QR
Preble, Micheal Charles. “Beyond Performance: A Diagnostic Method for Evaluating Human-Advancement Claims in AI-Mediated Systems” Manuscript v2.3. Available at SSRN 7416901, 2026.
research.perfinitive.com/publications/beyond-performance · no QR image is generated by this build.