Research Poster · Beyond Performance

Beyond Performance: A Diagnostic Method for Evaluating Human-Advancement Claims in AI-Mediated Systems

Micheal Charles Preble · Independent Researcher · SSRN 7416901

Problem · Background · Research Question

Problem

Improvement in an AI-mediated system does not by itself establish that the human participant advanced.

Background

Builds on human-AI teaming and meaningful-control literatures (Vaccaro et al. 2024; Santoni de Sio & van den Hoven 2018) without claiming to replace them.

Research question

What evidence is required before AI-mediated performance improvement can support a stronger human-advancement claim?

Framework

Three configurations describe how capability is arranged: autonomous capability, governed extended capability, ungoverned dependency. Observed change is classified into three patterns: coupled improvement, performance-only acceleration, regressive decoupling.

Central diagram: the decision table

Beyond Performance is a diagnostic method, not a linear narrative, so its central representation is the decision table that maps claim type to required evidence — not a flowchart.

Claim typeRequired contrastFails ifUnproven if
Assisted performanceTreated vs. control, or tool-on vs. tool-offDeclared output measure does not improveNo valid performance contrast
Human learning or skillDelayed unaided, transfer, or equivalent domain-valid measureAssisted performance rises while declared human capability materially declinesNo delayed or transfer measure
Meaningful controlObserved override, escalation, correction, or decision-right evidenceResponsibility stays human while practical influence becomes ineffectiveOnly formal/documentary rights measured
Governed extensionAssisted function plus continuity/recovery contrastFunctioning depends on a system the person cannot evaluate, correct, substitute, or recover fromContinuity and governability not tested
Broad social progressPopulation- and role-disaggregated outcomes over a declared windowMean improvement coexists with material deterioration in a relevant subgroupOnly aggregate or short-term evidence available

Propositions / Observable Implications

  • Predeclaration: the six declared elements (functioning, population, scope, contrast, durability, failure signs) should be fixed before evidence is interpreted.
  • Disposition consistency: independent evaluators applying the same declared claim should reach similar succeed/unproven/fail dispositions.
  • Directional deterioration: assisted performance rising while a predeclared transfer measure falls is sufficient to trigger failure of a learning claim, without needing a universal effect-size threshold.

Evidence Required

Bastani et al. (2025, PNAS): preregistered field experiment, ~1,000 HS math students — standard GPT-4 condition: +48% assisted practice, −17% unassisted exam vs. control; guarded tutor removed the penalty without a positive gain. Strömberg, Lei & Wu (2026, CEPR): 26,811 Chinese secondary students over 30 months — +18% homework score, −30% homework time, −20% closed-book exam within 6 months; losses concentrated among ~80% of users with an outsourcing-consistent usage pattern. These are third-party findings the method is applied to, not Preble's own data.

Limitations

Not a validated instrument. Material-deterioration thresholds vary by domain. Depends on predeclaration; vulnerable to post hoc rationalization if claims are redefined after seeing data. Inherits the limits of the evidence it evaluates.

Falsifiers / Open Questions

The method's reliability would be undermined if independent evaluators applying the same declared claim and failure criteria reached materially different succeed/unproven/fail dispositions. Cross-domain application (public administration, professional work) is untested.

Citation / QR

Suggested citation
Preble, Micheal Charles. “Beyond Performance: A Diagnostic Method for Evaluating Human-Advancement Claims in AI-Mediated Systems” Manuscript v2.3. Available at SSRN 7416901, 2026.

research.perfinitive.com/publications/beyond-performance · no QR image is generated by this build.