Research basis
Measuring human advancement in AI-mediated systems
How can we distinguish stronger AI-assisted performance from durable human capacity and retained human authority?
Output is being measured. The human is not.
AI systems are increasingly evaluated through productivity, quality, adoption, and performance. Those measures do not establish what happened to the human who produced the output alongside the system.
A configuration can look stronger every quarter while the person inside it learns less, decides less, and retains less than the numbers imply.
Four things routinely get collapsed into one.
Institutions asking whether AI is helping people tend to ask a single question where at least four are required.
Each can move independently of the others. A person can produce stronger work, retain nothing of it, direct none of the process, and hold no standing to decide the result is sound—all at once.
The literature already establishes important pieces of this problem.
The remaining challenge is integrating those distinctions into a practical, longitudinal measurement method—not discovering that the distinctions exist.
- Noy & Zhang — productivity effects of generative AI
- Brynjolfsson et al. — AI-assisted performance and skill gaps
- Bastani et al. — AI tutoring and learning outcomes
- Wu, Liu et al. — assisted performance and later intrinsic motivation
- Vaccaro et al. — human–AI collaboration performance
- Di Santi — automation and skill retention
- Meaningful Human Control (Santoni de Sio & van den Hoven)
- Distributed and extended cognition
- Relevant UN, ILO, and UNESCO governance frameworks
Institutional context, not primary evidence: Stanford HAI’s accounts of human-centered scientific discovery and AI-accelerated discovery help frame the policy problem. The empirical claims above remain tied to the underlying studies.
Full citations and context for each source are maintained on Related Work, credited there as external scholarship—not Perfinitive's own findings.
Ten dimensions. Not one score.
A single index would hide exactly the distinctions this brief is arguing for. The diagnostic keeps ten dimensions separate.
An evidence-status layer, not a leaderboard.
Each dimension is reported with its own evidence status—observed, self-reported, inferred, or unmeasured—rather than folded into an aggregate. The diagnostic is designed to be legible to the person it describes, not only to the institution reading it.
Full architecture: Measurement.
A bounded 90-day pilot, published evidence first.
- Begin from published studies and existing datasets before collecting anything new.
- No intrusive collection of personal conversations or private interaction logs.
- Stage the ten dimensions as observable, contestable tasks—not silent scoring.
- Submit the method and early results for independent replication.
No results are reported yet. This is a proposed pilot design, not a completed study.