Perfinitive
Research
Policy & research brief
2027

Research basis

Measuring human advancement in AI-mediated systems

How can we distinguish stronger AI-assisted performance from durable human capacity and retained human authority?

01 — The problem

Output is being measured. The human is not.

AI systems are increasingly evaluated through productivity, quality, adoption, and performance. Those measures do not establish what happened to the human who produced the output alongside the system.

A configuration can look stronger every quarter while the person inside it learns less, decides less, and retains less than the numbers imply.

02 — The distinction

Four things routinely get collapsed into one.

Institutions asking whether AI is helping people tend to ask a single question where at least four are required.

AI-assisted performance durable human capacity human agency practical authority

Each can move independently of the others. A person can produce stronger work, retain nothing of it, direct none of the process, and hold no standing to decide the result is sound—all at once.

03 — What existing research already establishes

The literature already establishes important pieces of this problem.

The remaining challenge is integrating those distinctions into a practical, longitudinal measurement method—not discovering that the distinctions exist.

  • Noy & Zhang — productivity effects of generative AI
  • Brynjolfsson et al. — AI-assisted performance and skill gaps
  • Bastani et al. — AI tutoring and learning outcomes
  • Wu, Liu et al. — assisted performance and later intrinsic motivation
  • Vaccaro et al. — human–AI collaboration performance
  • Di Santi — automation and skill retention
  • Meaningful Human Control (Santoni de Sio & van den Hoven)
  • Distributed and extended cognition
  • Relevant UN, ILO, and UNESCO governance frameworks

Institutional context, not primary evidence: Stanford HAI’s accounts of human-centered scientific discovery and AI-accelerated discovery help frame the policy problem. The empirical claims above remain tied to the underlying studies.

Full citations and context for each source are maintained on Related Work, credited there as external scholarship—not Perfinitive's own findings.

04 — What should be measured

Ten dimensions. Not one score.

A single index would hide exactly the distinctions this brief is arguing for. The diagnostic keeps ten dimensions separate.

Performanceobserved output
Capacityretained human capability
Agencydirection and choice
Judgmentevaluation of the work
Authorityoperative control
Persistencestate across time
Correctionrepair and propagation
Dependencereliance on assistance
Provenancetraceable origin
Portability / exitability to leave
05 — Proposed diagnostic

An evidence-status layer, not a leaderboard.

Each dimension is reported with its own evidence status—observed, self-reported, inferred, or unmeasured—rather than folded into an aggregate. The diagnostic is designed to be legible to the person it describes, not only to the institution reading it.

Measurement flow
Interaction record Ten-dimension diagnostic Evidence-status tags visible to the person measured

Full architecture: Measurement.

06 — Validation pilot

A bounded 90-day pilot, published evidence first.

  1. Begin from published studies and existing datasets before collecting anything new.
  2. No intrusive collection of personal conversations or private interaction logs.
  3. Stage the ten dimensions as observable, contestable tasks—not silent scoring.
  4. Submit the method and early results for independent replication.

No results are reported yet. This is a proposed pilot design, not a completed study.

07 — Why it matters

The same gap runs through six institutions at once.

EducationWhether AI-assisted learning produces durable capability or dependent output.
WorkWhether faster output this quarter comes at the cost of retained skill next year.
Public-sector AIWhether case workers and citizens retain judgment and standing, not just throughput.
Professional practiceWhether licensed judgment is being exercised or merely rubber-stamped.
Human developmentWhether AI-rich societies are building capability or renting it.
GovernanceWhether institutions can tell the difference before they regulate on the wrong variable.