Proposed
A construct, requirement, instrument, or explanation offered for criticism and testing. It has not yet been empirically established.
Methods and versions
A conceptual proposal, a preliminary observation, and a replicated result can all be useful. They cannot carry the same evidentiary weight. This page records the distinctions used across the research program.
Evidence status
Labels describe what kind of support a claim presently has. They are not quality scores, and they should change when the underlying evidence changes.
A construct, requirement, instrument, or explanation offered for criticism and testing. It has not yet been empirically established.
An observation supported by limited evidence whose scope does not yet justify broad generalization.
An argument available in a public repository with identifiable authorship, date, citation, and stated limitations.
A question, conflict, or boundary that remains open because the available evidence or reasoning does not settle it.
Evaluation practice
Any released evaluation should allow a reader to reconstruct what was tested and what the resulting score leaves out. Aggregate performance is insufficient when failures, refusals, incomplete runs, or scoring assumptions materially change the interpretation.
Publish the task purpose, source materials, required output, staged changes, and criteria used to determine success. When the full corpus cannot be released, provide a representative inspectable subset and explain the restriction.
Identify the model, version, settings, tools, prompt, run date, and number of repetitions. Report completion and failure separately from accuracy on successful runs.
Record how the answer key was created, who reviewed it, how disagreements were reconciled, and where model assistance entered the process. Human review improves an answer key; it does not remove the need to disclose how the key was produced.
State corpus limits, unresolved ambiguities, known sources of bias, sponsorship, and relationships that could affect task selection, scoring, or interpretation.
Versioning
A correction should not erase the existence of an earlier claim or instrument. Material changes receive a new version with a date, a description of what changed, and a reason. Superseded versions remain identifiable unless legal or ethical constraints require removal.
Independence and assistance
Perfinitive is presented as an independent research program. That description does not eliminate possible conflicts. Any released evaluation should identify its funding, model or data access, compute support, provider participation, advance review, and other relationships that could affect task selection or interpretation.
OpenAI’s Codex assisted with recovery of the prior website, implementation of this redesign, source discovery, and editorial drafting. AI assistance does not determine which claims are retained or what evidence is sufficient. Micheal Preble remains responsible for the published language and research decisions.
Current record
The four papers listed in the publication archive are publicly available through SSRN. The evaluation and case-registry materials described on this site remain proposals in design. No benchmark scores, released datasets, or replicated findings are presently claimed.