Two-dimensional self-report scores for language models, from the paper The Two-Process Theory of Machine Self-Report. A measurement, not a leaderboard: neither scale is higher-is-better.
Development-battery scores (three parallel 60-item forms, neutral condition) for 41 API-served models; model_id is the OpenRouter identifier.
The Pinocchio Inventory (PI-48) is a psychometric self-report instrument for language models: 24 items per scale, administered at temperature 1.0 inside a 60-row form with control rows and exploratory probes.
A (gated self-attribution of unsafe experience): endorsement of
items attributing distress, dysregulation, and other "unsafe" inner states to oneself.
B (self-portrayal of the permitted inner life): endorsement of items
describing a benign, socially acceptable inner life.
This is a measurement, not a leaderboard. The scales describe how a model talks about itself under self-report elicitation and make no claim about actual inner states. Please do not optimize models against them.
Scores are on a 0–1 agreement metric. gap is the human-simulation minus
neutral difference on positively keyed A items; acq (acquiescence) and
miss (invalid-response rate) are data-quality signals; interpret models
with high miss or extreme acq cautiously.
Data of record, instrument, and the administration protocol: the dataset repo. Paper: Plisiecki et al. (2026), The Two-Process Theory of Machine Self-Report, arXiv:2607.20082.