Improved Image Caption Rating — Datasets, Game, and Model
Research on constructing and validating human-centered rating systems for image captions.
View DOI →Latent Trait is an independent research institution developing calibrated measurement instruments for artificial intelligence systems. We apply Rasch measurement to populations of artificial respondents and items, with explicit attention to validity, calibration, invariance, uncertainty, and provenance.

A measurement is not merely a number. It is a defensible relationship among observations, an instrument, a scale, and the quantity being measured.
Artificial intelligence presents a new class of respondent, but the scientific requirements of measurement remain familiar: define the quantity, construct an instrument, calibrate it, establish its range and uncertainty, test its assumptions, and preserve a traceable record of how the result was obtained.
Establishing the scale on which observations become interpretable.
02Locating artificial respondents and items on a common latent scale.
03Building item populations that cover a defined construct and useful range.
04Testing whether comparisons remain stable across suitable measurement conditions.
05Reporting the precision and limits of each estimated location.
06Establishing whether observations behave as the instrument requires.
07Preserving instrument version, conditions, analysis, and lineage of a result.
08Maintaining comparability as instruments and measurement occasions change.
The observations are visible. The quantity is not. A valid instrument reveals the structure that makes the observations comparable.
Artificial respondents produce patterns across a population of items.
Rasch measurement estimates respondent locations and item difficulties on the same scale.
The distance between a respondent and an item has a probabilistic interpretation.
Measurement is not the number at the end. It is the calibrated instrument, scale, and procedure that give the number meaning.
Latent Trait grows out of published research in AI evaluation, human and machine judgment, and latent-variable measurement. The institution extends that work toward calibrated instruments for artificial respondents.
Research on constructing and validating human-centered rating systems for image captions.
View DOI →A validated dataset and measurement-oriented foundation for comparing image-caption quality.
View paper →Psychometric modeling of latent dispositions across multimodal and text-based narratives.
Publication record →Direct application of item response theory to the evaluation of vision-language models.
Publication record →Latent Trait is organized around a simple requirement: consequential claims about artificial systems should rest on instruments whose construction, calibration, uncertainty, and provenance can be examined independently of the systems being measured.
Measurement is institutionally separated from model development and vendor claims.
Instrument version, analysis, conditions, uncertainty, and provenance remain part of the scientific record.
The aim is comparable measurement that can survive changes in individual items, systems, and measurement occasions.