Independent measurement researchlatenttrait.ai
Latent TraitLatent TraitMeasurement Science for Artificial Intelligence
Research

A research program in measurement.

The work leading to Latent Trait spans psychometrics, artificial intelligence evaluation, human and machine judgment, validated datasets, and accessible media.

Program

Population-based measurement of artificial systems.

The current program treats artificial systems as respondent populations and calibrated observations as measurement instruments.

Latent Trait focuses on Rasch measurement because it makes the requirements of measurement explicit: construct coherence, item calibration, respondent location, fit, uncertainty, invariance, and comparability across suitable instruments and occasions.

Research lineage
01

Accessible AI systems

Work on automated video description and machine perception established an early concern with systems whose outputs must be useful to people in consequential settings.

02

Validated evaluation

Image-caption rating work moved toward validated datasets and explicit evaluation procedures rather than ad hoc scoring.

03

Human & machine judges

Later work studies how VLMs and other models rate outputs, how those judgments compare with humans, and how their performance can be quantified.

04

Psychometric measurement

IRT and Rasch methods provide the measurement framework for separating respondent location from item difficulty and constructing common latent scales.

Selected work
2026
Discourse Processes

Disentangling local versus global processing dispositions in autistic cognition using item response theory models

Psychometric modeling of latent processing dispositions across multimodal and text-based narratives.

Public record

Evaluation of Vision Language Models with Item Response Theory

Direct application of item response theory to the evaluation of vision-language models.

2026
Preprint

Toward Scalable Audio Description Quality Control

A workflow for evaluating human and VLM raters in a shared quantitative framework.

2023
NeurIPS D&B

Validated Image Caption Rating Dataset

A validated dataset for systematic comparison of image-caption quality.

Complete publication record →