Observational Equivalence of LLM and Human Annotation

Poster presentation at the Society for Political Methodology (PolMeth) Annual Meeting, July 2026. With Kentaro Nakamura and George Yean.

Presented under the earlier title “No ‘Human Labels as Gold Standard’: When LLMs Disagree, Humans Do Too”. We revised the paper after the conference and posted the current version to SSRN in September 2026.

Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and recent LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.

Preprint: Observational Equivalence of LLM and Human Annotation