Anchoring pulls judgment toward first impressions
In model evaluation, humans and LLMs alike judge later outputs in the shadow of earlier context, a pervasive bias that shapes scores in ways that may not reflect intrinsic quality.