LLM judges are easier to talk down than up: a Bayesian item response analysis

Publication
NeurIPS 2026 Workshop on Trust in AI Evaluation (TAE)

Related