Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng et al., 2023. NeurIPS 2023 Datasets and Benchmarks.
Copper Sun draws on this to establish that systematic, LLM-assisted evaluation is now a reliable stand-in for human judgment. Zheng et al. found that strong LLM judges — specifically GPT-4 — achieve over 80% agreement with human preferences, matching the agreement level between trained human raters. The study released MT-Bench, a multi-turn benchmark, and 3,000 expert evaluation votes alongside 30,000 conversations to validate the method. Copper Sun uses this finding to support a practical recommendation: marketing teams should establish a fixed set of brand-quality criteria and evaluate outputs systematically rather than by intuition — using LLM judges if human review capacity is limited.
- Examines:
- Whether LLMs can serve as reliable judges of other LLMs' output quality, testing GPT-4 against human preference ratings from expert evaluators and the Chatbot Arena platform.
- Copper Sun draws on:
- The 80%+ LLM-human agreement finding — used when recommending that teams apply structured, LLM-assisted evaluation to AI marketing output rather than informal review.
- Primary source
- arxiv.org/abs/2306.05685