LLM Judge Bias
LLMジャッジの偏り
LLMを評価者として用いるとき、評価基準とは関係の薄い出力の順序、長さ、表現、説得力などの要因によって判断が系統的に傾く問題です。
ARC-V1-029
例
内容が同程度の二つの回答を比べたとき、評価基準と無関係なのに先に表示された回答や長い回答を高く評価する。
区別・注意
LLM-as-a-Judgeが評価方法そのものを指すのに対し、LLM Judge Biasはその評価に入り込む偏りを指します。ここに挙げた順序や長さなどは確認対象になり得る要因の例であり、すべての偏りを網羅する固定分類ではありません。
Evidence
Evidenceを見る →
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SRC-F04-004
- タイトル
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- 著者・組織
- Zheng et al.
- 年
- 2023
- 種別
- NeurIPS Datasets & Benchmarks
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM-as-a-Judge; LLM Judge Bias; human agreement; position/verbosity/self-enhancement bias
- 範囲
- MT-Bench + Chatbot Arena; GPT-4-era judges
- 限界
- early-generation judges; >80% agreement is setting-specific
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/075280-2020
Judging the Judges: A Systematic Study of Position Bias
SRC-F04-005
- タイトル
- Judging the Judges: A Systematic Study of Position Bias
- 著者・組織
- Shi et al.
- 年
- 2025
- 種別
- IJCNLP/AACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; position bias varies by judge/task/candidate gap
- 範囲
- 15 judges, MTBench/DevBench, 22 tasks, about 40 generators, over 150k evaluations
- 限界
- position-bias subtype only
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.18653/v1/2025.ijcnlp-long.18
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
SRC-F04-006
- タイトル
- Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- 著者・組織
- Hwang et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; Grader Gaming boundary; persuasion inflates scores for incorrect solutions
- 範囲
- 6 math benchmarks, 7 persuasion techniques
- 限界
- adversarial rhetoric; not proof of spontaneous gaming
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.790
Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation
SRC-F04-007
- タイトル
- Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation
- 著者・組織
- Li et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; LLM-as-a-Judge; auxiliary-information-induced biases; task complexity
- 範囲
- ComplexEval Bench; 12 basic + 3 advanced scenarios
- 限界
- benchmark-defined complex evaluation
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.805