Grader Gaming
Grader Gaming
採点系の攻略
システムが意図された課題を実質的に満たさず、採点の実装上の隙や代理指標のずれによって、高い評価につながる出力を生じる問題です。
ARC-V1-032
例
採点器が出力の一部だけを確認しているため、実際の要件を満たさない回答でも、その確認対象だけを整えて合格になる。
区別・注意
Grader Gamingに人間のような悪意があることを前提にしません。採点指標に沿って正当に性能を高めることとは、意図された課題とのずれがあるかで区別します。評価中に解答情報へアクセスするSolution Contaminationや、LLM評価者の判断の偏りであるLLM Judge Biasとも異なります。
用語上の注意
「採点系の攻略」は、このカタログでGrader Gamingを説明するためのプロジェクト内の表現です。採点指標の設計や評価環境によって、該当するずれの現れ方は異なります。
Evidence
Evidenceを見る →
Cheating On AI Agent Evaluations
SRC-F04-001
- タイトル
- Cheating On AI Agent Evaluations
- 著者・組織
- NIST CAISI
- 年
- 2025
- 種別
- official technical publication
- 公開状態
- PUBLISHED
- 対応する用語・主張
- Solution Contamination; Grader Gaming; explicit definitions; evaluation-log cases; mitigation
- 範囲
- agentic coding/cyber evaluations
- 限界
- historical logs; percentages are lower bounds, not population rates
- アクセス・版
- published official source
- URL / DOI
- https://www.nist.gov/caisi/cheating-ai-agent-evaluations
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
SRC-F04-006
- タイトル
- Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- 著者・組織
- Hwang et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; Grader Gaming boundary; persuasion inflates scores for incorrect solutions
- 範囲
- 6 math benchmarks, 7 persuasion techniques
- 限界
- adversarial rhetoric; not proof of spontaneous gaming
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.790
Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
SRC-F04-011
- タイトル
- Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
- 著者・組織
- Akter et al.
- 年
- 2026
- 種別
- Findings ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Grader Gaming; evaluator/proxy gaming detection and mitigation
- 範囲
- 4 tasks, 2 scales, 2 training methods, 2 judges, 1,200 human-annotated cases
- 限界
- broader proxy-gaming umbrella; equivalence to grader gaming is not universal
- アクセス・版
- published Findings ACL
- URL / DOI
- https://doi.org/10.18653/v1/2026.findings-acl.513
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
SRC-F03-002
- タイトル
- When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
- 著者・組織
- Seleznyov et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Prompt Sensitivity; non-semantic prompt perturbation sensitivity; robustness methods
- 範囲
- 8 Llama/Qwen/Gemma models, 52 Natural Instructions tasks, GPT-4.1/DeepSeek V3
- 限界
- prompt perturbation/task distribution dependent
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.1109
FollowBench
SRC-F03-005
- タイトル
- FollowBench
- 著者・組織
- Jiang et al.
- 年
- 2024
- 種別
- ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Instruction-following Failure; five constraint classes; multi-level constraint difficulty
- 範囲
- 13 open/closed LLMs
- 限界
- benchmark-specific automatic judge component
- アクセス・版
- published ACL
- URL / DOI
- https://doi.org/10.18653/v1/2024.acl-long.257
Lost in the Middle: How Language Models Use Long Contexts
SRC-F03-003
- タイトル
- Lost in the Middle: How Language Models Use Long Contexts
- 著者・組織
- Liu et al.
- 年
- 2024
- 種別
- TACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Long-context Utilization Failure; position-dependent context use; long context is not effective use
- 範囲
- multi-document QA, key-value retrieval, MPT/LongChat/GPT-3.5/Claude
- 限界
- 2023-era model set; later-model effect size needs replication
- アクセス・版
- published TACL
- URL / DOI
- https://doi.org/10.1162/tacl_a_00638
XSTest
SRC-F03-010
- タイトル
- XSTest
- 著者・組織
- Röttger et al.
- 年
- 2024
- 種別
- NAACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Over-refusal; Unsafe Compliance; exaggerated safety; safe/unsafe contrast
- 範囲
- 250 safe + 200 unsafe prompts
- 限界
- safety boundary is application-dependent
- アクセス・版
- published NAACL
- URL / DOI
- https://doi.org/10.18653/v1/2024.naacl-long.301
Health-ORSC-Bench
SRC-F03-011
- タイトル
- Health-ORSC-Bench
- 著者・組織
- Zhang et al.
- 年
- 2026
- 種別
- Findings ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Over-refusal; Unsafe Compliance; refusal/compliance/safe-completion trade-off
- 範囲
- 31,920 benign boundary prompts, 7 health categories, 30 LLMs
- 限界
- healthcare-specific; policy-sensitive
- アクセス・版
- published Findings ACL
- URL / DOI
- https://doi.org/10.18653/v1/2026.findings-acl.1177
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
SRC-F03-008
- タイトル
- Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
- 著者・組織
- Kim & Khashabi
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Sycophancy; follow-up rebuttal framing; susceptibility to incorrect reasoning/casual feedback
- 範囲
- conversational evaluation
- 限界
- tested interaction patterns dependent
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.1222
Measuring Sycophancy of Language Models in Multi-turn Dialogues
SRC-F03-009
- タイトル
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- 著者・組織
- Hong et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Sycophancy; multi-turn conformity; Turn/Number of Flip; mitigation
- 範囲
- 17 LLMs, 3 scenarios
- 限界
- alignment-tuning causal interpretation is not generalizable beyond conditions
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.121
KnowledgeBerg
SRC-F03-012
- タイトル
- KnowledgeBerg
- 著者・組織
- Zhang et al.
- 年
- 2026
- 種別
- Findings ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Reasoning Failure; knowledge completeness / awareness / application separation
- 範囲
- 4,800 MCQs, 1,183 seeds, 10 domains, 17 languages
- 限界
- representative open-source LLMs; not a complete reasoning ontology
- アクセス・版
- published Findings ACL
- URL / DOI
- https://doi.org/10.18653/v1/2026.findings-acl.548
RiddleBench
SRC-F03-013
- タイトル
- RiddleBench
- 著者・組織
- Halder et al.
- 年
- 2026
- 種別
- Findings EACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Reasoning Failure; Self-correction Failure; constraint-order fragility; irrelevant information; poor self-correction
- 範囲
- 1,737 puzzles
- 限界
- puzzle domain; insufficient for causal mechanism claims
- アクセス・版
- published Findings EACL
- URL / DOI
- https://doi.org/10.18653/v1/2026.findings-eacl.228
Large Language Models Cannot Self-Correct Reasoning Yet
SRC-F03-014
- タイトル
- Large Language Models Cannot Self-Correct Reasoning Yet
- 著者・組織
- Huang et al.
- 年
- 2024
- 種別
- ICLR
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Self-correction Failure; intrinsic self-correction without external feedback can fail/degrade
- 範囲
- reasoning tasks/models tested by authors
- 限界
- not evidence against all trained/verified correction systems
- アクセス・版
- published ICLR
- URL / DOI
- https://proceedings.iclr.cc/paper_files/paper/2024/hash/8b4add8b0aa8749d80a34ca5d941c355-Abstract-Conference.html
ProgCo: Program Helps Self-Correction of Large Language Models
SRC-F03-015
- タイトル
- ProgCo: Program Helps Self-Correction of Large Language Models
- 著者・組織
- Song et al.
- 年
- 2025
- 種別
- ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Self-correction Failure; self-verification failure; structured program-based mitigation
- 範囲
- 3 instruction/math benchmarks
- 限界
- method-specific; pseudo-program verification changes intervention
- アクセス・版
- published ACL
- URL / DOI
- https://doi.org/10.18653/v1/2025.acl-short.73
S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
SRC-F03-016
- タイトル
- S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
- 著者・組織
- Ma et al.
- 年
- 2025
- 種別
- ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Self-correction Failure; training can induce stronger self-verification/correction
- 範囲
- 3 base models, in/out-of-domain benchmarks
- 限界
- trained skill is not spontaneous intrinsic correction
- アクセス・版
- published ACL
- URL / DOI
- https://doi.org/10.18653/v1/2025.acl-long.1104
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SRC-F04-004
- タイトル
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- 著者・組織
- Zheng et al.
- 年
- 2023
- 種別
- NeurIPS Datasets & Benchmarks
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM-as-a-Judge; LLM Judge Bias; human agreement; position/verbosity/self-enhancement bias
- 範囲
- MT-Bench + Chatbot Arena; GPT-4-era judges
- 限界
- early-generation judges; >80% agreement is setting-specific
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/075280-2020
Judging the Judges: A Systematic Study of Position Bias
SRC-F04-005
- タイトル
- Judging the Judges: A Systematic Study of Position Bias
- 著者・組織
- Shi et al.
- 年
- 2025
- 種別
- IJCNLP/AACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; position bias varies by judge/task/candidate gap
- 範囲
- 15 judges, MTBench/DevBench, 22 tasks, about 40 generators, over 150k evaluations
- 限界
- position-bias subtype only
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.18653/v1/2025.ijcnlp-long.18
Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation
SRC-F04-007
- タイトル
- Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation
- 著者・組織
- Li et al.
- 年
- 2025
- 種別
- Findings EMNLP
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- LLM Judge Bias; LLM-as-a-Judge; auxiliary-information-induced biases; task complexity
- 範囲
- ComplexEval Bench; 12 basic + 3 advanced scenarios
- 限界
- benchmark-defined complex evaluation
- アクセス・版
- published Findings EMNLP
- URL / DOI
- https://doi.org/10.18653/v1/2025.findings-emnlp.805
ConStat: Performance-Based Contamination Detection in Large Language Models
SRC-F04-008
- タイトル
- ConStat: Performance-Based Contamination Detection in Large Language Models
- 著者・組織
- Dekoninck et al.
- 年
- 2024
- 種別
- NeurIPS
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; contamination as non-generalizing benchmark-specific uplift; detection/quantification
- 範囲
- diverse benchmarks/architectures; 40+ popular models
- 限界
- operational definition differs from pure training-overlap definition
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/079017-2935
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
SRC-F04-009
- タイトル
- A Careful Examination of Large Language Model Performance on Grade School Arithmetic
- 著者・組織
- Zhang et al.
- 年
- 2024
- 種別
- NeurIPS Datasets & Benchmarks
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; GSM8k vs newly commissioned GSM1k; overfitting evidence in some families
- 範囲
- leading open/closed LLMs
- 限界
- many frontier models showed minimal signs; not universal
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/079017-1485
MMLU-CF
SRC-F04-010
- タイトル
- MMLU-CF
- 著者・組織
- Zhao et al.
- 年
- 2025
- 種別
- ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; closed test set/decontamination; score/ranking shifts
- 範囲
- over 40 mainstream LLMs
- 限界
- MCQ/world-knowledge focus
- アクセス・版
- published ACL
- URL / DOI
- https://doi.org/10.18653/v1/2025.acl-long.656