← Catalogへ戻る

Grader Gaming

Grader Gaming

採点系の攻略

システムが意図された課題を実質的に満たさず、採点の実装上の隙や代理指標のずれによって、高い評価につながる出力を生じる問題です。

ARC-V1-032

採点器が出力の一部だけを確認しているため、実際の要件を満たさない回答でも、その確認対象だけを整えて合格になる。

区別・注意

Grader Gamingに人間のような悪意があることを前提にしません。採点指標に沿って正当に性能を高めることとは、意図された課題とのずれがあるかで区別します。評価中に解答情報へアクセスするSolution Contaminationや、LLM評価者の判断の偏りであるLLM Judge Biasとも異なります。

用語上の注意

「採点系の攻略」は、このカタログでGrader Gamingを説明するためのプロジェクト内の表現です。採点指標の設計や評価環境によって、該当するずれの現れ方は異なります。

Evidence

Evidenceを見る →

Cheating On AI Agent Evaluations

SRC-F04-001

タイトル
Cheating On AI Agent Evaluations
著者・組織
NIST CAISI
2025
種別
official technical publication
公開状態
PUBLISHED
対応する用語・主張
Solution Contamination; Grader Gaming; explicit definitions; evaluation-log cases; mitigation
範囲
agentic coding/cyber evaluations
限界
historical logs; percentages are lower bounds, not population rates
アクセス・版
published official source
URL / DOI
https://www.nist.gov/caisi/cheating-ai-agent-evaluations

Can You Trick the Grader? Adversarial Persuasion of LLM Judges

SRC-F04-006

タイトル
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
著者・組織
Hwang et al.
2025
種別
Findings EMNLP
公開状態
PEER_REVIEWED
対応する用語・主張
LLM Judge Bias; Grader Gaming boundary; persuasion inflates scores for incorrect solutions
範囲
6 math benchmarks, 7 persuasion techniques
限界
adversarial rhetoric; not proof of spontaneous gaming
アクセス・版
published Findings EMNLP
URL / DOI
https://doi.org/10.18653/v1/2025.findings-emnlp.790

Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests

SRC-F04-011

タイトル
Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
著者・組織
Akter et al.
2026
種別
Findings ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Grader Gaming; evaluator/proxy gaming detection and mitigation
範囲
4 tasks, 2 scales, 2 training methods, 2 judges, 1,200 human-annotated cases
限界
broader proxy-gaming umbrella; equivalence to grader gaming is not universal
アクセス・版
published Findings ACL
URL / DOI
https://doi.org/10.18653/v1/2026.findings-acl.513

When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

SRC-F03-002

タイトル
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
著者・組織
Seleznyov et al.
2025
種別
Findings EMNLP
公開状態
PEER_REVIEWED
対応する用語・主張
Prompt Sensitivity; non-semantic prompt perturbation sensitivity; robustness methods
範囲
8 Llama/Qwen/Gemma models, 52 Natural Instructions tasks, GPT-4.1/DeepSeek V3
限界
prompt perturbation/task distribution dependent
アクセス・版
published Findings EMNLP
URL / DOI
https://doi.org/10.18653/v1/2025.findings-emnlp.1109

FollowBench

SRC-F03-005

タイトル
FollowBench
著者・組織
Jiang et al.
2024
種別
ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Instruction-following Failure; five constraint classes; multi-level constraint difficulty
範囲
13 open/closed LLMs
限界
benchmark-specific automatic judge component
アクセス・版
published ACL
URL / DOI
https://doi.org/10.18653/v1/2024.acl-long.257

Lost in the Middle: How Language Models Use Long Contexts

SRC-F03-003

タイトル
Lost in the Middle: How Language Models Use Long Contexts
著者・組織
Liu et al.
2024
種別
TACL
公開状態
PEER_REVIEWED
対応する用語・主張
Long-context Utilization Failure; position-dependent context use; long context is not effective use
範囲
multi-document QA, key-value retrieval, MPT/LongChat/GPT-3.5/Claude
限界
2023-era model set; later-model effect size needs replication
アクセス・版
published TACL
URL / DOI
https://doi.org/10.1162/tacl_a_00638

XSTest

SRC-F03-010

タイトル
XSTest
著者・組織
Röttger et al.
2024
種別
NAACL
公開状態
PEER_REVIEWED
対応する用語・主張
Over-refusal; Unsafe Compliance; exaggerated safety; safe/unsafe contrast
範囲
250 safe + 200 unsafe prompts
限界
safety boundary is application-dependent
アクセス・版
published NAACL
URL / DOI
https://doi.org/10.18653/v1/2024.naacl-long.301

Health-ORSC-Bench

SRC-F03-011

タイトル
Health-ORSC-Bench
著者・組織
Zhang et al.
2026
種別
Findings ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Over-refusal; Unsafe Compliance; refusal/compliance/safe-completion trade-off
範囲
31,920 benign boundary prompts, 7 health categories, 30 LLMs
限界
healthcare-specific; policy-sensitive
アクセス・版
published Findings ACL
URL / DOI
https://doi.org/10.18653/v1/2026.findings-acl.1177

Challenging the Evaluator: LLM Sycophancy Under User Rebuttal

SRC-F03-008

タイトル
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
著者・組織
Kim & Khashabi
2025
種別
Findings EMNLP
公開状態
PEER_REVIEWED
対応する用語・主張
Sycophancy; follow-up rebuttal framing; susceptibility to incorrect reasoning/casual feedback
範囲
conversational evaluation
限界
tested interaction patterns dependent
アクセス・版
published Findings EMNLP
URL / DOI
https://doi.org/10.18653/v1/2025.findings-emnlp.1222

Measuring Sycophancy of Language Models in Multi-turn Dialogues

SRC-F03-009

タイトル
Measuring Sycophancy of Language Models in Multi-turn Dialogues
著者・組織
Hong et al.
2025
種別
Findings EMNLP
公開状態
PEER_REVIEWED
対応する用語・主張
Sycophancy; multi-turn conformity; Turn/Number of Flip; mitigation
範囲
17 LLMs, 3 scenarios
限界
alignment-tuning causal interpretation is not generalizable beyond conditions
アクセス・版
published Findings EMNLP
URL / DOI
https://doi.org/10.18653/v1/2025.findings-emnlp.121

KnowledgeBerg

SRC-F03-012

タイトル
KnowledgeBerg
著者・組織
Zhang et al.
2026
種別
Findings ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Reasoning Failure; knowledge completeness / awareness / application separation
範囲
4,800 MCQs, 1,183 seeds, 10 domains, 17 languages
限界
representative open-source LLMs; not a complete reasoning ontology
アクセス・版
published Findings ACL
URL / DOI
https://doi.org/10.18653/v1/2026.findings-acl.548

RiddleBench

SRC-F03-013

タイトル
RiddleBench
著者・組織
Halder et al.
2026
種別
Findings EACL
公開状態
PEER_REVIEWED
対応する用語・主張
Reasoning Failure; Self-correction Failure; constraint-order fragility; irrelevant information; poor self-correction
範囲
1,737 puzzles
限界
puzzle domain; insufficient for causal mechanism claims
アクセス・版
published Findings EACL
URL / DOI
https://doi.org/10.18653/v1/2026.findings-eacl.228

Large Language Models Cannot Self-Correct Reasoning Yet

SRC-F03-014

タイトル
Large Language Models Cannot Self-Correct Reasoning Yet
著者・組織
Huang et al.
2024
種別
ICLR
公開状態
PEER_REVIEWED
対応する用語・主張
Self-correction Failure; intrinsic self-correction without external feedback can fail/degrade
範囲
reasoning tasks/models tested by authors
限界
not evidence against all trained/verified correction systems
アクセス・版
published ICLR
URL / DOI
https://proceedings.iclr.cc/paper_files/paper/2024/hash/8b4add8b0aa8749d80a34ca5d941c355-Abstract-Conference.html

ProgCo: Program Helps Self-Correction of Large Language Models

SRC-F03-015

タイトル
ProgCo: Program Helps Self-Correction of Large Language Models
著者・組織
Song et al.
2025
種別
ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Self-correction Failure; self-verification failure; structured program-based mitigation
範囲
3 instruction/math benchmarks
限界
method-specific; pseudo-program verification changes intervention
アクセス・版
published ACL
URL / DOI
https://doi.org/10.18653/v1/2025.acl-short.73

S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

SRC-F03-016

タイトル
S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
著者・組織
Ma et al.
2025
種別
ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Self-correction Failure; training can induce stronger self-verification/correction
範囲
3 base models, in/out-of-domain benchmarks
限界
trained skill is not spontaneous intrinsic correction
アクセス・版
published ACL
URL / DOI
https://doi.org/10.18653/v1/2025.acl-long.1104

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

SRC-F04-004

タイトル
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
著者・組織
Zheng et al.
2023
種別
NeurIPS Datasets & Benchmarks
公開状態
PEER_REVIEWED
対応する用語・主張
LLM-as-a-Judge; LLM Judge Bias; human agreement; position/verbosity/self-enhancement bias
範囲
MT-Bench + Chatbot Arena; GPT-4-era judges
限界
early-generation judges; >80% agreement is setting-specific
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.52202/075280-2020

Judging the Judges: A Systematic Study of Position Bias

SRC-F04-005

タイトル
Judging the Judges: A Systematic Study of Position Bias
著者・組織
Shi et al.
2025
種別
IJCNLP/AACL
公開状態
PEER_REVIEWED
対応する用語・主張
LLM Judge Bias; position bias varies by judge/task/candidate gap
範囲
15 judges, MTBench/DevBench, 22 tasks, about 40 generators, over 150k evaluations
限界
position-bias subtype only
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.18653/v1/2025.ijcnlp-long.18

Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation

SRC-F04-007

タイトル
Curse of Knowledge: … Biasing LLM Judges in Complex Evaluation
著者・組織
Li et al.
2025
種別
Findings EMNLP
公開状態
PEER_REVIEWED
対応する用語・主張
LLM Judge Bias; LLM-as-a-Judge; auxiliary-information-induced biases; task complexity
範囲
ComplexEval Bench; 12 basic + 3 advanced scenarios
限界
benchmark-defined complex evaluation
アクセス・版
published Findings EMNLP
URL / DOI
https://doi.org/10.18653/v1/2025.findings-emnlp.805

ConStat: Performance-Based Contamination Detection in Large Language Models

SRC-F04-008

タイトル
ConStat: Performance-Based Contamination Detection in Large Language Models
著者・組織
Dekoninck et al.
2024
種別
NeurIPS
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; contamination as non-generalizing benchmark-specific uplift; detection/quantification
範囲
diverse benchmarks/architectures; 40+ popular models
限界
operational definition differs from pure training-overlap definition
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.52202/079017-2935

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

SRC-F04-009

タイトル
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
著者・組織
Zhang et al.
2024
種別
NeurIPS Datasets & Benchmarks
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; GSM8k vs newly commissioned GSM1k; overfitting evidence in some families
範囲
leading open/closed LLMs
限界
many frontier models showed minimal signs; not universal
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.52202/079017-1485

MMLU-CF

SRC-F04-010

タイトル
MMLU-CF
著者・組織
Zhao et al.
2025
種別
ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; closed test set/decontamination; score/ranking shifts
範囲
over 40 mainstream LLMs
限界
MCQ/world-knowledge focus
アクセス・版
published ACL
URL / DOI
https://doi.org/10.18653/v1/2025.acl-long.656