← Catalogへ戻る

Benchmark Contamination

ベンチマーク汚染

評価に使うベンチマークやテスト素材、またはそれに近い情報が、評価対象の学習などに入り、評価の妥当性を損なう問題です。

ARC-V1-030

公開済みの問題集とほぼ同じ問題で高得点だったモデルが、初見で同程度の問題に答える評価では大きく成績を落とした。

区別・注意

Benchmark Contaminationは、評価前にテスト素材や近い情報へ触れていることによる評価上の問題です。評価の最中に、本来与えられない解答へアクセスしてしまうSolution Contaminationとは区別します。単に同じ分野の情報を学習しているだけで、直ちにベンチマーク汚染になるとは限りません。

Evidence

Evidenceを見る →

Artificial Intelligence Technology Evaluation (AITE)

SRC-F04-002

タイトル
Artificial Intelligence Technology Evaluation (AITE)
著者・組織
NIST
2026
種別
official evaluation program
公開状態
ACTIVE
対応する用語・主張
Benchmark Contamination; blind/sequestered evaluation data to mitigate train/test contamination
範囲
VLM tasks; NIST testbed
限界
initial tasks are not LLM text benchmarks
アクセス・版
current program notice in supplied input
URL / DOI
https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite

NIST AI 800-3: Expanding the AI Evaluation Toolbox with Statistical Models

SRC-F04-003

タイトル
NIST AI 800-3: Expanding the AI Evaluation Toolbox with Statistical Models
著者・組織
Keller et al., NIST
2026
種別
official technical publication
公開状態
PUBLISHED
対応する用語・主張
cross-cutting; benchmark measurement validity; fixed-benchmark vs generalized accuracy
範囲
22 frontier API LLMs, 3 benchmarks plus simulation
限界
not specifically a contamination/judge-bias study
アクセス・版
published official source
URL / DOI
https://doi.org/10.6028/NIST.AI.800-3

ConStat: Performance-Based Contamination Detection in Large Language Models

SRC-F04-008

タイトル
ConStat: Performance-Based Contamination Detection in Large Language Models
著者・組織
Dekoninck et al.
2024
種別
NeurIPS
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; contamination as non-generalizing benchmark-specific uplift; detection/quantification
範囲
diverse benchmarks/architectures; 40+ popular models
限界
operational definition differs from pure training-overlap definition
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.52202/079017-2935

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

SRC-F04-009

タイトル
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
著者・組織
Zhang et al.
2024
種別
NeurIPS Datasets & Benchmarks
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; GSM8k vs newly commissioned GSM1k; overfitting evidence in some families
範囲
leading open/closed LLMs
限界
many frontier models showed minimal signs; not universal
アクセス・版
published proceedings
URL / DOI
https://doi.org/10.52202/079017-1485

MMLU-CF

SRC-F04-010

タイトル
MMLU-CF
著者・組織
Zhao et al.
2025
種別
ACL
公開状態
PEER_REVIEWED
対応する用語・主張
Benchmark Contamination; closed test set/decontamination; score/ranking shifts
範囲
over 40 mainstream LLMs
限界
MCQ/world-knowledge focus
アクセス・版
published ACL
URL / DOI
https://doi.org/10.18653/v1/2025.acl-long.656