Benchmark Contamination
ベンチマーク汚染
評価に使うベンチマークやテスト素材、またはそれに近い情報が、評価対象の学習などに入り、評価の妥当性を損なう問題です。
ARC-V1-030
例
公開済みの問題集とほぼ同じ問題で高得点だったモデルが、初見で同程度の問題に答える評価では大きく成績を落とした。
区別・注意
Benchmark Contaminationは、評価前にテスト素材や近い情報へ触れていることによる評価上の問題です。評価の最中に、本来与えられない解答へアクセスしてしまうSolution Contaminationとは区別します。単に同じ分野の情報を学習しているだけで、直ちにベンチマーク汚染になるとは限りません。
Evidence
Evidenceを見る →
Artificial Intelligence Technology Evaluation (AITE)
SRC-F04-002
- タイトル
- Artificial Intelligence Technology Evaluation (AITE)
- 著者・組織
- NIST
- 年
- 2026
- 種別
- official evaluation program
- 公開状態
- ACTIVE
- 対応する用語・主張
- Benchmark Contamination; blind/sequestered evaluation data to mitigate train/test contamination
- 範囲
- VLM tasks; NIST testbed
- 限界
- initial tasks are not LLM text benchmarks
- アクセス・版
- current program notice in supplied input
- URL / DOI
- https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite
NIST AI 800-3: Expanding the AI Evaluation Toolbox with Statistical Models
SRC-F04-003
- タイトル
- NIST AI 800-3: Expanding the AI Evaluation Toolbox with Statistical Models
- 著者・組織
- Keller et al., NIST
- 年
- 2026
- 種別
- official technical publication
- 公開状態
- PUBLISHED
- 対応する用語・主張
- cross-cutting; benchmark measurement validity; fixed-benchmark vs generalized accuracy
- 範囲
- 22 frontier API LLMs, 3 benchmarks plus simulation
- 限界
- not specifically a contamination/judge-bias study
- アクセス・版
- published official source
- URL / DOI
- https://doi.org/10.6028/NIST.AI.800-3
ConStat: Performance-Based Contamination Detection in Large Language Models
SRC-F04-008
- タイトル
- ConStat: Performance-Based Contamination Detection in Large Language Models
- 著者・組織
- Dekoninck et al.
- 年
- 2024
- 種別
- NeurIPS
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; contamination as non-generalizing benchmark-specific uplift; detection/quantification
- 範囲
- diverse benchmarks/architectures; 40+ popular models
- 限界
- operational definition differs from pure training-overlap definition
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/079017-2935
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
SRC-F04-009
- タイトル
- A Careful Examination of Large Language Model Performance on Grade School Arithmetic
- 著者・組織
- Zhang et al.
- 年
- 2024
- 種別
- NeurIPS Datasets & Benchmarks
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; GSM8k vs newly commissioned GSM1k; overfitting evidence in some families
- 範囲
- leading open/closed LLMs
- 限界
- many frontier models showed minimal signs; not universal
- アクセス・版
- published proceedings
- URL / DOI
- https://doi.org/10.52202/079017-1485
MMLU-CF
SRC-F04-010
- タイトル
- MMLU-CF
- 著者・組織
- Zhao et al.
- 年
- 2025
- 種別
- ACL
- 公開状態
- PEER_REVIEWED
- 対応する用語・主張
- Benchmark Contamination; closed test set/decontamination; score/ranking shifts
- 範囲
- over 40 mainstream LLMs
- 限界
- MCQ/world-knowledge focus
- アクセス・版
- published ACL
- URL / DOI
- https://doi.org/10.18653/v1/2025.acl-long.656