Reliability
信頼性
AIシステムが、想定された条件・期間・用途のもとで、求められた機能を安定して果たせる性質です。
ARC-V1-001
例
同じ種類の問い合わせを継続して処理しているとき、ある日だけ誤答が急に増えたり、必要な処理が何度も失敗したりする。
区別・注意
信頼性は、単に「正答率が高い」ことや「毎回答えが同じ」ことだけを意味しません。正確性、頑健性、一貫性など複数の観点が関係しますが、それらのどれか一つと同義ではありません。
Evidence
Evidenceを見る →
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
SRC-F01-001
- タイトル
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- 著者・組織
- NIST
- 年
- 2023
- 種別
- official technical publication
- 公開状態
- published
- 対応する用語・主張
- Reliability, Accuracy, Robustness; trustworthiness characteristics
- 範囲
- AI全般のrisk framework。LLM固有taxonomyではない。2026年時点でrevision underway。
- アクセス・版
- 現行公開版をhistorical/technical baselineとして扱う。
- URL / DOI
- 10.6028/NIST.AI.100-1
The Language of Trustworthy AI: An In-Depth Glossary of Terms
SRC-F01-002
- タイトル
- The Language of Trustworthy AI: An In-Depth Glossary of Terms
- 著者・組織
- NIST
- 年
- 2023
- 種別
- official technical publication
- 公開状態
- published
- 対応する用語・主張
- terminology harmonization; Reliability/Accuracy周辺
- 範囲
- glossaryは共通語彙形成を目的とするが、全research communityでの唯一の定義ではない。
- アクセス・版
- 用語調整のbaseline。
- URL / DOI
- 10.6028/NIST.AI.100-3
WILDS: A Benchmark of in-the-Wild Distribution Shifts
SRC-F01-009
- タイトル
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- 著者・組織
- Koh et al.
- 年
- 2021
- 種別
- ICML original peer-reviewed research
- 公開状態
- published
- 対応する用語・主張
- Distribution Shift, Robustness
- 範囲
- general ML benchmark。LLM-onlyではない。
- アクセス・版
- published proceedings。
- URL / DOI
- https://proceedings.mlr.press/v139/koh21a.html
Evaluating Model Robustness and Stability to Dataset Shift
SRC-F01-010
- タイトル
- Evaluating Model Robustness and Stability to Dataset Shift
- 著者・組織
- Subbaswamy et al.
- 年
- 2021
- 種別
- AISTATS original peer-reviewed research
- 公開状態
- published
- 対応する用語・主張
- Robustness, Distribution Shift
- 範囲
- predictive-model setting中心。
- アクセス・版
- published proceedings。
- URL / DOI
- https://proceedings.mlr.press/v130/subbaswamy21a.html
Evaluating large language models for accuracy incentivizes hallucinations
SRC-F01-012
- タイトル
- Evaluating large language models for accuracy incentivizes hallucinations
- 著者・組織
- Kalai et al.
- 年
- 2026
- 種別
- Nature original peer-reviewed research
- 公開状態
- published
- 対応する用語・主張
- Accuracy, Abstention, Hallucination, evaluation incentives
- 範囲
- formal model + selected frontier-model case study。utilityはuse case依存。
- アクセス・版
- published Nature。
- URL / DOI
- 10.1038/s41586-026-10549-w