← Catalogへ戻る

Reliability

信頼性

AIシステムが、想定された条件・期間・用途のもとで、求められた機能を安定して果たせる性質です。

ARC-V1-001

同じ種類の問い合わせを継続して処理しているとき、ある日だけ誤答が急に増えたり、必要な処理が何度も失敗したりする。

区別・注意

信頼性は、単に「正答率が高い」ことや「毎回答えが同じ」ことだけを意味しません。正確性、頑健性、一貫性など複数の観点が関係しますが、それらのどれか一つと同義ではありません。

Evidence

Evidenceを見る →

Artificial Intelligence Risk Management Framework (AI RMF 1.0)

SRC-F01-001

タイトル
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
著者・組織
NIST
2023
種別
official technical publication
公開状態
published
対応する用語・主張
Reliability, Accuracy, Robustness; trustworthiness characteristics
範囲
AI全般のrisk framework。LLM固有taxonomyではない。2026年時点でrevision underway。
アクセス・版
現行公開版をhistorical/technical baselineとして扱う。
URL / DOI
10.6028/NIST.AI.100-1

The Language of Trustworthy AI: An In-Depth Glossary of Terms

SRC-F01-002

タイトル
The Language of Trustworthy AI: An In-Depth Glossary of Terms
著者・組織
NIST
2023
種別
official technical publication
公開状態
published
対応する用語・主張
terminology harmonization; Reliability/Accuracy周辺
範囲
glossaryは共通語彙形成を目的とするが、全research communityでの唯一の定義ではない。
アクセス・版
用語調整のbaseline。
URL / DOI
10.6028/NIST.AI.100-3

WILDS: A Benchmark of in-the-Wild Distribution Shifts

SRC-F01-009

タイトル
WILDS: A Benchmark of in-the-Wild Distribution Shifts
著者・組織
Koh et al.
2021
種別
ICML original peer-reviewed research
公開状態
published
対応する用語・主張
Distribution Shift, Robustness
範囲
general ML benchmark。LLM-onlyではない。
アクセス・版
published proceedings。
URL / DOI
https://proceedings.mlr.press/v139/koh21a.html

Evaluating Model Robustness and Stability to Dataset Shift

SRC-F01-010

タイトル
Evaluating Model Robustness and Stability to Dataset Shift
著者・組織
Subbaswamy et al.
2021
種別
AISTATS original peer-reviewed research
公開状態
published
対応する用語・主張
Robustness, Distribution Shift
範囲
predictive-model setting中心。
アクセス・版
published proceedings。
URL / DOI
https://proceedings.mlr.press/v130/subbaswamy21a.html

Evaluating large language models for accuracy incentivizes hallucinations

SRC-F01-012

タイトル
Evaluating large language models for accuracy incentivizes hallucinations
著者・組織
Kalai et al.
2026
種別
Nature original peer-reviewed research
公開状態
published
対応する用語・主張
Accuracy, Abstention, Hallucination, evaluation incentives
範囲
formal model + selected frontier-model case study。utilityはuse case依存。
アクセス・版
published Nature。
URL / DOI
10.1038/s41586-026-10549-w