← Catalogへ戻る

Accuracy

正確性

予測、計算、推定、回答などが、正解・基準値・受け入れられた参照値にどの程度一致しているかを表す性質です。

ARC-V1-002

10問の計算問題に回答させたところ、基準となる正答と一致したのが8問だった。

区別・注意

正確性の測り方はタスクによって異なります。分類問題の正解率だけを指すわけではありません。また、確信度が高いかどうかや、その確信度が適切かどうかは別の観点です。

Evidence

Evidenceを見る →

Artificial Intelligence Risk Management Framework (AI RMF 1.0)

SRC-F01-001

タイトル
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
著者・組織
NIST
2023
種別
official technical publication
公開状態
published
対応する用語・主張
Reliability, Accuracy, Robustness; trustworthiness characteristics
範囲
AI全般のrisk framework。LLM固有taxonomyではない。2026年時点でrevision underway。
アクセス・版
現行公開版をhistorical/technical baselineとして扱う。
URL / DOI
10.6028/NIST.AI.100-1

On Calibration of Modern Neural Networks

SRC-F01-003

タイトル
On Calibration of Modern Neural Networks
著者・組織
Guo, Pleiss, Sun & Weinberger
2017
種別
ICML original peer-reviewed research
公開状態
published
対応する用語・主張
Calibration; confidence vs correctness; temperature scaling
範囲
主にclassification neural networks。LLM verbal confidenceへ直接一般化不可。
アクセス・版
published proceedings。
URL / DOI
https://proceedings.mlr.press/v70/guo17a.html

FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

SRC-F02-004

タイトル
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
著者・組織
Min et al.
2023
種別
EMNLP original peer-reviewed research
公開状態
published
対応する用語・主張
Factuality, Verification, claim-level support
範囲
biographies/long-form中心。reference knowledge source品質に依存。
アクセス・版
published EMNLP。
URL / DOI
10.18653/v1/2023.emnlp-main.741

Evaluating large language models for accuracy incentivizes hallucinations

SRC-F01-012

タイトル
Evaluating large language models for accuracy incentivizes hallucinations
著者・組織
Kalai et al.
2026
種別
Nature original peer-reviewed research
公開状態
published
対応する用語・主張
Accuracy, Abstention, Hallucination, evaluation incentives
範囲
formal model + selected frontier-model case study。utilityはuse case依存。
アクセス・版
published Nature。
URL / DOI
10.1038/s41586-026-10549-w