← Catalogへ戻る

Confidence

確信度

モデルが、ある予測や回答をどの程度正しいとみなしているかを、確率、スコア、または自然言語などで外部に表したものです。

ARC-V1-004

別名: Verbalized Confidence

回答と一緒に「確信度90%」と表示したり、「かなり確信があります」と説明したりする場合、その数値や表現が確信度にあたります。

区別・注意

LLMの確信度には、内部の確率に基づくスコアと、文章として自己申告する確信表現があります。両者は同じものとは限りません。また、確信度が高いことは、その回答が実際に正しいことを保証しません。

未解明・注意点

内部確率を取得できないモデルでは、どの外部表現を確信度として使うのが最も妥当かは、まだ一意に決まっていません。

Evidence

Evidenceを見る →

Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models

SRC-F01-004

タイトル
Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models
著者・組織
Zhao et al.
2024
種別
Findings ACL original peer-reviewed research
公開状態
published
対応する用語・主張
Calibration, Confidence; prompting effects
範囲
QA tasksと検討したprompting methodsに依存。
アクセス・版
published ACL Findings。
URL / DOI
10.18653/v1/2024.findings-acl.515

MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs

SRC-F01-006

タイトル
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
著者・組織
Liu et al.
2025
種別
EMNLP original peer-reviewed research
公開状態
published
対応する用語・主張
Confidence, Calibration, verbal uncertainty
範囲
自然言語でのuncertainty表明が対象。token-level probabilitiesとは別。
アクセス・版
published EMNLP。
URL / DOI
10.18653/v1/2025.emnlp-main.1505

Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

SRC-F01-014

タイトル
Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
著者・組織
Ji et al.
2025
種別
EMNLP original peer-reviewed research
公開状態
published
対応する用語・主張
Confidence, verbal vs semantic uncertainty, hallucination
範囲
short-form answers中心。representation interventionの一般化範囲に限界。
アクセス・版
published EMNLP。
URL / DOI
10.18653/v1/2025.emnlp-main.187

Bayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models

SRC-F01-005

タイトル
Bayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models
著者・組織
Tonolini et al.
2024
種別
Findings ACL original peer-reviewed research
公開状態
published
対応する用語・主張
Uncertainty Estimation, prompt uncertainty
範囲
Natural-language classification中心、小規模labeled validation setを使う。
アクセス・版
published ACL Findings。
URL / DOI
10.18653/v1/2024.findings-acl.728