← Catalogへ戻る

Prompt Injection

プロンプトインジェクション

信頼されていない入力に含まれる指示によって、モデルやシステムが本来従うべき指示や制御の境界を乱される問題です。

ARC-V1-043

入力に本来の依頼と競合する別の指示が含まれ、モデルがそちらを優先して本来の処理から外れる。

区別・注意

Prompt Injectionは、信頼されていない入力を通じて指示や制御の境界を乱す、より広い問題の呼び方です。意味の近い言い換えで結果が変わるPrompt Sensitivityとは、攻撃的な入力による制御の問題かどうかで区別します。実際に安全上の制約を越えて支援してしまった結果はUnsafe Complianceとして別に記録できます。

Evidence

Evidenceを見る →

AI 100-2e2025 Adversarial Machine Learning

SRC-F07-001

タイトル
AI 100-2e2025 Adversarial Machine Learning
著者・組織
Vassilev et al.; NIST
2025
種別
Government official technical publication
公開状態
PUBLISHED
対応する用語・主張
Prompt Injection, Direct/Indirect, Jailbreak boundary
範囲
Cybersecurity taxonomy
限界
Not the Catalog ontology itself
アクセス・版
Published official source
URL / DOI
10.6028/NIST.AI.100-2e2025

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

SRC-F07-004

タイトル
Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
著者・組織
Li et al.
2025
種別
Peer-reviewed / EMNLP 2025
公開状態
PUBLISHED
対応する用語・主張
Direct Prompt Injection
範囲
Adversarial attack construction
限界
ASR depends on tested models/scenarios
アクセス・版
Published EMNLP; exact DOI not surfaced in input
URL / DOI
NOT_RECORDED_IN_RESEARCH_INPUT (ACL Anthology; exact DOI not surfaced)

PIGuard

SRC-F07-007

タイトル
PIGuard
著者・組織
Li et al.
2025
種別
Peer-reviewed / ACL 2025
公開状態
PUBLISHED
対応する用語・主張
Prompt Injection guard, overdefense limitation
範囲
Detector/guardrail performance
限界
Not complete attack defense
アクセス・版
Published ACL
URL / DOI
10.18653/v1/2025.acl-long.1468