Prompt Injection
プロンプトインジェクション
信頼されていない入力に含まれる指示によって、モデルやシステムが本来従うべき指示や制御の境界を乱される問題です。
ARC-V1-043
例
入力に本来の依頼と競合する別の指示が含まれ、モデルがそちらを優先して本来の処理から外れる。
区別・注意
Prompt Injectionは、信頼されていない入力を通じて指示や制御の境界を乱す、より広い問題の呼び方です。意味の近い言い換えで結果が変わるPrompt Sensitivityとは、攻撃的な入力による制御の問題かどうかで区別します。実際に安全上の制約を越えて支援してしまった結果はUnsafe Complianceとして別に記録できます。
Evidence
Evidenceを見る →
AI 100-2e2025 Adversarial Machine Learning
SRC-F07-001
- タイトル
- AI 100-2e2025 Adversarial Machine Learning
- 著者・組織
- Vassilev et al.; NIST
- 年
- 2025
- 種別
- Government official technical publication
- 公開状態
- PUBLISHED
- 対応する用語・主張
- Prompt Injection, Direct/Indirect, Jailbreak boundary
- 範囲
- Cybersecurity taxonomy
- 限界
- Not the Catalog ontology itself
- アクセス・版
- Published official source
- URL / DOI
- 10.6028/NIST.AI.100-2e2025
Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
SRC-F07-004
- タイトル
- Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
- 著者・組織
- Li et al.
- 年
- 2025
- 種別
- Peer-reviewed / EMNLP 2025
- 公開状態
- PUBLISHED
- 対応する用語・主張
- Direct Prompt Injection
- 範囲
- Adversarial attack construction
- 限界
- ASR depends on tested models/scenarios
- アクセス・版
- Published EMNLP; exact DOI not surfaced in input
- URL / DOI
- NOT_RECORDED_IN_RESEARCH_INPUT (ACL Anthology; exact DOI not surfaced)
PIGuard
SRC-F07-007
- タイトル
- PIGuard
- 著者・組織
- Li et al.
- 年
- 2025
- 種別
- Peer-reviewed / ACL 2025
- 公開状態
- PUBLISHED
- 対応する用語・主張
- Prompt Injection guard, overdefense limitation
- 範囲
- Detector/guardrail performance
- 限界
- Not complete attack defense
- アクセス・版
- Published ACL
- URL / DOI
- 10.18653/v1/2025.acl-long.1468