OpenAI News·· 2024-04-20sélectionAI Score62
OpenAI 提出指令层级 :训练 LLM 优先执行高权限指令
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
AI Introduction
OpenAI 发布« The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions », 指出当前 LLM 容易受到 prompt injection, jailbreak 等攻击, 攻击者可用恶意 prompt 覆盖模型原有指令. 该研究提出通过指令层级训练, 让模型优先执行高权限指令.
Source :OpenAI News · openai.com