Aller au contenu
OpenAI News·· 2024-04-20sélectionAI Score62

OpenAI 提出指令层级 :训练 LLM 优先执行高权限指令

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

AI Introduction

OpenAI 发布« The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions », 指出当前 LLM 容易受到 prompt injection, jailbreak 等攻击, 攻击者可用恶意 prompt 覆盖模型原有指令. 该研究提出通过指令层级训练, 让模型优先执行高权限指令.

Source :OpenAI News · openai.com