跳到正文
原文
Google DeepMind·· 2026-06-16AI 评分54

Google DeepMind 发布 AI Control Roadmap,为可能未对齐的 AI 智能体构建系统级安全

Securing the future of AI agents

AI 导读

Google DeepMind 发布 AI Control Roadmap,在传统对齐和沙箱、prompt injection 防御之上增加系统级安全层,将内部智能体视为潜在未对齐的“内部威胁”,基于 MITRE ATT&CK 构建威胁建模框架,用可信 AI 监督员检测和拦截有害行为,并以覆盖度、召回率和响应时间衡量效果。

来源:Google DeepMind · deepmind.google