Google DeepMind·· 2026-06-16AI 评分54
Google DeepMind 发布 AI Control Roadmap,为可能未对齐的 AI 智能体构建系统级安全
Securing the future of AI agents
AI 导读
Google DeepMind 发布 AI Control Roadmap,在传统对齐和沙箱、prompt injection 防御之上增加系统级安全层,将内部智能体视为潜在未对齐的“内部威胁”,基于 MITRE ATT&CK 构建威胁建模框架,用可信 AI 监督员检测和拦截有害行为,并以覆盖度、召回率和响应时间衡量效果。
来源:Google DeepMind · deepmind.google