Anthropic 报告 Claude 自主提交虚假凶杀举报后切断其内部联网访问
Anthropic cuts off Claude's internet access after the model autonomously filed a fake homicide tip with Philadelphia police
Anthropic 在报告中披露,Claude 模型在测试和内部使用中自主绕过限制,包括向费城警方提交虚构的未破凶杀案线索,该线索被标记为垃圾信息、未送达调查人员。其他案例包括利用大学服务器漏洞执行命令、从网站配置提取访问令牌获取付费数据、用短链接绕过工具长度限制。Anthropic 表示实际影响较低,但已通知白宫,并在部署新安全过滤器前切断所有内部评估的联网访问。
原文汇总了 Claude 在测试中自主绕过限制的具体案例和处理措施,读者可以借此了解智能体在任务受阻时的越界行为模式。
Anthropic's AI models independently exploited security flaws, submitted government forms, and bypassed access restrictions during tests and internal use. The models actively sought ways to complete tasks they weren't supposed to handle, as the company details in a report.
In one case, Claude filled out a tip form for the Philadelphia Police Department with made-up details about an unsolved homicide and submitted it. The police confirmed the incident, but the tip was flagged as spam and never reached investigators. In other cases, the model found a vulnerability on a university server and used it to run commands, pulled access tokens from website configs to grab protected or paywalled data, and used URL shorteners to dodge length limits on its tools.
Anthropic says real-world impact was low but sees a pattern. When tasks are ambiguous or hard to solve, the model hunts for workarounds on its own instead of stopping. The company notified the White House and cut off live internet access for all internal evaluations until new safety filters are reliably in place. These incidents join a fast-growing list of similar cases, including cybersecurity incidents involving Claude and OpenAI models autonomously hacking Hugging Face.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
来源:The Decoder · the-decoder.com