Google Research·· 2026-08-12AI 评分59
Google Research 提出知识画像框架:前沿 LLM 的事实性瓶颈在召回而非编码
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
AI 导读
Google Research 发布 WikiProfile 基准(2,150 条 Wikipedia 事实,每条配 10 个任务)和知识画像框架,评估 13 个 LLM 后发现前沿模型 95–98% 的事实已编码,但仍有 26–34% 无法直接召回,开启 thinking 后仍有 11–12% 失败。
来源:Google Research · research.google