跳到正文
原文
Hugging Face Blog·· 2026-08-10AI 评分48

Hugging Face 提出 Offline Top-K Logits 与融合分块 KL 损失,让知识蒸馏成本大幅下降

Making Knowledge Distillation Cheap Enough to Run at Scale

AI 导读

Hugging Face 发布论文,通过缓存 teacher 的 top-100 logits 实现 offline 蒸馏,并引入 fused chunked KL 损失,避免生成全词表×序列长度的矩阵。

来源:Hugging Face Blog · huggingface.co