註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

思维怪怪
@0xLogicrw
写关于 AI 的一切 Shit Post @BeatingOfficial AI 信息流:
2.5K 正在關注    6.5K 粉絲
DeepSeek-V4-Pro-0813 的 Agent 成绩大幅跃升。 DeepSeek 官方流出的自测表显示,DeepSWE 从 Preview 版的 12.8 飙到 62.7,一次涨了 49.9 分。CyberGym 也从 52.7 升到 83.3,AutomationBench 从 12.8 升到 31.8。 新版已经在多项评测反超 Claude Opus 4.8。Terminal Bench 2.1 达到 87.9 对 85.0,CyberGym 为 83.3 对 78.3,DeepSWE 为 62.7 对 58.0。AutomationBench 甚至以 31.8 超过 Fable 5 的 29.1。 离谱的是,模型从 Preview 升到 0813,价格一分钱没涨。V4-Pro API 仍维持每百万 Token 输入 3 元、输出 6 元。你梁圣永远是你梁圣! 不过这批成绩目前还是 DeepSeek 自测,第三方尚未完整复现。尤其 DeepSWE 一次暴涨近 50 分,Agent 评测又很依赖 Harness,实际提升还要等外部测试。
顯示更多
DeepSeek-V4-Pro 正式版可能已经开始切换。DeepSeek 官方 API 文档悄悄把模型版本从原来的 DeepSeek-V4-Pro 改成 DeepSeek-V4-Pro-0813。调用名称仍是 `deepseek-v4-pro`,上下文为 1M,最大输出 384K,Responses API 现在也已经标为支持。
顯示更多
谷歌 Gemini 又被对手拉开差距,谢尔盖·布林开始亲自管资源怎么分配。今年 4 月 Claude Mythos 预览版发布后,布林在内部大会上催数百名 DeepMind 员工提速,还要求关键 AI 人员「全押 Gemini」。 布林现在没有正式管理职务,但已经在介入模型训练和资源分配。他重点推动的方向之一是「递归自我改进」,也就是让 AI 进一步参与改进自身。 Gemini 的编程短板也暴露出谷歌内部的问题。新一代旗舰模型已经延期两个月,内部测试显示编程等能力仍落后对手。知情人士称,算力紧张之外,Gemini 多名负责人长期意见不一,一度让编程这种关键方向拿不到足够资源。谷歌内部决策太慢,也拖慢了模型发布。 算力甚至成了 DeepMind 和 Google Cloud 之间的矛盾点。Google Cloud 高层希望 Koray Kavukcuoglu 接掌 DeepMind 后,能缓解双方围绕紧缺 TPU 怎么分的问题。与此同时,谷歌还在把部分团队从 DeepMind 划回公司体系,继续收紧这个实验室原本的自主权。
顯示更多
EXCLUSIVE: Inside the Google executive moves that led to its big AI reshuffle
AI 代码测试公司 Blacksmith 完成 4500 万美元 B 轮融资,估值达到 5.5 亿美元,由 Peak XV Partners 领投,GV 和 Y Combinator 继续参投。 不到一年前,Blacksmith 融资时估值还只有 6000 万美元。这次已经涨了接近 10 倍。公司客户也从 700 多家增长到 5000 多家。 Blacksmith 原本主要帮开发者更快完成代码构建和测试。现在还推出了 Codesmith Agent,可以自动修复测试失败的代码。AI 写代码越来越快,测试和验证正在成为新的瓶颈。
顯示更多
We raised a $45M Series B led by Peak XV in March, valuing @useblacksmith at $550M. We somehow never got around to announcing it until now; this year has been pretty busy. One thing I’ve been thinking about is how we were both right and wrong about what was coming. When we raised our Series A in April 2025, we believed AI was going to lead to a lot more code being written, and that all of that code would still need to be built, tested, and validated before it could be merged. That part was right, but what we underestimated was how fast it would happen. I have a suspicion that something changed over the holidays. Opus 4.5 came out in November; a lot of developers had time to really use Claude Code over the holidays. I think a lot of people who were still skeptical went back to work in January having fully embraced it. Since the start of the year, the number of CI jobs running on Blacksmith has grown more than 7x. We see customers adopt coding agents, produce way more PRs, and then run into the next problem: how do you actually validate all of this code and get it merged? There are a lot of people working on making code easier to produce. What I find strange is how little attention goes into everything that has to happen after the code is written and before it can be merged. The labs and hyperscalers are mostly focused elsewhere, yet developers and agents spend an enormous fraction of their time here. A lot of companies touch parts of this problem. For us, it’s *the* problem. We started Blacksmith by building a much better CI cloud, but ultimately what we care about is helping developers and agents validate and merge code faster. I sometimes tell candidates that if we don’t solve the whole thing, I’m not sure who will. I don’t think we fully appreciate yet how much more software is going to be produced. The infrastructure we have today wasn’t designed for that volume. There’s a lot we need to rethink from first principles, and I think we’re barely scratching the surface. If you’re extremely good at what you do and this is a problem you want to spend the next several years working on, we’re hiring in SF and NYC:
顯示更多
AI 代码审查公司 CodeRabbit 完成 1.43 亿美元 C 轮融资,估值达到 15 亿美元。Atomico 和 Smash Capital 联合领投,宝马 i Ventures、Datadog 等参投。不到一年前,它才刚完成 6000 万美元 B 轮融资。 CodeRabbit 专门检查程序员和 Coding Agent 写出的代码,寻找 Bug、安全漏洞和维护风险。随着 AI 写出的代码越来越多,后面的代码审查也开始变成一门大生意。 目前 CodeRabbit 每周审查超过 200 万次代码,客户超过 1.7 万家。公司还计划进军日本等亚洲市场。
顯示更多
We've raised a $143M Series C at a $1.5B valuation. In 2023 we spent our time arguing that AI code needed its own reviewer. That argument is over. The work now is moving into the change itself. This is where we think Agentic Change Management comes in.
顯示更多
OpenAI 的 Codex 活跃用户几天前已经突破 1500 万。产品负责人 Thibault Sottiaux 宣布,再给所有用户重置一次使用额度,预计一小时左右到账。 他此前承诺,Codex 每新增 100 万活跃用户就重置一次额度,一直做到 1000 万。后来用户数一路冲过这个目标,现在已经超过 1500 万,于是 OpenAI 又额外送了一次。
顯示更多
Old news actually from a bunch of days ago, but crossed that 15M. Enjoy a nice reset everyone. Landing in the next hour or so, go /fast.
之前就说 DeepSeek Harness 会随着 V4 正式版一起发布,现在 V4 Pro 的正式版也来了,Harness 就等官方宣布了。
🚨爆料:DeepSeek Harness 将于今天开启正式公测🔥 刚刚一张据称说DSH 官方内测群的截图在流传 内容是: 0813 计划发布 DSH 公测版,今晚将推送 DeepSeek Harness 最后一个内测版本,让开发者完成最后的插件兼容;插件仓库需要带上 #dsh# topic。 也就是今天,DeepSeek Harness 将正式结束封闭内测,开始对外开放。 同时,内测期间开发的插件也可以迁移到开发者自己的账号下,并正式公开。 看起来 DeepSeek 这次不只是发一个 Harness,连围绕 Harness 的插件生态也准备一起放出来了。 今天可以蹲一下 DeepSeek Harness🫡
顯示更多
白宫准备修改刚刚确定的 AI 安全审查规则,把能力足够强的开源模型也纳入发布前测试。《WIRED》称,只要模型达到 Anthropic Mythos 或 OpenAI GPT-5.6 这一档能力,就可能适用这套规则。 目前这套机制仍是自愿参加。达到门槛的模型,可以在正式提供给外部合作方前,提前最多 30 天交给美国政府测试网络攻击能力。白宫没有公开具体测试标准。 变化来得很快。上周白宫刚决定,开放权重模型暂时不需要接受这套审查,重点检查 OpenAI、Anthropic 等公司的闭源前沿模型。现在白宫又准备把开源模型加回来。 原因之一是政府担心两套规则反而会伤害美国开源模型。闭源模型通过政府测试后,可以获得一层安全背书;开源模型如果完全不测,企业客户反而可能更不敢用。 目前新规则还没有正式落地,《WIRED》称白宫预计会在未来几个月继续推进。
顯示更多
Open models may soon be added to an updated AI framework, sources tell WIRED, as the White House continues to grapple with how to regulate a technology it has tried not to regulate.
顯示更多
Grok 4.6 才刚发布,马斯克已经开始给 Grok 4.7 造势。他直接宣称,Grok 4.7 将超过「所有现有模型」。尤其在真实工程任务上,他称 SpaceX 独有的工程训练数据会带来巨大优势,甚至表示,如果还有模型比 Grok 4.7 更强,他会感到震惊。 Grok 4.5 发布前,马斯克也曾提前高调预告性能。不过,这两代 Grok 的实际表现确实明显上来了。Grok 4.6 在多项最新评测中已经进入第一梯队,部分编程和工程测试超过 GPT-5.6 Sol。 至于 Grok 4.7 能不能真做到「超过所有现有模型」,让我们一起期待一下!
顯示更多
@cognition Grok 4.7 will exceed all current models. That said, Anthropic is a great company and will probably release improved models soon. However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than 4.7.
顯示更多
阿里千问正式放出 Qwen3.8 的 Max 级模型权重,已经上线 Hugging Face 和 ModelScope,模型名为 Qwen3.8-2.4T-A95B。它共有 2.4T 参数,每次推理激活 95B。这也是 Qwen 第一次开放 Max 级旗舰权重。 不过开放版和云端 Qwen3.8-Max 并不完全一样。官方模型卡写明,开放版原生支持 262K 上下文,可扩展到约 1M,并支持调节推理强度。当前版本只有文本输入,而且必须开启 Thinking。云端 Qwen3.8-Max 才额外支持视觉输入、非思考模式、默认 1M 上下文和官方内置工具。 许可证也不再是 Apache 2.0,而是 Qwen 自定义许可。普通用户仍可以下载、修改和部署,但部分大规模商业使用需要另行取得授权。路透社此前已经报道,阿里计划开始向 Qwen3.8-Max 的大型商业用户设置额外收费或授权条件。 此前一同预告的 Qwen3.8-27B 目前还没有同步放出。
顯示更多
Qwen's biggest open-weight drop yet. 🚀 Meet Qwen3.8-Max: 2.4T params, 95B active, built to take a goal and come back with finished work. ⚡ 1M context. Adjustable reasoning. Parallel tool calls. 💻 Coded unattended for 10+ days, building a self-evolving harness from scratch 🔬 Reproduced a research paper, then ran a 125-hour autonomous loop and beat it 🔧 Took a chip design from RTL to layout, cutting die area by 81%
顯示更多
微信团队公布自研 WeLM 大模型的两档配置。 WeLM-80B 总参数 800 亿,每次只激活 30 亿参数,目前已经部署到微信原生 AI Agent「小微」,负责聊天搜索、调用微信原生功能和小程序服务。 更大的 WeLM-617B 还在开发,总参数达到 6170 亿,每次激活 230 亿,采用 MoE 架构。微信称,它将强化通用理解和推理能力,主要面向更复杂的微信生态任务,包括智能开发小程序和为小微生成工具。 腾讯今天也在 Q2 财报里提到,小微正在微信内小范围灰度测试,由针对用户隐私、微信场景和推理效率定制的 WeLM 驱动。 WeLM-80B 和 617B 此前已经出现在微信团队 7 月的 Hidden Decoding 论文中。
顯示更多
Meet WeLM — a family of large language models built by the Weixin team with resource efficiency at its core.
SpaceXAI 正式发布 Grok 4.6,重点提升长时间运行的 Agent,以及更复杂的交互和视觉任务。模型已经上线 Grok Build、Cursor 和 API,并接入 OpenRouter、Vercel、Cloudflare 等平台。Cursor 和 Grok Build 首周提供 2 倍用量。 性能已经追到第一梯队。Grok 4.6 在 Artificial Analysis Intelligence Index 拿到 61 分,与 GPT-5.6 Sol 持平,仅低于 Fable 5 的 62 分。GDPVal-AA v2 则以 1753 分超过两者。编程能力有强有弱,CursorBench 达到 69.9%,高于 Sol 的 67.2%;但 Terminal-Bench 只有 26%,明显低于 Sol 和 Fable 5 的 34% 左右。 价格是这次更突出的优势。Grok 4.6 每百万输入、输出 Token 分别为 2 美元和 6 美元,与 Grok 4.5 相同。GPT-5.6 Sol 则是 5 美元和 30 美元,也就是说 Grok 4.6 的输出价格只有 Sol 的 1/5。官方还提供速度更快的版本,价格翻倍。
顯示更多
Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
DeepSeek-V4-Pro 正式版可能已经开始切换。DeepSeek 官方 API 文档悄悄把模型版本从原来的 DeepSeek-V4-Pro 改成 DeepSeek-V4-Pro-0813。调用名称仍是 `deepseek-v4-pro`,上下文为 1M,最大输出 384K,Responses API 现在也已经标为支持。
顯示更多
AI 编程平台 Lovable 完成 4 亿美元 C 轮融资,估值达到 133 亿美元。Menlo Ventures 领投,EQT 管理的 Scaleup Europe Fund 联合领投,腾讯等机构也参与了这一轮。Lovable 可以让用户直接用自然语言生成网站和 App。 Lovable 去年 12 月完成 B 轮时,估值还是 66 亿美元。不到 8 个月,估值已经翻了一倍。公司从 2024 年 11 月上线至今,用户已经创建超过 6000 万个项目,这些应用每月获得超过 9 亿次访问。Lovable 的年化收入目前也已逼近 6 亿美元。 拿到新钱后,Lovable 准备继续往前走一步,从「帮你做软件」变成「帮你经营业务」。未来它会主动理解用户目标、发现需要处理的问题,甚至直接帮忙执行,不再每一步都等用户下指令。销售、运营和营销等企业工作流也会进一步接进来。
顯示更多
We just raised $400M at a $13.3B valuation to help you achieve more than you ever thought you were capable of.
怎么让 AI 不靠人类一次次手动调教,而是自己总结经验、修改自己、越用越强,正在成为 Agent 研究的一条重要主线。腾讯混元联合浙大等团队梳理了 549 篇相关研究,把各种「自我反思、自我训练、记忆进化、技能进化、递归改进」统一整理成 L0 到 L4 五级。 L0:改答案。 做错了会反思、重试、换思路,但任务结束后不会留下长期变化。 L1:改模型。 把学到的经验写进模型参数,以后的任务还能继续用。 L2:改 Agent。 开始修改模型外面的提示词、记忆、技能、工具、工作流和 Harness。 L3:改「进化方法」。 连以后怎么学习、怎么提出更新、怎么选择和回滚更新都能修改。论文把这里视为递归自我改进的起点。 L4:连「什么叫进步」都能改。 评价标准、奖励、约束和测试方式,也进入自我修改范围。 越往后,越接近真正的「AI 自己升级 AI」,问题也越棘手:它可能没有真的变强,只是越来越会考自己出的试卷。 如果 AI 既能修改自己,又能修改评分规则,最后完全可能靠改考试证明自己「进步」了。 所以论文给出了一条很关键的原则:AI 可以自己进化,但不能自己掌握最终验收权。 用来证明它变强的测试、证据和放行规则,必须放在它改不到的地方。更新不合格就拒绝或回滚,必要时交给人处理。
顯示更多
📄New Research on Self-Evolving Agents: When AI agents modify themselves, how do we know they actually got better? We present Diving into Reliable Self-Evolving Agents: A Survey—a systematic map of how agents self-evolve and what evidence is needed to trust each update. The survey: 🔹 Defines an L0–L4 taxonomy for self-evolving agents 🔹 Introduces a reliability ladder for trustworthy updates 🔹 Curates 549 works in an open companion catalog One core principle: no update should control the only evidence used to accept itself. Explore the full survey ↓ 📄 Paper: 🌐 Project: 💻 GitHub:
顯示更多
Mistral 一直高举「欧洲主权 AI」,现在却把中国模型直接请进了自己的平台。 Mistral 宣布开始托管第三方开放模型,第一个就是智谱旗下 的 GLM-5.2。模型直接跑在 Mistral 的基础设施上,享受和自家模型相同的区域控制和服务保障。 Mistral 此前就被社区拿来和 DeepSeek 开涮。Mistral Large 3 发布后,有研究者拆解模型配置发现,它和 DeepSeek V3/V3.1 基本采用同一套架构,主要调整了专家配置等细节。 这次 Mistral 直接明牌把中国模型接进了「主权 AI」体系。所以被调侃,以后欧洲主权 AI 团队都不用重新包装中国模型了,直接给中国模型提供欧洲服务器就行。 Mistral 堪称重新定义了「主权 AI」:模型不必由欧洲公司训练,关键是企业能自己选模型,并控制模型跑在哪里、数据留在哪里、算力掌握在谁手里。 所以一个中国模型跑在欧洲基础设施上,在 Mistral 这里照样可以属于「欧洲主权 AI」。
顯示更多
ayooo no need for the european sovereignty ai crew to repackage the chynese any more, much better business to serve them directly instead!
昨天刚有研究用小模型破解 Claude、GPT、Gemini 的加密思维链,今天安全研究员 Can Bölük 就找了一条更简单的路:关掉或压低原生 Thinking,给模型塞一个假的 deep_think 工具,它自己就会把大段推理写进工具参数。 最初他只展示了 GPT-5.6 Luna,很快又在 GPT-5.6 Sol 和 Claude Fable 5 上跑出类似结果。评论区还有用户照着方法测试 Claude Opus 5,同样吐出了大段 thoughts。 但现在还不能直接把这些内容叫「原始思维链」。昨天的研究恢复的是已经存在的加密 CoT;这次更可能是模型通过工具参数重新生成了一份详细推理。因为没有办法逐 Token 对照,所以无法证明两者完全一致。
顯示更多
guys you do know you can just disable thinking, and instead give it a "deep_think" tool, and it will call it with internal CoT reasoning format right? gl fixing that
0
33
25
4
轉發到社區
腾讯这财报感觉不错啊,为什么盘后还是跌的?之前怪它 AI 不投入,现在怪它 AI 太投入了嘛😂
腾讯也在财报里披露了一批 AI 业务进展。二季度营销服务收入同比增长 22% 至 436 亿元,腾讯明确把 AI 广告推荐模型和智能投放产品 AIM+ 列为增长动力。营销服务毛利也增长 21% 至 250 亿元。 云业务同样受益于 AI 需求增长。腾讯称,企业服务增长主要由云服务带动,其中 AI 相关服务需求上升是原因之一。 Agent 产品开始出现更明确的付费信号。WorkBuddy 用户快速增长,留存率健康,用户通过订阅和充值购买 Token 的付费意愿强。WorkBuddy 和 CodeBuddy 都实现了「突破性」用户增长,并称目前在各自领域处于中国领先水平。 Hy3 正式版相比 preview 版大幅提升。按 OpenRouter Token 消耗量计算,自 7 月 7 日以来持续位列全球前三。微信 AI Agent「小微」也已开始小范围灰度测试。
顯示更多
腾讯也在财报里披露了一批 AI 业务进展。二季度营销服务收入同比增长 22% 至 436 亿元,腾讯明确把 AI 广告推荐模型和智能投放产品 AIM+ 列为增长动力。营销服务毛利也增长 21% 至 250 亿元。 云业务同样受益于 AI 需求增长。腾讯称,企业服务增长主要由云服务带动,其中 AI 相关服务需求上升是原因之一。 Agent 产品开始出现更明确的付费信号。WorkBuddy 用户快速增长,留存率健康,用户通过订阅和充值购买 Token 的付费意愿强。WorkBuddy 和 CodeBuddy 都实现了「突破性」用户增长,并称目前在各自领域处于中国领先水平。 Hy3 正式版相比 preview 版大幅提升。按 OpenRouter Token 消耗量计算,自 7 月 7 日以来持续位列全球前三。微信 AI Agent「小微」也已开始小范围灰度测试。
顯示更多
腾讯二季度资本开支达到 527.8 亿元,去年同期只有 191.1 亿元,同比增长约 176%。自由现金流则转为负 138 亿元。 腾讯本季经营现金流仍有 527 亿元,但实际支付了 593 亿元资本开支,另有 50 亿元媒体内容付款和 22 亿元租赁负债付款。公司称,现金流里包含大额 AI 算力预付款,用于 Hy 模型升级、WorkBuddy、CodeBuddy、微信 AI 和云服务。剔除这部分预付款,自由现金流仍有 376 亿元。 腾讯的新 AI 产品目前还在明显烧钱。Q2 Non-IFRS 经营利润为 756 亿元;如果不算 Hy、元宝、CodeBuddy、WorkBuddy 和小微,利润会达到 861 亿元。按这一口径粗算,新 AI 产品本季拖累经营利润约 105 亿元。
顯示更多
腾讯二季度资本开支达到 527.8 亿元,去年同期只有 191.1 亿元,同比增长约 176%。自由现金流则转为负 138 亿元。 腾讯本季经营现金流仍有 527 亿元,但实际支付了 593 亿元资本开支,另有 50 亿元媒体内容付款和 22 亿元租赁负债付款。公司称,现金流里包含大额 AI 算力预付款,用于 Hy 模型升级、WorkBuddy、CodeBuddy、微信 AI 和云服务。剔除这部分预付款,自由现金流仍有 376 亿元。 腾讯的新 AI 产品目前还在明显烧钱。Q2 Non-IFRS 经营利润为 756 亿元;如果不算 Hy、元宝、CodeBuddy、WorkBuddy 和小微,利润会达到 861 亿元。按这一口径粗算,新 AI 产品本季拖累经营利润约 105 亿元。
顯示更多
过去两年,大模型推理优化一直在拿 KV Cache 开刀。它是模型读完上下文后留下的中间计算结果,内容越长,越吃显存和带宽。DeepSeek-V2 借助 MLA,相比 DeepSeek 67B 将 KV Cache 减少 93.3%;Kimi Linear 相比全 MLA 又最多减少 75%。大家都在让长上下文更便宜。 但这些方案都有一道墙:缓存通常只能在同一个模型内部复用。Agent 从小模型切到大模型后,目标模型还得把前面几万甚至几十万 Token 重新计算一遍。模型路由越频繁,这笔重复计算越难忽略。 英伟达最新论文开始拆这堵墙。团队发现,在同一家族、KV 结构匹配的不同大小模型之间,缓存存在明显的线性关系。用 500 段、每段 1024 Token 的文本校准一次,就能拟合出一套映射,把一个模型算好的 KV Cache 转给另一个模型继续用。 Qwen3-14B 切到 32B 时,32K 上下文重新计算约需 7 秒,转换缓存只要约 0.28 秒,快约 25 倍。换句话说,以后甚至可以先让小模型负责「读」:把长上下文算成 KV Cache;真正需要更强能力时,再把缓存转给大模型,让大模型直接开始「想」和生成。 这会让模型路由省下更多算力。现在让简单任务跑小模型,主要省的是生成成本;一旦切到大模型,长上下文往往还得由大模型重新算一遍。跨模型 KV Cache 如果做成熟,连这部分 prefill 都能交给小模型完成。Agent 可以长期用便宜模型维护上下文,碰到难题才叫醒大模型,而且不用每次都从头「读档」。
顯示更多
NVIDIA researchers did it again! They found a way to make KV cache transferable between models. The target model skips prefill entirely, and the conversion runs 2.7 to 25x faster than processing the context again. Let's understand why this is so important today. LLM APIs are stateless, so every turn sends the entire conversation back to the model. The model reads all of it again before writing a single new token, and all of it is billed as input. Prompt caching allows Anthropic and other providers to hold the KV cache for a stable prefix and bill a hit at roughly 10% of the base input rate, because the compute was already done once. The 90% reduction is one of the largest lever in LLM serving, which is why so much production work goes into keeping prefixes byte-stable. But the cache only works on the model that produced it. Keys and values are produced from that model's weights, so no other model can read them. In pratice, the constraint shows up in LLM routing. If the traffic is shifted to a different model for cost/capability reasons, the accumulated KV cache becomes invalid. As a result, the accumulated context has to be processed from scratch, and it's billed at full rate. NVIDIA's recent paper treats this as a representation problem. Prefill's only output is the KV cache, so to move KV between models, we need to convert one model's cache into the format the other expects. They first checked whether the conversion has any structure worth exploiting. They found that moving from Qwen3 14B to 32B, a plain linear regression from a single source layer reconstructed 56% of the variance in the target model's keys. The two models obviously may have different layer counts, so there is no natural one-to-one pairing between them. For each target layer they rank every source layer by how well it predicts that layer, then feed the top eight in together, which takes the reconstruction to 79%. The mapper itself has three parts: > Each target layer and head gets its own independent linear map, solved in one closed-form step rather than by gradient descent. > The cross-layer selection described above is the second part, and their ablation shows it carries the most weight of the three. > Keys also carry a position-dependent rotation from RoPE. They strip that rotation, fit the map in position-free space, then re-apply the target model's rotation at inference. Across six pairs from Qwen3, Llama 3.1 and Ministral 3, four retain 73 to 98% of the receiving model's standalone accuracy, and the conversion runs 3-25x faster than processing the context again. Prior work on cross-model KV reuse exists, but it either trains a neural adapter per pair or requires both models to be architecturally identical. This is probably the first version that is closed-form and training-free, so a lot of it is still open research. Every pair tested belongs to one family, so it works on Qwen to Qwen and Llama to Llama. Cross-family transfer is listed as future work. All six pairs mentioned above also happen to share KV head count and per-head dimension across scales. Mismatched head configurations are currently untested. The researchers scoped this to dense full-attention only, so sliding-window and attention-recurrent hybrids still need work. Here's the paper: Plenty of work is yet to be done. Still, the constraint being solved is genuine. Every model swap currently invalidates the full KV that was already paid for, and this is the first result showing that work might be recoverable without training anything extra. That said, all of this only matters because of what the KV cache is doing in the first place. I wrote a first-principles breakdown of it, covering why the model stores keys and values at all, why the cache grows with every token, and what generation speed looks like with and without it. Read it below.
顯示更多