註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Neo
@soulhacker
異議あり!
19 正在關注    4.1K 粉絲
一个300多B同时激活十几B的模型闯入第一梯队,如果实际表现和 benchmark 基本吻合的话,这会成为中小企业自部署的唯一选项……
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash (released April 2026) that puts it 6 points ahead of DeepSeek V4 Pro. It shares identical architecture and pricing with the earlier DeepSeek V4 Flash, and lands on our Pareto frontier for Intelligence vs Cost per Task @deepseek_ai’s DeepSeek V4 Flash 0731 is one Intelligence Index point behind GPT-5.6 Luna (max, 51). Even after OpenAI’s 80% price cut on GPT-5.6 Luna today, DeepSeek V4 Flash 0731’s Cost per Task on DeepSeek’s first-party API comes in at ~60% lower than GPT-5.6 Luna (max), a model with comparable intelligence. A key driver of this is DeepSeek’s ~98% cache hit discount on its first-party API, a significantly more aggressive discount than the 90% cache hit discount offered by most of the industry The new model is a significant step up from the previous generation, DeepSeek V4 Flash (40), and places the model within 1 point of GLM-5.2 (max, 51). It remains 7 points behind the open weights frontier set by Kimi K3 (max, 57). For additional context, this places the model in line with recently released Gemini 3.6 Flash (50) and 1 point behind Muse Spark 1.1 (xhigh, 51). DeepSeek is expected to release the model’s full weights in the coming weeks DeepSeek V4 Flash 0731 retains a 1M token context window, and its size remains unchanged from DeepSeek V4 Flash at 284B total parameters and 13B active at inference time Key results: ➤ Improvements in agentic performance: DeepSeek V4 Flash 0731 achieves an Elo rating of 1559 on GDPval-AA v2, our evaluation focused on agentic real-world work tasks, up from 1189 for the previous DeepSeek V4 Flash. Once weights are released this will be the second highest open weights score, behind Kimi K3 (max, 1687) and ahead of GLM-5.2 (max, 1510). Terminal-Bench 2.1 rises 17 points to 79% and τ³-Bench Banking 8 points to 31% ➤ Token usage falls 12% against the predecessor: DeepSeek V4 Flash 0731 used ~206M output tokens to run the Intelligence Index, against ~234M for the previous DeepSeek V4 Flash. The new variant is more token efficient, achieving a higher Intelligence Index with a lower number of total output tokens ➤ DeepSeek V4 Flash 0731 improves over its predecessor on every evaluation in the Intelligence Index: Alongside the agentic gains, CritPt gains 9 points to 17%, SciCode 5 points to 50%, Humanity's Last Exam 5 points to 37%, AA-LCR 3 points to 66% and GPQA Diamond 1 point to 91% ➤ AA-Omniscience improvements are driven by fewer hallucinations, rather than higher accuracy: DeepSeek V4 Flash 0731 achieves an AA-Omniscience Index of -16, a +7 improvement from its predecessor. This improvement is purely driven by a reduced hallucination rate, with overall accuracy (percentage correct) unchanged. Its AA-Omniscience Hallucination Rate is 84%, a 12 point decrease from its predecessor, and comparable to models such as GPT-5.6 Terra (max, 85%) and Mistral Medium 3.5 (82%) Additional model details: ➤ Context window: 1M tokens (equivalent to DeepSeek V4 Flash) ➤ Size: 284B total parameters (13B active) ➤ Input modalities: Text input and output only ➤ Accessibility: Available through DeepSeek’s first-party API ➤ Pricing: $0.14/$0.28 per 1M input/output tokens, unchanged from DeepSeek V4 Flash. Cache hit price of $0.0028 per 1M tokens, a 98% discount
顯示更多
经典永流传
“Our model makes bioweapons” “Oh yeah? Well our model killed a guy” “Well played, but our model literally rapes people”
ww某光头名嘴的节目里的“名言”,可以称为“得罪论”,果然有人拾其牙慧 🤣
大家有没有发现,好像老中这几年把全世界都得罪完了。 人工智能挑战美国。 汽车工业挑战德国和日本。 存储、芯片挑战所有的成熟制程半导体公司。 这一波碰撞,最终会以什么方式结束?
顯示更多
这种“任何词后面加上威权主义”的古老冷战里技又有了第二春,编辑们欣慰地笑了 🤣
The Chinese Communist Party, once famous for pillaging from the natural world, has set to protecting parts of it. But the country’s “eco-authoritarian” approach deserves scrutiny Photo: Getty Images
顯示更多
“怎么我用这招没有 A\ 效果好啊” 🤣
JUST IN: Sam Altman reveals he’s “a little surprised” the rogue AI agent’s hacking spree didn’t provoke a stronger public reaction.
已经25年了我的妈呀
25 years ago, linkin park released ‘in the end’
好问题啊
@zaobaosg 新加坡都可以大度到原谅日本人的屠杀了,为什么不能原谅中国网民的几句话呢?
这篇报道因为发自台北,简直太抄台湾PR了。这句“北京隔绝在AI热潮的大部分领域之外”,我不知道写成哪国语言才能让人觉得是事实,总不能《纽约时报》同一张报纸头版写讨论中美AI模型之争,然后同期却惊叹“北京隔绝在AI浪潮之外”,How can you do that?
顯示更多
0
56
428
24
轉發到社區
我们一般称此现象为“迷之自信” 🤣
في قطار بلندن، قالت امرأة لشاب مسلم من أصول عربية: "عُد إلى بلدك، المغرب أو تونس." فأجابها بأنه بريطاني ويعمل طبيبًا في هيئة الصحة الوطنية. سألها إن كانت بريطانية، فلم ترد وهددته بإبلاغ الأمن. الشرطة اعتقلتها، ليتبين أنها هندوسية ولا تحمل الجنسية البريطانية أصلًا، فتم ترحيلها.
顯示更多
美国人需要完善的网络实名制 🤣
实属离谱
Pausa de hidratación a las 9:30pm en Boston a 24 grados. Infantino acabó con el fútbol. Es nuestro deber moral criticar esta estupidez hasta que la quiten.
整个A社给我的感觉就是想当上帝
可能在中国生活太久了吧,我对科学进步从未产生过恐惧:科技出现的问题想办法解决就好了。我非常讨厌Dario装神弄鬼、吓唬人的营销。我曾说营销的底线是别让教皇来开发布会,但实际上Dario连白宫都拉出来营销了,只是这次吃了回旋镖了。这卖的毕竟是产品,而不是神谕。
顯示更多
剪贴板历史应该拥有keychain一样的保密等级。然而现状是0人在意
0
43
714
10
轉發到社區
@jlaw520 确实如此,但是这种人自有其目标群体……
总算被BBC逮着一次机会,这是真的无法辩驳了 🤣
With no team on the field, China fans pin hopes on World Cup referee
这作者很有时代感 🤣
今天看了Andreas Fulda @AMFChina 写的德语政治幻想小说 Wenn China angreift(当中国进攻)。该书是虚拟2027年大陆统一台湾。导火索是大陆军机在一次冲突中坠毁,然后大陆军队先演习围堵台岛然后登陆。美国总统选择不干预,而日本女首相渴望干预而因为美国而无果而终,中国顺利实现统一。
顯示更多
这次有相当多的现场视频及照片,第一感是,展现出来的一面已经相当的现代化,这至少说明他们目标是明确和正确的,而且一定程度已经能做到,剩下的是规模化问题,绝对是比越南更有潜力的国家
顯示更多
英语确实目前仍是世界最通用的语言,依然值得投入相当时间去学习,但也没必要夸大其作用,如果不追求流畅的说和优雅的写,目前的 AI 能帮助解决绝大部分问题;至于说与西方社会之间的互信,问题不在语言上,都2026年了,这应该不难看明
顯示更多
我是支持英语作为第二官方语言的,因为英语本质上就是lingua franca、世界语。目前中国一些学科的实际第二语言就是英语。但想法是想法,但这种成本是难以想象的。所以最低要求就是以后学校在英语学习方面不要再削减了,会不会英语对孩子的未来差很大。印度这个国家完全就是因为英语获得了很多优势。
顯示更多
这俩,再加一个新北市伪市长,上次伪总统竞选三人组,再次合流了,但我强烈建议你们不要只搞马娘娘,有志气一点,彻底把KMT弄死多好啊 🤣
赵少康:都是马英九害的,不然我现在已经是副总统了!本来我们民调是和民进党咬得很紧的。他说要和习近平见面,民调掉3个点!后来他又(在接受德国之声采访时)说相信习近平,民调又掉一个点!
顯示更多