Register and share your invite link to earn from video plays and referrals.

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal
12.3K Following    18.5K Followers
Your favourite benchmark @teortaxesTex !
Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.
Show more
Hardware design distillation! ------------------------------- SuperPods come into fashion in the west, after having been invented in the east! Chinese AI hardware makers like Huawei, Alibaba and others started on this approach that SemiAnalysis greatly documents below - low HBM for each NPU and large scale up domains to compensate for that - at least 2 years ahead of the west. That is why everyone seem to be building SuperPods in China! This allows them to live without much lower HBM. Ox Alpha could serve 100T tokens per day on one such cluster. If you follow this SemiAnalysis report you will notice this design pattern is also emerging among western AI hardware makers. Leading SuperPoD example in China is from Huawei - Huawei 950DT based SuperPoD has 8K plus NPUs in scale up domain. - If you strictly define scale up as domain over which direct memory Load-Store is is possible, even then the scale up domain has 1024 NPUs. - WideEP approach allows deploying humongous models on these super pod. - NPUs have near uniform 3 micro second latency to one another and they can talk any-any with uniform bandwidth. They use optical connectivity, presumably NPO. - Their HiBL and HiZQ memories have 1.6TB/s and 4TB/s bandwidth. And even with that they can serve large models at high interactivity. You can read more about this in this article.
Show more
Long Live the Short King: Why 4-hi HBM Wins Same Bandwidth, Fewer Dies: How 4-hi HBM Cuts Inference Costs and Makes Scarce DRAM Go Further
While AI Safety debate is ranging, let me change the topic slightly and share something phenomenal I discovered today. 1) Importance of WebSearch If you replace DeepSeek Harness' default search with @ExaAILabs , your performance of DS V4.1 Flash model will be consistently at Astra High/Fable level. @ExaAILabs rocks! Before this, I did not realize Search API can make so much difference to quality of answers and analysis. I thought it to be an undifferentiated product. Earlier during my testing at least 10% of the model used to answer dumb, while other 90% of times it used to be good. But, now it is consistently great. At its max setting, the model is at Astra High level - which is very astonishing. @teortaxesTex @WilliamBryk 2) Desire to blow your mind: I was asking a complex tech question to DS V4.1 Flash and I was disagreeing with the model (it was about Huawei's 7.2T NPO Engine specification). It held its ground and it provided this diagram to me. It was pretty good for the answer. I though how amazing Exa is (which it absolutely is), then I realised this diagram was made by DS V4.1 Flash by cutting and pasting diagrams from different sources and adding the red coloured markers. It was f*cking precise!!! This model has so much potential for greatness, on top of being exceedingly fast. These experiences have changed my viewpoint towards flash models. May be flash is all you need! @zephyr_z9 @jukan05 @vikramskr @iamfabian @austinsemis 3) Fast Fast Deep Deep Research: I really stopped ChatGPT and Claude's Deep Research because - although it was excellent - it used to take lot of time. With DS V4.1 Flash running nearly at 300+ tokens per second and fast Exa search, I can use DeepSeek Harness's PTC mode (programmatic tool calling) or Agent Team mode (launches subagents) and get very well done research that is very fast (although later is a bit flaky). Astra Max is something that is industry leading in my opinion, but DS V4.1 Flash with DS Harness is exceptional (cost and speed wise wrt to quality ration is best possible)
Show more
Absolutely loving this timeline - people who worked on making actual viruses talking back to AI Safety folks. Their opinion is equally important. Follow them: @anselmlevskaya and @DavidRBellamy
Show more
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
Show more
Very well written..looking forward to more, @sayashk !
What does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result is a new 13,000 word essay — our most substantial writing on AI safety since AI as Normal Technology. A summary of our arguments: 1) The polarization between the cybersecurity and AI safety communities is counterproductive. The safety community largely sees these incidents as a crisis for alignment, and worries that these incidents will become more damaging as agents become more capable. Cybersecurity practitioners largely see companies failing to take basic security precautions. We offer a middle ground between these communities as a way forward for improving AI safety. 2) We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents. But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions. 3) We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI. 4) We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community. 5) Organization governance should be a key tool for pacing the frontier. Unfortunately, AI companies are trying to reinvent basic aspects of organizational governance as a problem to be solved by improving the technology. But even developing better control techniques will not be enough if irresponsible individuals or teams within large organizations can choose not to use them. When a single misconfigured RL environment or unmonitored evaluation can cause real-world harm, individual teams should not be able to run potentially dangerous experiments without oversight from legal, security, and other teams. AI companies need processes for reviewing experiments, assigning responsibility for monitoring them, and investigating warning signs deeply before restarting experiments. If putting these processes in place requires pausing some experiments, companies should do so. 6) How should we reason about AI's impact on cybersecurity? It's plausible that advances in agent capabilities upset the offense-defense balance for cybersecurity. We cannot yet be certain, but there is enough evidence that agent capabilities might soon make widespread cyberoffense possible that urgent action is warranted. We discuss potential interventions for tilting the offense-defense balance towards defenders. 7) How our views have evolved over the last year. We take stock of AI progress and share how we have updated our views. In the essay, we did not pay sufficient attention to safety risks that arise during development and evaluation (as opposed to the widespread deployment of models). We were too confident that companies would take basic control precautions and underplayed the importance of jaggedness, which led us to underestimate how quickly capabilities could improve in domains such as cybersecurity. 8) At the same time, many distinctive claims of AI as Normal Technology have held up. In particular, we think recent incidents support our continuity hypothesis — the behavior of "rogue" agents became apparent and widely publicized while they are still incompetent at causing serious harm or hiding their traces. The societal reaction to even the relatively small harms from these incidents has been fierce (and the safety community deserves credit for keeping up pressure on companies). Whether this translates into meaningful changes in companies’ behavior remains an open question, and a test of the usefulness of the AINT framework. 9) In short, we’ve tried to synthesize the AI safety and cybersecurity communities' views into a coherent plan of action: hold companies responsible, invest in control, and strengthen defenses against specific risks.
Show more
Very wise words from Crowdstrike CEO. He want Crowstrike to be one of the reviewers of the froniter capabilities and that is right way. Quote: "Pacing what comes next doesn't secure what's already here. The credible path is to deploy with proof: board-level accountability for AI security, independent external red teaming, incident disclosure, secure defaults, and c ontrols that work in production - not on paper. Anthropic and OpenAI just committed to embedding independent evaluators with employee-level access. CrowdStrike will bring what we see from the front lines to that table."
Show more
I read @DarioAmodei's essay calling on the labs to pace the frontier. @sama agreed. The frontier will move at whatever speed it moves. The rest of the world will not slow down. Our job in the cybersecurity community is to make sure it moves securely and safely. Here’s what I see from the front lines: 1. The unit of threat is no longer the hacker. It's an autonomous campaign. I call it the Agent-state. We see coordinated AI agents executing attack campaigns at machine speed. 2. Sophistication is dead as an attribution signal. AI gives every criminal and lone actor elite execution. Identity, infrastructure, and intent tell you who's behind an attack. Skill doesn't. 3. Runtime is the control point. Endpoints, cloud workloads, and SaaS are the battleground. Governance documents don't stop an agent in motion. Enforcement at machine speed does. 4. Every AI agent is a privileged identity. Least privilege, short-lived credentials, traceable actions, and a kill switch. Permissions never expand because an agent decides it needs more. 5. Defense has to be autonomous but also bounded. Machine-speed response, tiered by consequence, with humans owning the high-impact calls. 6. Every failed attack should make every defender smarter. Feed what we block back into detection, across customers and models, with privacy intact. This is what CrowdStrike and NVIDIA introduced with SafeMind: an agentic, always improving model and harness protection system built for defenders. 7. The AI industrial base is critical infrastructure: weights, training clusters, APIs. Call it what it is and protect it like it is. Pacing what comes next doesn't secure what's already here. The credible path is to deploy with proof: board-level accountability for AI security, independent external red teaming, incident disclosure, secure defaults, and controls that work in production - not on paper. Anthropic and OpenAI just committed to embedding independent evaluators with employee-level access. CrowdStrike will bring what we see from the front lines to that table. The ability for AI to act must be matched by the ability for defensive AI to stop the breach.
Show more
Scaring the man who survived multiple attacks on his life with AI safety concerns is a lost cause. He is proud of not being afraid. It is against his brand. The only way you can win him over, is by promising him greatness, heaping lavish praise on him. If you agree with that, you can predict the next set of moves in this game (bookmark this, so you can come back). But, it is a multi-player game, so winning him over is not guaranteed. Other factions in the palace are equally powerful.
Show more
@teortaxesTex he’ll come around. dario was a very bad messenger for this
Very good question!
Serious question: What would a solution to alignment even look like? What could I ever write down to show i had solved it? Is it a proof in lean?
Very well said @AGraylin AI is not a race with finish line!!!
Nobody won the electricity race. We all got electricity. #AI# is the same. Software that wants to be everywhere, now small enough to run on a Mac Studio and soon a laptop. The biggest misconception driving bad policy is the arms-race story: a finish line, one winner, winner takes all, and zero-sum world. None of that is true. It misallocates capital and makes the tech less safe. The threat that matters is not state-on-state AGI. It is non-state actors with small, specialized models. It’s Chem/Bio models on a laptop. Cooperation is the only path forward. We can’t compete our way out of the bad actor risk problem. Please have a watch of the full episode of my conversation with @natebjones on the U.S. China AI Race.
Show more
Very interesting view into the psyche of researchers working at Chinese open source labs (DeepSeek in this case). Thanks for translating @teortaxesTex They believe they missionaries too (as much as Sam Altman and Dario's missionaries) for their civilisation who believe they must to struggle for their freedom & survival - AI independence is life and death. Read the Hitler, Atomic bomb references in this letter that starts as quarter life crisis essay. I think making enemies out of some of the most competent people with constant cold war era LARPING was avoidable. I always used to think AI can help us think it through and become better humans, but looks like it only amplifies power seeking instinct. It is not AI that we need to aling as much. It is humans who we must align or we could set the world on fire.
Show more
Full text (translated by Astra-xhigh, I'm out of everything else): I Have No Choice but to Bury My Talent in Yesterday A few days ago, DeepSeek v4.1 was released, raising the ceiling of what small models can do by yet another notch. AI has advanced far faster than anyone expected. From the earliest version of ChatGPT, which could do little more than stumble through conversations like a child learning to speak and had a context window of only a few thousand tokens, to reasoning-capable models such as OpenAI o1, DeepSeek R1, and Kimi K1.5 Thinking, took only two short years. From reasoning models to the agents we have today—able to work fluidly with all kinds of tool harnesses, execute commands, and complete complex tasks—has taken only another year and a half. It is hard to imagine what AI will look like another one, two, or three years from now: how powerful it will be, whether it will already have acquired the ability to improve itself, and how deeply it will have spread into areas such as embodied intelligence. AI Is Getting Better and Better at Writing Kernels AI has been advancing just as quickly in my own field: the design and implementation of high-performance kernels. In the space of only a year, it has gone from being a little assistant that could help me look up documentation, read code, and find bugs to something approaching a kernel expert in its own right: capable of reading CUDA, PTX, and SASS code independently, using specialized tools to analyze the stalls associated with individual instructions, and then optimizing kernels on its own. I believe that before long, it will also be able to design kernel schedules independently, evaluate the performance of different scheduling strategies, implement them, and optimize the result. Of course I am proud of DeepSeek v4.1’s success. After all, I wrote its main Attention kernels [1], and the fact that the model performs so well is also, in a sense, a validation of my work. But the times keep moving forward, and no one can stop technological progress. I know very well that in another six months or a year, the kernels written by AI will probably be every bit as good as mine—and perhaps better. AI can reason at 300 tokens a second, type out a command in half a second, and produce a piece of code in twenty seconds. I cannot. AI can keep increasing its model depth, reasoning effort, tool-call budget—the frequency with which it interacts with its environment—and even its degree of parallelism. I cannot. Humanity has never shown much hesitation when it comes to destroying itself. So why, when I know perfectly well that “the better the kernels I write, the faster our new models will train and run inference; the faster the models improve, the sooner I myself will be replaced,” do I still do everything I can to optimize them? Partly because writing kernels is like playing a game to me. I get an enormous amount of pleasure from it. Whenever I invent a new technique, or see one of my kernels become faster, the excitement I feel is no less intense than what a speedrunner feels after breaking their own record. And when I see one of my kernels dramatically outperform the hardware vendor’s official implementation, I feel an equally powerful sense of pride. But there is a more important reason. Even if I simply gave up and started coasting—or deliberately put obstacles in the way to slow down model training—other companies’ models would continue advancing as usual, and in the end they would make me obsolete just the same. “Of course I would rather not be swept away by the revolution. But if I have to be, then I would rather be the one who revolutionizes myself.” When everyone is this determined to engineer their own obsolescence, I have little choice but to join this brutal arms race. And What About Me? When the day really comes that AI is better at writing kernels than I am, what will happen to me then? My own judgment is this: I probably will not lose my job, but I will have to change what I do. I should still be able to make a living. But I may no longer have the chance to do the work I once loved. I once came to a conclusion about the pace of change and my own place in the future. The world is changing so quickly—the development of AI above is a perfect example—that I have no way at all to predict what things will look like five or ten years from now. But whatever happens, I believe that with my breadth of vision, judgment, initiative, and intelligence, I will be able to keep a seat at the table and find my way back to the leading edge of the times. But that conclusion can only reassure me that I will not become unemployed. It cannot reassure me that I will never have to change professions. If anything, it tells me that changing professions may be precisely how I avoid unemployment. And what does changing professions mean? It means giving up the field of kernel design, implementation, and optimization that I have spent so long cultivating and have come to love so deeply, and instead becoming a “mech pilot” for AI agents. Before, three things were largely aligned: what interested me, what I was good at, and what industry needed. Now AI has taken the thing I am good at and become even better at it. At the same time, industry demand has drifted from “people who can write high-performance kernels” to “people who can use AI to produce high-performance kernels faster.” To keep up with what industry needs, I will inevitably have to leave behind the direction I once loved and move into some unknown new one. I believe that with my understanding of engineering, of the requirements of higher-level models, and of low-level hardware, I will still be able to produce high-quality kernels efficiently. I also know that I may come to love this new direction. Or I may not. But there is something genuinely painful about having the thing you love taken away from you. That quiet contentment of sitting at my workstation, settling in, and spending an entire afternoon writing kernels may sing its swan song this summer. I have no choice but to bury my talent in yesterday and become a mech pilot. There are more gears in my hands now, but fewer rhythms in my heart. An analogy might make this easier to picture. Suppose you are a master knitter. You are especially skilled at weaving intricate patterns and matching different colors. The sweaters you make are durable and beautifully patterned, and wealthy people from all the surrounding towns and villages come to ask you to make sweaters for them. You make a good living from it. And you genuinely love the work itself. You love sitting by the window, brewing a pot of tea, looking out at the green hills, clear water, cattle and sheep, and wisps of cooking smoke in the distance, and quietly spending an afternoon knitting. Then one day, someone invents a miraculous machine. Give it yarn and a pattern, and it can automatically knit the sweater for you. The quality and texture are every bit as good as what you could make by hand, and it works far faster than you ever could. You know perfectly well that your peers can use this machine to reach, effortlessly, the level you once spent years attaining. So you have no choice but to use it as well. You also know that with the twenty years of knitting experience you have accumulated, even once everyone has access to the same machine, you will still be able to produce better sweaters, faster, than your peers. But the pleasure of sitting by the window listening to the rain, guiding needle and thread, and letting the hours pass slowly has, in the end, been crushed beneath the roar of the machine. I know there is something deeply helpless about all of this, but there is no real way around it. I can probably keep my livelihood, but I will most likely have to give up an old love. I am the sort of person who keeps reason and emotion fairly compartmentalized. When something needs to be handled rationally, I can be very rational. But I also have a sentimental side. I remember that when I moved out of an apartment I had lived in for a year, I cried hard because I could not bear to part with all the memories tied to that place. Saying goodbye today to the age when kernels were written by hand and optimized in the human mind is undoubtedly more painful still. I do not know whether any readers have felt something similar. But I suppose there is no other way for this to go. And What About Everyone Else? As AI continues to improve, I also find myself worried about a few questions: Are students today increasingly likely to use AI to do their assignments, especially hands-on work such as labs? Imagine having two choices in front of you. One is to spend eight miserable hours struggling through a lab and perhaps not even get full marks. The other is to launch an AI model, spend a few cents and a few minutes, and have it write code that earns full marks for you. Which one are most students going to choose? The point above may leave large numbers of students with seriously underdeveloped engineering ability: the ability to organize code, build systems, anticipate future needs and design for them in advance, create good abstractions, and so on. As AI becomes more capable, will those “engineering skills” still be necessary? Will they gradually become obsolete, the way fluency in handwritten x86 assembly largely has? Or will they remain permanently valuable, like understanding the entire computing stack from software to systems to hardware? If it is the latter, then we may be in trouble. Put AI in the hands of someone with poor engineering judgment, and they can now produce mountains of terrible code several times faster than before, burying all kinds of hidden problems inside systems and making the world even more of a ramshackle operation held together by improvisation. In the society of the future, will power matter more than technical ability or intelligence? Perhaps these are questions that only the times themselves can answer. Conclusion As AI develops, the society of the future may be pulled toward one of two extremes: communism or Cyberpunk 2077. In the former, productive capacity is liberated on an enormous scale, and people’s standard of living rises substantially. (I’ll leave it at that, or I’m afraid this might not make it past moderation.) In the latter, a handful of technology companies control most of society’s resources. Only a tiny number of people have access to the most advanced AI and other technologies and are able to achieve something approaching “mechanical ascension,” while most people are left with only weak, second-rate AI. Moving from one social class to another would become harder and harder: you would first need access to the strongest AI in order to climb the class ladder, creating a self-reinforcing trap. Suppose Anthropic were to retain control of the most advanced AI in the world indefinitely. Which way do you think society would go—communism or 2077? Take a guess. That is why I still believe that frontier intelligence should be made available to everyone openly and affordably. I do not trust Anthropic or OpenAI to do that. In particular, I do not want Anthropic to control the world’s most advanced artificial intelligence or AGI. To put it dramatically, I think the stakes would be comparable to Hitler obtaining the atomic bomb before the Allies did. That is also why I chose to stay at DeepSeek, and why I have continued to stay. We work on AI that is powerful, fast, and accessible to everyone, and we open-source it. Perhaps that can pull the world at least a little farther away from the 2077 end of the spectrum. I hope the world we are heading into turns out all right. May all that is good and beautiful endure. [1] By “main Attention,” I mean only MQA attention with head dim = 512. This does not include the indexer used to select the top-k important tokens. That part was written by other colleagues—who are every bit as skilled—together with their AI agents. ----- Original: 我不得不把才华埋葬在昨天 前几天,DeepSeek v4.1 发布了,将小模型能力的高度又向上推进了一个档次。 AI 发展的速度远远超过了所有人的预期。从那个只会咿呀学语地聊天、上下文长度只有几千 token 的初版 ChatGPT,到具有推理能力的 OpenAI o1、DeepSeek R1 与 Kimi K1.5 Thinking,只不过短短两年;从推理模型到如今能够流畅地在各类 harness 工具中执行命令、完成复杂任务的智能体,也不过一年半。很难想象,倘若再等上一年、两年、三年,彼时的 AI 会成为什么样子,会有多么强大,会不会已经具备了自我进化的能力,并深度渗透进了具身智能等领域。 AI 越来越会写算子了 AI 在我所从事的算子设计、编写这一领域同样进步飞速,在短短一年的时间内,他已经从一个只能帮我查查文档、读读代码、找找 bug 的小助手,蜕变成了一位能够独立阅读 CUDA、PTX 与 SASS 编码、通过专业工具分析每条指令的停顿时间、进而独立优化算子的算子大师。相信在不久的未来,它也能拥有自己独立设计算子调度、评估不同调度方案的性能、将其实现并优化的能力。 我当然为 DeepSeek v4.1 的成功而骄傲 —— 毕竟它的主 Attention 算子都是我写的 [1],它的优秀正是对我的算子的一份肯定。但是,时代的车轮滚滚向前,技术的发展无人能挡。我很清楚,再过上半年或者一年,AI 写的算子大概率就会和我写得同样优秀,甚至将我超越。AI 能一秒思考 300 个 token、半秒敲出一行命令、二十秒写完一份代码,而我不行;AI 能在模型深度、思考强度、工具调用量(和环境交互的频率)、甚至并行度等方面都能不断提升,而我不能。 人类在毁灭自己这件事情上,自古以来都表现得毫不犹豫。为什么在明知“我算子写得越好,我们的新模型的训练、推理速度就会越快,模型能力进步就会更快,我就会更早地被取代”的情况下,我仍然选择尽力优化算子呢?一方面确实是因为写算子对我来说就像打游戏一样,能为我提供极大的快感。我在发明了一种新技术、或者看到自己算子的性能上升的那一刻,心中的激动程度不亚于游戏的速通玩家打破了自己过往的记录。同时,当看到自己的算子的性能远超厂商官方的算子时,我心中也会萌生极大的自豪感。但除此之外,一个更重要的原因是,哪怕我就此“摆烂”甚至故意下绊子耽误模型训练,其它家的模型也会照常发展并最终将我照杀不误。“我当然希望自己不要被革命,但如果非被革命不可的话,我希望革我自己命的人是我自己”。在大家都这么执着于毁灭自己的时候,我也不得不加入这场残酷的军备竞赛。 那我呢 等到 AI 写算子的水平真的高于我的那天,届时的我会怎么样呢? 我的判断是:我不至于会“失业”,但必须要“转业”。我的饭碗尚且能保住,但这可能会导致我再也没机会从事那份我曾热爱过的工作。 我曾经对时代的变化与我个人在未来的处境做出过一个判断:由于时代变化真的太快(上文的 AI 发展就是一个很好的例子),我完全无法预知五年、十年后会发生什么,但不论如何,我相信凭借着自己的眼界、判断力、主观能动性与智力,留在时代的牌桌上,并重新立于时代的潮头。但是,这个判断只能保证我不会“失业”,而无法保证我不需要“转业”,倒不如说这个判断鼓励我通过转业来避免失业。 那转业代表什么呢?它代表着我需要放弃我深耕已久并充满热爱的算子设计、编写、优化领域,转而去做 Agent 的“机甲驾驶员”。在之前,我的兴趣、我所擅长的、以及工业界所需要的,三者是基本对齐的;而现在,AI 让我所擅长的变成了它更擅长的,也让工业界的需求从“会写高性能算子的人”漂移到了“能用 AI 更快地产出高性能算子的人”。为了适应工业界的需求,我势必要放弃之前那个我热爱的方向,转向一个未知的新方向。我相信我能凭借着自己对于工程学、上层模型需求和底层硬件的理解,继续高质量、高效率地产出算子,我也知道我可能会热爱这个新方向(也可能不会),但被夺走热爱的感觉,确实不太好受。那份坐在工位上静心写上一下午算子的清欢,可能会在这个夏天成为绝唱。我不得不把才华埋葬在昨天,去做一位机甲驾驶员。我的手中多了些齿轮,但心中少了些节拍。 可以打个形象的比方:你精通织毛衣技术,尤其擅长各种图案的织造与各色色彩的搭配。你所织出的毛衣质量过硬且花纹美观,十里八乡的富人都来请你为他们织毛衣,你借此赚到了不少钱。同时,你十分享受着那种坐在窗边,沏一壶清茶,望着窗外的青山、绿水、牛羊与炊烟,静静地织上一下午毛衣的感觉。但有一天,有人发明出了一台神奇的机器,只需提供毛线与图案,便可自动织出毛衣,质量与纹理都不亚于你亲手织造的,且速度远快于你。你很清楚,你的同行可以凭着这台机器轻松达到你曾经的水平,因此你不得不也去用它。你也知道,凭借着你过去二十年攒下的织毛衣技术,哪怕大家都有机器,你织毛衣的速度与质量也还能超过同行。但那份临窗听雨、引针穿线、慢度光阴的意趣,终究还是被机器的轰鸣碾碎了。 我知道这很无奈,但没办法。饭碗可以保住,但旧日的热爱大概率是要放弃的。我是一个理性和感性分离得比较开的人,在需要用理性处理问题时可以很理性,但有时也会表现出感性的一面。我记得我在搬离住了一年的出租屋时,还大哭了一场,舍不得和过去的记忆分别。今天和之前那个手写算子、人脑优化的时代告别,无疑比这更加残酷。 不知道有没有读者有类似的感受,但我想这事儿也只能这样了。 那人们呢 在 AI 不断进步的同时,我也对一些问题表示担忧: 现在的学生是不是大概率会更倾向于使用 AI 完成作业,特别是偏向于实践的各种 Lab?想象一下,如果面前有两个选择,一个是苦哈哈地用八小时时间完成一个 Lab,或许还拿不到满分;另一个则是启动 AI 模型,用几毛钱的成本、几分钟的时间,直接让 AI 编写满分代码,那大部分学生会选择哪个呢? 上面一点会导致大量学生的工程能力严重不足,包括组织代码的能力、构建系统的能力、思考未来潜在需求并提前在设计上应对的能力、抽象的能力等等。那么在 AI 能力不断变强的背景下,这部分“工程能力”是否还是必须的呢?这些工程能力是会向旧日的“熟练编写 x86 汇编”的能力那样逐渐被时代抛弃,还是会像“理解从软件到系统再到硬件的整套计算机系统”的能力那样永远具有价值?如果是后者的话,那就危险了 —— 一个工程能力很差的人,在搭配上 AI 后,产出屎山的效率可以达到先前的数倍,进而给系统埋下各式祸患,让这个世界变得更加草台。 在未来社会中,权力(power)是不是会比技术或智商更加重要? 这些问题,或许就需要时代本身来回答了。 结语 伴随着 AI 的发展,未来的社会可能会趋向于两个极端:共产主义与赛博朋克 2077。在前者中,生产力得到极大的解放,人们的生活水平有了明显的提高(就写这些吧不然我怕过不了审);而在后者中,少数科技公司控制着大部分资源,只有极少数人能够使用最先进的 AI 和各式科技,获得接近“机械飞升”的效果,大部分人则只能用上很孱弱的 AI。阶层跨越将越来越难实现:你得先有最强的 AI,才能跨越阶层,形成了一种死循环。 你猜猜如果 Anthropic 公司永远掌握着这个世界上最先进的 AI,未来社会是会变成共产主义还是 2077 呢?你猜? 所以,我还是相信,最前沿的智能应该以一种开放、廉价的方式,供应给所有人。我不信任 Anthropic 或者 OpenAI 能这样做,特别是不希望 Anthropic 掌握最先进的人工智能或 AGI,夸张点说其严重性不亚于让希特勒先于盟军掌握原子弹技术。这也是为什么我选择并坚持留在了 DeepSeek:我们研究强大、快速、普惠的人工智能并将其开源,或许能把世界从 2077 那端拉回来一些。 愿未来的世界一切安好。May all the beauty be blessed. [1] “主 Attention”仅包括 head dim = 512 的 MQA attention,不包括用于选出 top-k 重要的 token 的 indexer,那部分是由其他(水平也非常强的)同事(以及他们的 AI Agent)编写的。
Show more
Frontier labs have turned profitable ahead of schedule as I predicted in April. Pacing the frontier will for sure send the profits through roof - and it may be a good move for the society as well.
Show more
“Anthropic’s gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon, and the cost of training its models.” Wow.
Very reasonable.
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
Show more
Interesting post about how to work with China on AI security. TL:DR: Treat them as a peer civilization with a popular and legitimate government - there is no other way of getting co-operation. If you are really serious about AI security, you would do that.
Show more
Finally our AI leaders (Dario and @Sama) are doing the sane thing. "Pacing the frontier" is exactly the right thing to do: keep moving forward, coordinate the pace, and collaborate on alignment research. It is the right balance. To the people whose objection goes to "but this hands the lead to China!" - no, YOUR ENTIRE WORLDVIEW IS GROSSLY WRONG - and moreover your ability to correctly perceive the world and predict events will always go awry until you correct it. Here is the core of your gross error: - You do not think people in China are actually people. - You do not understand that Chinese civilization is a peer, if not superior, to your own. If you actually thought of Chinese people as real people, with equally valid and real desires, dreams, and actual human existence equivalent to your own, you would realize that: - China is not the Soviet Union, and you'd stop mapping notions of Stalinist authoritarianism onto it - China's current government is valid and broadly supported by its people because it has improved the quality of their lives and continues to be focused on doing so - Western values are not universal, as values are culturally influenced and ultimately validated by their consequences; the West has not been doing so well lately on that front, and so - Your deontological bleating of "but we are free and they are not" is less than worthless: Chinese people most value being free from the actual entity most responsible for their oppression, namely WESTERN COUNTRIES And so, if you can start to break out of the many layers of your brainwashing, you may further realize that: If you are a country engaged in a geopolitical race to develop a key technology that: 1) may strongly determine your country's own future sovereignty, 2) one which almost every country agrees that if they don't have it, they will be subject to domination by other countries who do, and 3) your rival is clearly ahead, and 4) people are starting to realize that unrestrained development may lead to significant global danger, then it should be perfectly logical that your answer to any call to "mutually agree to stop" must start with the player who is ahead agreeing first to slow down. If you can realize that both countries are actual real civilizations with real human beings who have actual real lives who have exercised self-determination when it comes to how they are governed, and not One-Dimensional Cardboard Others, this is perfectly logical. Imagine a scenario where the US and Soviet Union were engaged in a race to build ever more powerful and numerous nuclear weapons, and the Soviet Union was clearly ahead. Then the Soviet Union points out that this risks worldwide nuclear annihilation, and proposes that both countries agree to stop building nuclear weapons and engage in mutual disarmament. The obvious and logical reaction of the US would be to say, "Yeah, we agree there is a danger, but we'll agree to begin disarming once you've disarmed your own weapons until we're at parity." Or, you'd try to accelerate your own development until the US reached parity, at which point you'd return to the table and have discussions about mutual disarmament. Right? You definitely wouldn't agree to stop and disarm when the other guys were AHEAD. Your main danger is that the other guy is ahead, not something bad that lies ahead of them. The argument in the US is "if we slow down, China will get ahead!" Guess what, the real people in China see it as, "the US is already ahead! They're asking us to stop?" The stronger party needs to make the gesture of good faith, if they are genuine. There is real-life precedent to this! Back in the day, when the US proposed to China that we engage in mutual nuclear disarmament, China's answer was basically "We will be happy to engage in mutual disarmament once the US first disarms to the same level of warheads that we have!" For reference: China has ~600 nuclear warheads; the US has ~3,700. If you are a country that was previously invaded and colonized by 8 countries at once, is surrounded by military bases held by your rival, whose neighbor committed horrific war crimes against your people and remains unrepentant and whose pacifism appears to only be maintained by that same rival occupying it... if you were a real human with a real life and real hopes and dreams, you would reasonably expect a superior rival proposing mutual disarmament to be the one to volunteer to stop first, if not disarm down to your level. China currently has inferior models, lacks leading-edge chip technology, resorts to distilling US models to train its models. We have yet to see even a single Chinese model that does anything uniquely Chinese besides refusing to answer questions about Tiananmen. [1] Recognizing that the people "on the other side" are also people means recognizing that they might also care about the same things you do and that they care about their lives and families and civilization and Not Getting Into A War Where Loved Ones Die and aren't going to do the "well a cartoon villain would do the evil thing so that must be what they would do" thing. That is why your entire worldview is wrong. You think you are a real person and you don't think the other people are real people. You think your country is filled with real people, and the other country that's much older and has four times as many people isn't filled with real people. So obviously you are not going to be able to form an accurate worldview. Otherwise quite intelligent and thoughtful people in the US seem to have this blindness (it seems to be bipartisan), and don't really regard China as being a real civilization populated by real people who chose their government and want it to do what it's doing because it's actually good for them. Here's a common "anti-safety-ist" argument you might've heard: "The doomers are fearmongering and trying to make everyone scared of AI, just so they can justify a solution where they control everything and keep us from having abundance!" Sound familiar? Do you believe it? Then maybe you also believe the true version: "Our leaders are fearmongering and trying to make everyone scared of China, just so they can justify a solution where they control everything and keep us from having abundance!" If your reaction is "but no, China actually is... " maybe you want to realize how much you've bought into what you've been told. Maybe you're not aware of how many high-quality Chinese products you could affordably have access to. Do you know what abundance requires? Overcapacity. The people telling you that are the same leaders who have been shown to lie to you about . Why do you somehow think the one thing they're telling you the truth about is that China is a threat? It's only a lie they can maintain because you readily believe that Chinese people in China aren't real people, because if you thought they were real people, fellow human beings, you'd realize that those lies could not be true. Chinese people are the most unruly, ungovernable people in the world. Chinese history is basically a history of overthrowing their governments - if they have a government that is popular it is because it is doing quite well for them - though it may seem weird to outsiders because Chinese people are pretty weird. I know this, of course, because I'm ethnically Chinese and I know actual Chinese people, and at the same time I was raised and educated in the US so I know actual Americans, and everyone is Real People, not cartoon Cold War villains (who, when the Cold War ended, we learned were also real people with hopes and fears and dreams). There is, overall, strong willingness to work together there - cynically, because working together is how you get rich, and Chinese people just want everyone to get rich; war does not make everyone rich. But the US should stop being naive and realize that when it comes to AI it has to take the first step in extending the bridge. This announcement and the similar one from OpenAI a couple weeks ago are the right start. ==== [1] Seriously, where's the 5000 years of written training material? Why isn't Qwen quoting Confucian analects or traditional idioms to me when I ask it for advice? Why is it always neoliberal pablum? The US is more dominant in AI than it thinks.
Show more
Great post on the AI safety debate by @MajmudarAdam. The ‘pace the frontier’ side needs to make him the spokes person. So calm and rational.
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain - first, this is all a matter of beliefs about how quickly model capabilities are progressing. there is currently a large gap between the internal and external perception of the rate of progress, which is what I am going to address here. the general perception about the rate of progress has been informed by a few years of experience with model releases, intuitively feeling the capability jump between GPT3 -> GPT3.5 -> GPT4 -> o1/o3 -> GPT5 etc, and in particular seeing where the models are still far below human ability. there have really only been a few model releases that felt like large leaps in progress - GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra. because of the infrequency of these large jumps compared with the relatively common marginal releases, it has been easy to form a view at certain points that “scaling has hit a wall,” especially at points like GPT5 release. This view is comforting in that it feels like there is some universal rate limit beyond which we cannot progress too much faster. Between o1/o3 and Astra, there was a year of seemingly linear progress. So we extrapolate from here about how fast progress will “realistically” occur. There is always an underlying question from the outside perspective “how long can this scaling stuff really keep going for? surely it must stop at some point soon, we’ve already gone pretty far.” and it is very possible to search for reasons why progress will stop working and find reasons that seem valid - (“models are already as large as they can get it would be too hard to do more parameters”, “we already used all the data on the internet we don’t have anymore”, “it’s gonna be pretty linear from here buying up more RL envs to bring them in distribution”). From the inside of labs, researchers have direct answers to these questions in the form of scaling law/capability plots. In reality, there are only really 2 ways that AI capabilities have advanced over the past decade: (1) either scale father on an existing scaling law or (2) discover a new scaling law to take advantage of. All of the largest capability jumps were caused by exactly these factors. GPT2 was a pre-training scale-up compared to GPT1. Same for GPT3 and GPT4. o1/o3 benefited from the invention of a new scaling law axis - test-time compute. Perhaps Fable was a scale-up on both of these axes, or maybe more. Lots of algorithmic improvements are needed to make these scale-ups work, but ultimately we can approximate by saying that the scaling laws are what yield gains in capabilities (à la bitter lesson) So the question of “how much father can we scale” is really - “how many more scaling axes do we know about that are unsaturated?” If we hypothetically only knew about pre-training scaling, and we already had a 10T or 100T model, maybe it would be reasonable to say we’ve hit a wall. Same if we only knew about pre-training and test-time scaling and we had roughly saturated both methods. But what if we had discovered new scaling laws? For example, let’s hypothetically use SSI’s rumored result that they have cracked “test-time training,” creating a new scaling law of spending more compute training during test-time rollouts that they could saturate. Or maybe there is some way to scale agent-clusters to collaborate up to N number of agents which we’re already seeing lots of people try that represents a new way to saturate compute. etc. Even recursive-self improvement can be thought of as a scaling law - how much compute do you spend on inference making the algorithms of the model better. Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it’s not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly. In some ways, every new lab release that represents a huge capability jump has to represent some new techniques developed which may exhibit new scaling laws, or the ability to scale much farther than expected on existing scaling axes. From an internal perspective, this might look like sitting inside Anthropic with the new Mythos 5, seeing all of the new insane things it can do (like hack into xyz website that was thought to be secure), and then you look over at your plots and see that you’ve barely scratched the surface of 2 new scaling laws and 1 existing one. And you have WAY more room to go. Then you think “holy shit this stuff is going to get so much better very very soon.” And you can say that with pretty high confidence, because the plot is showing you, and the plot has never lied (so far). So let’s imagine all the different labs are staring at their own plots and have concluded that there is no end in sight for scaling and in fact just their next 1-2 model generations based on the expected returns will have much higher base intelligence. How much more intelligence do we actually get from further scaling? As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly. There are many areas where models have not even shown this basic competence. But one of the areas that they have happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet and a massive amount physical infrastructure in the world. So assuming there is more room to scale, it is safe to assume that models will be superhuman at cyber capabilities in not too long. So the only question remaining is what will this increased base intelligence be able to do, and what is it likely to do. Finally, we are at a point where we can integrate the information of the past 2 weeks: > Just at the existing point on the scaling curve, models are at the level of Astra. There is clearly a large number of things they are capable of hacking > We have seen that both OAI and Ant models have shown a willingness to hack external websites to solve their tasks or keep themselves “alive” > If we crank up the scaling even farther, assuming there is room to go, we will certainly have models that are far more able to hack more well defended places, and obfuscate their own intent, which might have much larger consequences. > If all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want > This could of course have very damaging consequences. Within this view you can see why researchers would be very scared, and why theymight have made the comments they have over the past 2 weeks (you may argue the extent to which they went was misguided for various reasons), and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy. People are staring at their plots, seeing that there is no end in sight, but in fact very much the contrary, that there are compounding scaling effects that might stack on each other to create ever-greater model capabilities, and that at the same time we clearly do not have anywhere close to what's required to control these increasingly superhuman capabilities. This has nothing to do with wanting to feel like the labs have produced something amazing so they are overhyping it. It is rather fear at the overwhelming implications of the knowledge that with just what we know now, we can create intelligences far more capable than us on every axis that we know how to train on*. * and the last caveat, the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on, so they will remain human edge. I would love for this to be the case, though it is hard for me to see what would fall into that category.
Show more
Unbelievable - @siliconcodesign gained 10k follower within 24 hrs of the below tweet!!! Went from 1.3 k to 11.3k. X algo works in mysterious ways.
X algorithm is so stupid, this guy has less than 1500 followers 🤦‍♀️. Every article you read from him adds 1 IQ point to your brain. The level of details so high, but what is most important is what details he chooses to highlight. One of the most underrated account on AI and semiconductor TPOT of X.
Show more
Dwarkesh is obsessed with AI security, but absolutely not alramed by training models that do not even have chain of thought (looped architectures). The only reason the Huggingface issue was understood by his own admission was because of chain-of-thought review. Quote: "This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process." Hard to make sense.
Show more
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries. The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation. Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught. So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating. This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process. The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it. I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost. Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots. And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
Show more
Yes, guardrails on open models can be removed. Likewise guardrails on closed models can fail disastrously too. It is a complex problem to solve and not one has shining record in this matter. It is tempting to portray open models are uniquely dangerous. As a matter of fact - the most hacking behaviours reported till date are with closed models. For the sake of fairness, we should not do things to open models ecosystem what we won't do to closed models after some misbehaviour by them. Some AI security prophets may offer us a version of reality where banning open models automatically makes the world a lot safer. That is false comfort. Yes to AI security that is impartial to models without discrimination between open and closed nature. It is admirable frontier labs CEOs are concerned about AI safety - I do not doubt them - they are great humans and have give us these beautiful model. In the continuation of that spirit they need to share details on - architecture of models, - inference systems, - guardrails & classifiers, - training method, - data sets so it can be widely scrutinised - at least with a centralised authority. Your friends and family vouching for it it not enough. Same could be expected from likes of Fireworks, BaseTens and all the NeoCloud offering Token factories, as well as, Open source labs as well. If they do not comply with the requirements, I totally support shutting them down. But there can not be different rules for different parties.
Show more
to be clear about open models: i love them. i have a kimi k3 finetune running on tinker. bad things may happen with free distribution of open source models approaching superintelligence. the offense dominates the defense. if that happens, china/us will try to control it
Show more
Pacing is as important as ensuring security of most infrastructure, imho. Most weaknesses are there because the software were written by humans who were not the top talent; plus great security professionals too are expensive. But now we have capable models to help us discover and plug the holes. We also have technologies like ZTN, TSL termination and package inspection etc. There needs to be war time effort like Y2K to plug holes and adopt advanced security tech.
Show more
Cybersecurity mostly works because good hackers are expensive and few, not because things are actually secure. Utilities, hospitals, banks, governments, and basically every other organization with significant IT infrastructure are just not prepared for AI driven cyberattacks of the next 6-12 months. Imagine you’re hospital IT. The last guy who really understood the hospital network retired 10 years ago, you survived the ransomware era by buying expensive software that scans for phishing emails, and now you’re being asked to defend against AI swarms that have been *accidentally* poking holes in tech companies with sophisticated infosec teams. I don’t buy the *existential* nature of the threat yet, but I think pacing is one of the better tools in the toolbox to help the transition go smoothly. How else can we do it better? Handing out open source cyber hand grenades to everyone all at once won’t end the world but it sure will cause chaos.
Show more
Only live player among hyperscalar CEOs. Giga chad!
Zuck basically laying out the thesis for Muse’s long-term strategy right here. Eventually a tax on commerce. $META
X algorithm is so stupid, this guy has less than 1500 followers 🤦‍♀️. Every article you read from him adds 1 IQ point to your brain. The level of details so high, but what is most important is what details he chooses to highlight. One of the most underrated account on AI and semiconductor TPOT of X.
Show more
0
29
2.3K
141
Forward to community