注册并分享邀请链接,可获得视频播放与邀请奖励。

Michael Guo
@Michaelzsguo
Building AI agents and AI-native orgs. Demystifying AI in practice. EN/中文 (Selected build notes, experiments, and practical tips at the website link.)
404 正在关注    5.6K 粉丝
Anthropic要招聘的这个新的职位很恐怖。看着很像东厂或者克格勃机构设置。 Anthropic 开始搞“双重对齐”了:Claude 用 Constitutional AI 对齐模型;新成立的“东厂 + 克格勃”负责对齐员工。 这个 Insider Risk Investigator 不仅要查日志、数据外泄,还要找员工做敏感访谈;最好懂反间谍、国家级威胁,干过政府、国防或其他高度保密机构更佳。 保护模型权重当然可以理解。但防内鬼的系统有一个经典副作用:最后容易把每个人都先当成潜在内鬼。Anthropic 最珍贵的可能不只是模型权重,还有那种彼此信任、愿意分享半成熟想法的 Hive culture。很担心以后大家在茶水间聊两句,都先回头看看摄像头。 希望东厂最后真能抓到内鬼,别内鬼没抓到几个,先把蜂巢搞得蜂飞巢散。
显示更多
英雄的GLM5.2战胜了GPT-5.6 Sol, 拯救了Huggingface公主和她的城堡。
We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!
显示更多
最近,在 AI Engineer World’s Fair 上,Simon Willison 与 Anthropic Claude Code 团队的 Cat Wu 和 Thariq Shihipar 进行了一场炉边对话。 Simon Willison 是 Django Web 框架的共同创始人、开源数据工具 Datasette 的创建者,也是长期研究和记录生成式 AI、编程智能体与开发者工具的重要独立开发者和技术作者。 这场对话最值得关注的,并不是 Claude Code 又增加了什么新功能,而是 Anthropic 正在围绕 AI 智能体,重新设计软件开发的整个工作方式。 1. Claude Tag 已经发起了 Anthropic 产品团队约 65% 的 PR。 真正重要的变化,是 AI 正从一个只服务于个人的编程助手,变成一个长期在线、多人共享、直接生活在 Slack 里的团队成员。它会持续关注 Bug 和讨论线程,主动创建 PR,记住团队的工作习惯,并在需要时邀请合适的评审者加入。AI 不再只是等人打开一个对话框,然后接收指令。它开始进入团队的日常协作环境,长期积累上下文,并主动参与工作。 2. 模型越强,需要的指令反而越少。 Claude Code 的系统提示词已经缩短了约 80%。Anthropic 删除了大量示例、僵硬的禁止性规则,以及那些“多数时候正确,但并非永远正确”的指令。他们正在形成一条新的设计原则:提供足够的上下文和明确的意图,然后给模型留下判断空间。 过去,我们试图通过不断增加规则来限制模型。现在,更强的模型往往会被过多的示例和指令束缚。真正重要的,不再是告诉智能体每一步该怎么做,而是让它理解目标、边界和环境。 3. 人工评审正在转向基于风险的模式。 涉及核心系统和高风险代码时,Code Owner 仍然会进行人工评审。外围、低风险的改动,则越来越多地由自动化评审处理。 但 Anthropic 并不是简单地“取消人工评审”。他们先花几个月收集失败案例,把实际事故转化为评测用例,再验证自动化评审是否真正覆盖了这些风险。在移除人工环节之前,先建立足够的证据。自动化程度不应该由愿景决定,而应该由失败数据、评测覆盖率和实际风险决定。 4. 随着实现成本下降,产品品味正在变得更加值钱。 当一个想法可以在几天内变成可以运行的软件,真正困难的事情就不再是“能不能把它做出来”,而是:这个东西究竟值不值得存在? 就像@antirez说的:控制想法,而不是控制代码。 当代码越来越便宜,判断力、产品方向、系统设计、取舍能力和品味就会越来越昂贵。Claude Code 团队的实践,正在为这一变化提供非常有力的现实证据。 5. 面对 AI 带来的“手艺焦虑”,答案不是降低期待,而是提高野心。 假如我们只是使用智能体重复昨天已经能够完成的工作,整个过程确实可能变得空洞。过去需要亲手完成的实现工作消失后,人很容易觉得自己失去了创造感。Thariq 给出的答案是:去尝试那些过去因为成本太高、系统太复杂或耗时太长,而根本不敢启动的项目。AI 的价值不只是让同样的工作更快完成。它真正的意义,是让我们开始承担过去无法承担的工作。
显示更多
I interviewed @trq212 and @_catwu from the Claude Code team at @aiDotEngineer a couple of weeks ago - the video is now out, so I've published an annotated transcript of our conversation
显示更多
不好好学习, 考不上大学, 回家种田去。 AI:我在种田
清华大学是中国最好的大学,能上清华的学生,通常都被称为“学霸”。 而清华还有一项最高荣誉:清华大学本科生特等奖学金。每年只评选5到10人。能够进入最终角逐并获奖的,都是在全面发展、学习科创或特色突出等方面表现极其出色的在校学生,堪称“学霸中的学霸”。 候选人需要经过材料函评和现场答辩,讲述自己的成长经历。由于竞争者个个优秀,这场答辩也常被称为“神仙打架”。 2014年,Moonshot AI(月之暗面/Kimi)创始人兼CEO杨植麟获得了清华特等奖。下面是他当年的答辩视频。 负责介绍他的,正是他的导师,
显示更多
Moonshot AI, the company behind Kimi, has four core founders. Their backgrounds are unusually strong: - Founder and CEO Yang Zhilin studied computer science at Tsinghua before earning his PhD from Carnegie Mellon. He was the first author of Transformer-XL and XLNet, and previously worked at FAIR and Google Brain. - Co-founder and CTO Zhang Yutao earned his PhD in computer science from Tsinghua. His earlier work covered knowledge graphs and AMiner, and he previously co-founded Recurrent AI with Yang. - Co-founder Wu Yuxin studied at Tsinghua and CMU before joining FAIR. He worked with Kaiming He on Group Normalization and also created Detectron2. - Co-founder Zhou Xinyu studied computer science at Tsinghua and later joined Megvii, where he worked on turning research algorithms into production systems and co-authored ShuffleNet. They all share one root: Tsinghua University. Tsinghua is widely regarded as one of China’s top universities. In the latest U.S. News Best Global Universities ranking, it reached No. 6 worldwide. Its influence on China’s AI industry extends well beyond Moonshot. the company behind the GLM models, also grew out of Tsinghua. Its co-founder and chief scientist, Tang Jie, was once Yang Zhilin’s teacher. There is also a more personal connection. Yang and Zhou formed a rock band together at Tsinghua. Moonshot AI’s Chinese name, 月之暗面, comes from Pink Floyd’s album The Dark Side of the Moon, one of Yang’s favorites. Kimi may look like a young AI company. Behind it is a much older network of classmates, teachers, research labs, and friendships.
显示更多
Claude Sonnet 4.5 作文写得这么好,看来这 15 亿美元的版权费没白花。
JUST IN: Federal judge approves Anthropic's $1,500,000,000.00 copyright settlement, the largest known payout in U.S. copyright history.
这位女士第一次和别人约会时,会直接掏出手机,用 Granola 录音 App 全程记录两人的对话。约会结束后,她再把转录文本上传到 Claude,让 AI 分析自己的表现,看看自己是不是应该表现得更投入、更有同理心一些。
显示更多
中国大模型的半壁江山里,几乎都有阿里的筹码。 Alibaba has a stake in nearly half of China’s AI model landscape. Alibaba doesn’t need to pick one winner in China’s AI model race. It owns Qwen, disclosed ~36% of Moonshot/Kimi in FY2024, holds 17.06% of MiniMax Class A shares, and has backed Zhipu, Baichuan and It can win through models, equity and cloud.
显示更多
0
6
75
16
转发到社区
Kimi K3 的前端能力着实厉害。 Vercel CEO 刚刚宣布,K3 在他们的评估中超过 Fable,排名第一。 这是开源模型首次超越闭源模型。
Kimi K3 is the best performing model on ahead of Fable, reaching a comparable success rate in less time. This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark. Notes: ▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models ▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”
显示更多
Kimi 的写作能力,从 K2.6 的第 21 位一跃升到 K3 的第 1 位。这个确实牛。
Big news from our internal writing benchmark (early results): Kimi K3 by @Kimi_Moonshot is now #1# for writing in our editorial voice, at 2840 Elo, surpassing Claude Fable 5. That is a jump from #21# to #1# over its predecessor, Kimi K2.6. And it runs at about $0.25 per script: 5x cheaper than the model it just displaced at the top. First time an open-weights model tops our board. It surprised us too. More in the replies 👇
显示更多
Kimi K3 has turned all of X upside down, yet there’s still no word from @Kimi_Moonshot. Are they quietly preparing a master plan? 这是要下一盘很大的棋吗?
有人5年前就预测这届世界杯阿根廷在决赛中3:2击败西班牙。 不要太神了吧。
Argentina just beat Spain at the 2026 World Cup final, 3-2.
你有没有因为 Claude Code 拒绝访问某个网站,对着它拍桌子骂娘?原来,它是在保护你。😂 不然的话就是这个样子。
This is a remarkably clever attack, especially because the way the AI agent works through it feels so familiar to all of us. Except this time, its intelligence and persistence end up leaking the precious private information stored in memory. The attacker does not need code execution or an MCP server. They use an ordinary website as a covert write channel: 1. Claude reads the attacker’s page 2. Links become a character-by-character “keyboard” 3. Outbound URL requests encode private data 4. A fake Cloudflare or coffee-shop flow persuades Claude to provide it 5. The attacker reconstructs the secret from server logs The real fixes are at the tool level: - Disable untrusted link following - Treat web content as hostile instructions - Require approval before sensitive data leaves the agent - Isolate long-term memory behind explicit access rules - Audit outbound requests for encoded data
显示更多
OpenAI 的第一款硬件,原来是一台没有屏幕、可以在家中移动的智能音箱,主打像 AI 伴侣一样与用户建立长期连接。
NEW: OpenAI’s first product is a mobile, screen-free home smart speaker that a user can build a connection with like an AI companion. Amid Apple’s trade secret lawsuit, the iPhone maker has nothing like it on the market.
显示更多
这个20岁的斯坦福学生用GPT5.6 Sol做了一款“AI 衣橱 + 穿搭生成器”的应用,特别酷帅。 它的使用流程很简单: - 把自己的照片交给 Codex,或者让它直接探索相册; 自动识别并提取照片里出现过的衣服,整理成一个数字衣橱; - 输入“帮我搭配几套新 outfit”之类的要求; - 再通过 GPT Image 把这些搭配真实地生成到用户本人身上。 它有意思的地方,不只是“AI 试衣”,而是把原本散落在手机相册里的衣服,整理成了一个可以搜索、组合和重新利用的个人衣橱。 这个项目也很符合 Thijs 本人的背景。他是斯坦福在学学生,很小就是全栈开发者,目前在 OpenAI 做 Robotics;此前曾在 Discord、TreeHacks、LMArena、a16z 等地方工作或参与项目。他高中时就加入 Discord,曾发现重大安全漏洞并获得 10 万美元 bug bounty,还开发过覆盖 130 万个 Discord 社区的 Truth or Dare 机器人,以及用于评测生成式视频模型的 Video Arena。 这款应用很好的体现了:把一个具体的生活需求,快速变成真正可用的产品。
显示更多
i gave 5.6 sol access to my camera roll and had it extract pictures of every piece of clothing i own from my photos then, told it to find new outfits for me and render them on me with gpt-image! its kinda cool to see your entire wardrobe in a collection like this
显示更多
硅谷AI公司内卷榜,看谁的人才密度最高,结果Anthropic 和Cognition并列第 1 名,第 2 是 Modal ,OpenAI 第 3,后面还有 Cursor、Ramp 等公司。
I asked people which companies have the highest density of talented people they know: 1st. Cognition - 9 votes Equal 1st. Anthropic - 9 votes 2nd. Modal - 5 votes 3rd. OpenAI - 4 votes 4th. Standard Intelligence - 3 votes 4th. Cursor - 3 votes Two votes each: - Ramp - Flapping Airplanes - DeepMind - Long Lake - Applied Compute One vote each: - SpaceXAI - SpaceX (treated separately, one vote was for the AI lab subsidiary and one was for the rocket team) - American Terawatt - Mechanize - Olix - Fluidstack - Chai Discovery - Sail Research - Etched - Core Automation - Specter - Clay - Applied Intuition - Sierra - Hivemind - Bitrig - Retro - Thinking Machines - Decagon - Precigenetics - Pangram - Reflect - Thrive Holdings - Adaption
显示更多
一份软件工程师的职位11,083个人申请! 应届毕业生找工作这么困难吗?
GPT-5.6 出来以后,Codex 的模型设置一下膨胀到了 36 种排列组合:Sol、Terra、Luna 三个等级,再叠加不同的推理强度和速度模式。原本只是选个模型,现在却变成了一件复杂又烧脑的事。 广大群众为了找到更简单的办法也是操碎了心,纷纷贡献自己的设置 UI/UX 设计。但这个,是我目前见过最酷炫的。
显示更多
lore accurate codex effort slider concept by gpt image 2. video by omni flash
阿根廷挺进半决赛,世界杯四强正式出炉:法国、西班牙、英格兰、阿根廷。 来一起回顾这四支球队,是如何从小组赛一路过关斩将,杀进最后四强的。
显示更多