注册并分享邀请链接,可获得视频播放与邀请奖励。

yihong0618
@yihong0618
喜欢王小波,大概我们能成为朋友。 我的 2026 我的 2025 我的 2024 ..............
3.7K 正在关注    78K 粉丝
频道里的朋友推荐的他最近见过的最有人味儿的官网也是我见过的。
Vane 0.1.0 正式发布 🎉 今年 7 月中旬,我们决定将 Vane 开源,希望它能在真实的使用、讨论和协作中持续成长。这段时间,最让我感到惊喜的是在 库中有来自外部的贡献者,说明是真的有用户在尝试 Vane。 感谢每一个 Star、Issue、PR 和建议,也感谢所有参与开发、关注和支持 Vane 的朋友。 0.1.0 是一个新的起点。欢迎使用,也欢迎一起参与建设。
显示更多
🎉 Vane 0.1.0 Is Officially Released Vane Data is a high-performance multimodal data engine built for AI workloads. Forked from DuckDB, it provides native multimodal processing and a unified execution model for both local and distributed environments. 🔧 Core Features Distributed DuckDB execution engine Extends DuckDB with distributed physical plans, distributed Plan Fragments, FTE (Fault-Tolerant Execution) scheduling, and distributed operator execution, with Arrow Flight providing cross-worker Exchange/Shuffle data transport. Python UDFs Relation UDFs support row-wise map, Arrow Table-based map_batches, and one-to-many flat_map. Expression UDFs provide @vane.func, @vane.cls, and their corresponding .batch forms. Scalar, batch, and class UDFs can all be registered as SQL functions through vane.attach_function(). AI Functions Provides typed Prompt and Embed APIs across the Python Expression API, Relation API, and SQL. Prompt integrates with OpenAI, Anthropic, Google, and the native vLLM backend, while Embed supports OpenAI, Google, and SentenceTransformers. Structured outputs and image Prompt inputs are available where supported by the provider. Native vLLM batch execution Implements a native Physical VLLM operator and Actor Pool, with bounded task submission enforced through in-flight limits. Prompts are bucketed by shared prefixes and routed to actors using prefix-aware routing to improve opportunities for reusing the vLLM Prefix Cache. The native vLLM Prompt path currently supports text input only. Adaptive multimodal batching and backpressure The UDF and vLLM execution paths dynamically split or combine batches according to row count, data size, and in-flight limits. Resource admission control and object-stream backpressure limit the number of queued tasks and their memory consumption. Fault-Tolerant Execution Supports task retries, Worker failure detection and replacement, Split reassignment, Attempt Fencing, cancellation, and resource cleanup. Ray Runner and Local Runner The same SQL and Relation plan model can run through either the distributed Ray Runner or the local In-Process FTE Runner. Ray Runner is the default execution path and supports both single-machine and distributed execution. Local Runner targets lightweight, lower-overhead local execution without Ray. Local Runner is currently experimental. Multimodal benchmarks Provides comparable Vane, Ray Data, and Daft pipelines covering audio transcription, document embedding, image classification, and video object detection. The current benchmarks use local files on a single-GPU machine. They represent a single-node environment and are not a direct reproduction of the original distributed Ray Data benchmark. 👏 Thank You to Our Contributors @kaka11chen @WangErxi @caomaocao @hubgeter @liwuhen @pollychen-lab @figurant @zy-kkk @suxiaogang223 @jingdaws @freemandealer @liujiwen-up @StanleyXu512 🗺️ Roadmap 1. Distributed Extension for Ray Runner — Implement a distributed extension compatible with the Ray runner, building upon the existing DuckDB extension architecture. 2. Native Multimodal Type Support — Add first-class native type support for multimodal data. 3. C++ Embedded Functions for Multimodal Types — Implement C++ built-in/embedded functions operating on multimodal types. 4. Distributed Lance Read/Write — Enable distributed read and write capabilities for the Lance format. 5. Distributed Iceberg Read/Write — Enable distributed read and write capabilities for the Iceberg table format. 6. Turbopuffer Sink Implementation — Implement a data sink for Turbopuffer. 7. Distributed CSV/JSON Read/Write — Enable distributed read and write capabilities for CSV and JSON formats. 8. Dynamic Batch Size — Implement dynamic batch size adjustment. 9. Merge DuckDB 1.5.0 → 1.5.5 PRs — Cherry-pick and merge relevant PRs from DuckDB versions 1.5.0 through 1.5.5. 10. UDF Parameter Type Support — Add support for NumPy dict, cuDF, Pandas, and Tensor parameter types in UDFs. 📎 Learn more: Vane 0.1.0 Release · AstroVela/vane 🔗 Explore Vane 🌐 Website: ⭐ GitHub:
显示更多
@yihong0618 以前他们出身藤校,我自愧不如😮‍💨; 后来他们叱诧大厂,我留下泪水🥺; 现在他们套餐全包,我彻底失败😭;
以前出身 xx 代表你,我理解 后来你说你是 xx 毕业,代表你,我也理解 但是你现在天天说一天消耗 xx 亿 token, 代表你,甚至还写在简历里,我不理解。
0
12
125
8
转发到社区
现在配合上聪明的 gpt 5.6 sol 、 agentclientprotol 和 lody ,已经完全可以依赖 @bubdotbuild 进行日常的开发工作,不用担心缺少一个合适的 ui/ux 本来之前想周末加一下 steering 支持,结果发现 @frostming90 已经做掉了🤗
显示更多
听完了这期博客我就记住了一句话:男人不要亏待自己的屁股🤣。还好我经常练臀腿:做深蹲(负重蹲 箭步蹲,罗马尼亚蹲)、爬楼、臀桥等等,要明白翘臀不是目的,只是结果,臀大肌给你的稳定性才是目的。
显示更多
0
82
102
13
转发到社区
两个字竟然能有这么多组合。。。
最近听过的最好的一期博客。
谢谢你的耐心回复。我看完了,也认真思考了,认为你说得很有道理,同时纠正我之前的观点🙏🏻。
0
18
31
3
转发到社区
我们还给出了 Neo-Corp 这个 ai native 团队/企业协作的参考范式,我们按照实践和探索,把它放在这个虚拟的公司下边 再发一下,昨天居然贴错了域名
显示更多
PyCon China 2026 报名正式启动啦!欢迎大家到现场玩! ps:讲师,赞助商,合作社区,志愿者都在持续招募中,欢迎大家一起加入!! 报名链接: 感谢各位金主 🤩
显示更多
基本上每天晚上我都会和老婆下楼喂猫,可能因为小区里有人对猫不友好,这边的猫咪都跟人非常有距离感,我们差不多固定喂了两年的三只猫,也只有一只猫让摸。 昨天我们碰到了一个让你觉得没成年的小狸花,肚子特别大,一看就是怀孕了,她居然会主动蹭人,对于我们这么被对待的野外食堂大爷大妈,感动到热烈盈眶。 我老婆赶紧在小区猫咪救助群里问她的情况,然后大家都担心她耳朵的绝育标记的大肚子是病不是怀孕,我安慰老婆说可能是在绝育的前几天已经受精了。 惊讶的是今天群里竟然高效到已经抓到猫咪送到医院,x 光查看真的是怀孕,让我们松一口气,然后知道有的无良医院收了钱只给做剪尔不真的绝育,我真的震惊了。 另外一件事儿,我和老婆发现连续两天,三花都会跑到我俩的路灯影子下边,第一天以为是巧合,第二天我才意识到,她是因为天太热了,想乘凉!! 说起来三花是我最喜欢的小公主: - 即使吃饱了,也会跑出来象征性吃一口我们准备的吃视,不想其他几只就不来了 - 我们喂完猫有时候会去倒垃圾,她每次都会护送我们往返,保持一定距离,虽然好像没看我们很悠闲,但是实际上一直保持差不多的距离,跑跑听听,可爱至极 - 可惜她从不让我们靠近摸摸
显示更多
0
16
80
2
转发到社区
deepseek flash确实强啊,来感受一下这个长程能力。agentic能力现在强太多了。 甚至他自己发现了harness里面 agent swarm 的 tool call,并且做好了拆分和安排。 这个是提示词里面没有的,它自己发现并且组合使用subagent swarm模式的。 之前测的时候,flash太蠢了。别说调用复杂工具链进行切分和安排。就是tool call都容易出问题。 截图是Deepseek + maka,
显示更多
0
82
258
17
转发到社区
这周和 @suohawking 冲刺发布有点榨干了,各种通宵有点扛不住了 明天要靠凡人修仙传胶带期 + #spiderman# 回血
0
22
20
1
转发到社区
腾讯竟然是靠 workbuddy 杀出来的