가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

金融汪
@yuyy614893671
油管频道:金融汪 ,网址: 金融市场老兵,观察和理解这个世界每天在发生着什么,是我的兴趣也是工作。X上的内容只是分享,不是投资建议。
가입 November 2022
632 팔로잉 중    71.3K 팬
“全球最大的对冲基金在投资者每天进行的六项文档筛选任务上测试了 Gemini、Claude 和 GPT。朴素提示的分数大约是 50%。相当于抛硬币。专家编写的提示将准确率推高到 78%。投资者需要 80% 以上才会信任系统融入他们的工作流程,而没有一个前沿模型达到这个标准。GPT 5.4 的成本比 5.2 高出 43%,但准确率几乎没有提升。 于是他们改用 Tinker 在 Qwen3-235B 上进行微调。准确率 84.7%。比最佳前沿模型减少 29.8% 的错误。在推理成本仅为 1/14 的情况下。”
더 보기
Bridgewater just published numbers that should make every frontier lab nervous. The world's largest hedge fund tested Gemini, Claude, and GPT on six document filtering tasks its investors do every day. Naive prompts scored around 50%. A coin flip. Expert-written prompts pushed accuracy to 78%. Investors needed 80% before they'd trust the system in their workflow, and no frontier model cleared it. GPT 5.4 cost 43% more than 5.2 and was barely more accurate. So they fine-tuned Qwen3-235B on Tinker instead. 84.7% accuracy. 29.8% fewer mistakes than the best frontier model. At 1/14th the inference cost. The smartest part is buried in the middle of the paper. Their vendor-labeled training data was riddled with wrong labels, and expert labeling costs too much to run on everything. Their fix: train a model on the noisy dataset, then run it back over its own training data. Any example the model disagreed with got routed to senior investors, because either the example was genuinely hard or the label was wrong. The model's own confusion became a detector for bad labels. Prompting hit a ceiling for a structural reason. A prompt captures only the judgment an expert can put into words. Twenty years of taste about which central bank memo actually signals a rate move doesn't compress into instructions. It transfers through labeled examples. Every institution sitting on decades of expert decisions just learned that those archives can train a model that beats the frontier at their specific job. The alpha was in the filing cabinet the whole time.
더 보기