Register and share your invite link to earn from video plays and referrals.

Search results for TokenEfficiency
TokenEfficiency community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including TokenEfficiency
⚔ GLM-5.3-Flash Aims Straight at the Post-Hike DeepSeek Zhipu's GLM-5.3-Flash — revealed this week as the anonymous "Ox Alpha" — has been open-weighted and priced at roughly one-tenth of GLM-5.3. Much of the early discussion compares it to DeepSeek's V4 Flash, which recently raised prices. Zhihu contributor 起步十档, who ran Ox Alpha inside real workflows before the reveal, gives a practitioner's verdict in one line: it is built to kill the post-hike DeepSeek. His case rests on three legs — performance, token efficiency, and an architecture change that makes the price possible. 1️⃣ It clears the bar for long-horizon work Official scores put Flash between Grok 4.6 and GLM-5.3, and clearly ahead of DeepSeek V4 Flash. In the author's own testing, its frontend ability roughly matches an early gray-test build of DeepSeek V4 Pro, while its backend is noticeably weaker than GLM-5.3 — but still usable on long-horizon tasks as long as the connection holds. His rule of thumb: any model past the Claude Opus 4.6 line is workflow-ready for long tasks. Beyond that, differences come down to reasoning style and accuracy, not viability. One caveat he flags: during the anonymous test the deployment was unstable, and some believe it served a mid-training checkpoint rather than the final model. 2️⃣ The real weapon: token efficiency Comparing peak API prices against the post-hike DeepSeek V4 Flash, the author notes cached input is actually 2x more expensive, while regular input and output sit at roughly 26% of DeepSeek's price. Since cached input is often the bulk of the bill, he wants real-world tests before calling a winner on price alone. But his own usage points the same direction. In one to two hours of real work — reading and writing files, running tests — Ox Alpha burned barely over 100K tokens. He estimates DeepSeek would need 250-300K for the same workload. His prediction: same tasks, run on both APIs, will come out cheaper on Flash — with clearly better performance. 3️⃣ His unexpected advice: skip the Coding Plan The plan only triples your quota for Flash. Using Zhipu's own best-case math — maximum usage, off-peak hours, the official 0.8x API-equivalent rate — the plan works out to about 53% of pay-as-you-go API cost. Since that scenario is already extreme, he concludes the API is the better deal for almost everyone. 4️⃣ The architecture change behind the price From GLM-5 through 5.3, Zhipu used DSA — essentially an optimized full attention — which costs more than DeepSeek's CSA/HCA, Kimi's KDA, or Qwen's linear-global hybrid. That is why GLM used to be pricier than larger DeepSeek models. Flash is the first GLM to switch to linear attention plus an HCA-like compressed attention, trained with the HCA recipe as well. That brings it in line with mainstream domestic practice — and the price fell accordingly. The author expects a future GLM-5.5 can scale up without costing much more than 5.3. On top of that, Zhipu added native multimodality and leaned on domestic compute, which he reads as the reason for the generous free quotas during the anonymous test. 5️⃣ Zhipu is still the team to beat The author's closing line is unambiguous: Zhipu remains, in his words, the number-one Chinese model company. That is his judgment, not a benchmark result — but the cost argument underneath it is now easy to check yourself. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #DeepSeek# #LLM# #AIInference# #TokenEfficiency# #OpenWeights#
Show more
openai has definitely unlocked some secret sauce for token efficiency from their previous models to now Astra as well
Ling-3.0-flash is now free on OpenCode inclusionAI's latest model optimized for token efficiency
0
33
1.2K
36
Forward to community
Sam Altman reveals the benchmark that may matter more than model IQ: 54% better token efficiency on agentic coding "5.6 Sol, I think, is not only the best model in the world for most people." "It's also much more efficient than other models out in the world." "So it's 54% more token efficient on agentic coding tasks and also as good or better as the other best models out there." "And we're really seeing people now start to care about efficiency, understand their spend, get a great ROI." "So this is a great step forward for us." The model race is moving from leaderboard intelligence to cost per completed task. The same agentic work at radically lower token cost changes product margins, rate limits, and how much autonomy an enterprise can afford. The next moat may not be the smartest model. It may be the system that turns each dollar of inference into the most reliable work. - Sam Altman (@sama), CEO of OpenAI, on @CNBC
Show more
just finished an hour of sexual harassment course and was thinking we should have a token harassment course given a task, what's the right model/prompt/context teach employees token efficiency, context window management, model selection and prompting skills
Show more
⚡ Stop treating intelligence and efficiency as separate. GPT-5.6 maximizes intelligence per token to deliver equal-or-better performance more cheaply and quickly. Title: How GPT-5.6 fuses frontier intelligence with frontier efficiency URL: ⚡ Overview GPT-5.6 is trained to optimize both task success and efficiency, taking a more direct path through tasks. OpenAI calls it their greatest intelligence-per-token efficiency yet. 🧩 Problem Solved Frontier models are smart, but reasoning tokens, latency, and cost are the wall in production. GPT-5.6 makes efficiency a first-class goal, pushing the performance-vs-cost tradeoff outward. 🛠 Methodology & Lineup ・Sol: flagship for frontier reasoning and long-horizon agentic work ・Terra: everyday balanced model, GPT-5.5-competitive at about half the cost ・Luna: fastest and cheapest (~80% less than Sol) On serving: improved speculative decoding gives 15%+ better token generation, and GPU kernel improvements cut serving cost 20%. 📊 Results On the Artificial Analysis Coding Agent Index, Sol (max reasoning) sets a new SOTA of 80, beating Fable 5 by +2.8 while using under half the output tokens, half the time, and ~1/3 less cost. On ExploitBench it matches Mythos Preview using ~1/3 of the output tokens. #GPT56# #OpenAI#
Show more
Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose the model is) and reasoning tokens (how much the model thinks before giving an answer). Reasoning tokens in particular offer a way for models to use compute at inference time to improve responses. Output tokens are an important determinant of both cost and time per task. Various effort levels of GPT-5.6 Sol dominate the frontier - Terra and Luna produce comparatively more tokens for any level of intelligence.
Show more
Crazy, I haven’t seen a single negative post about Astra. Not here, not on Reddit. Everyone’s praising its speed, token efficiency, intelligence, and ability to get things done. This might genuinely be OpenAI’s best release ever.
Show more
0
126
2.1K
71
Forward to community
Sam Altman on model pricing: "We want to create incredibly capable, incredibly low cost, incredibly abundant intelligence, and I think Astra is a great step forward there. The token efficiency of Astra is so impressive to me, the amount of work you can get done with a relatively small amount of tokens and the kind of the price per dollar of a task. Only a few weeks ago, we cut the price of Luna, our small model, by 80%. Our goal is to have the best price performance at every level of intelligence. All the way along the curve, we will continue to keep dropping prices dramatically...I think the opportunity in front of us is so huge we will be able to continue to bring prices down and still have huge revenue." Via @BloombergTV
Show more