Register and share your invite link to earn from video plays and referrals.

Search results for multimodality
multimodality community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including multimodality
‘This special issue proposes multimodality as an integrated and versatile methodological and interpretative framework for museum research and practice, while showcasing innovative methods as a way of advancing multimodality research…’ Access issue here -
Show more
Missed any of the Multimodality Talks series organised in 2023? Catch up on the recordings online
‘Our semiotic landscape study is inspired by the framework of multimodality. We trace the industrial Tower resemiotized and remediated…’ Read the full paper (open access) here -
Show more
🚨 just dropped GLM-5.3 FlashX From what I can tell via Vercel it’s basically a really fast version of 5.3 Flash: • up to 200 tok/s • 1M context • native multimodality They’re apparently running this on 100,000 Chinese chips too, so pretty cool
Show more
🚨 DeepSeek V4.1 Flash API beta is live DeepSeek are testing an intermediate build with: • Model ID: deepseek-v4.1-flash-expires-on-0910 • new architecture + native multimodality • stronger + faster • It's already callable right now through their API From the slug this looks like a very short test before the real V4.1 release, looks interesting imo
Show more
🔥 Big News|Kimi Pre-IPO Allocation Now Live — Start from Just 100U 📈 Why Kimi? Kimi K3, just released on July 16, features 2.8 trillion parameters, native multimodality, and a 1M token context window — placing it among the world’s top-tier AI models. Limited Pre-IPO allocation available — first come, first served. 👉 Check details: 👉Subscribe here:
Show more
🚨 GLM-5.3-Flash just landed on LobeHub, and the architecture is kind of insane. It has roughly the same total parameter count as GLM-4.5: 355B → 320B total params But under the hood: → Active params: 32B → 18B → Layers: 92 → 45 → 30T-token multimodal pretraining corpus → 1M context window → Native multimodality + tool calling Basically, @Zai_org cut the compute almost in half and somehow made the model stronger. That’s why GLM-5.3-Flash can deliver frontier-level capability at a dramatically lower cost. This is what the next model war looks like: not just bigger models, but much more intelligence per dollar. Now available on LobeHub. ⚡
Show more
⚔ GLM-5.3-Flash Aims Straight at the Post-Hike DeepSeek Zhipu's GLM-5.3-Flash — revealed this week as the anonymous "Ox Alpha" — has been open-weighted and priced at roughly one-tenth of GLM-5.3. Much of the early discussion compares it to DeepSeek's V4 Flash, which recently raised prices. Zhihu contributor 起步十档, who ran Ox Alpha inside real workflows before the reveal, gives a practitioner's verdict in one line: it is built to kill the post-hike DeepSeek. His case rests on three legs — performance, token efficiency, and an architecture change that makes the price possible. 1️⃣ It clears the bar for long-horizon work Official scores put Flash between Grok 4.6 and GLM-5.3, and clearly ahead of DeepSeek V4 Flash. In the author's own testing, its frontend ability roughly matches an early gray-test build of DeepSeek V4 Pro, while its backend is noticeably weaker than GLM-5.3 — but still usable on long-horizon tasks as long as the connection holds. His rule of thumb: any model past the Claude Opus 4.6 line is workflow-ready for long tasks. Beyond that, differences come down to reasoning style and accuracy, not viability. One caveat he flags: during the anonymous test the deployment was unstable, and some believe it served a mid-training checkpoint rather than the final model. 2️⃣ The real weapon: token efficiency Comparing peak API prices against the post-hike DeepSeek V4 Flash, the author notes cached input is actually 2x more expensive, while regular input and output sit at roughly 26% of DeepSeek's price. Since cached input is often the bulk of the bill, he wants real-world tests before calling a winner on price alone. But his own usage points the same direction. In one to two hours of real work — reading and writing files, running tests — Ox Alpha burned barely over 100K tokens. He estimates DeepSeek would need 250-300K for the same workload. His prediction: same tasks, run on both APIs, will come out cheaper on Flash — with clearly better performance. 3️⃣ His unexpected advice: skip the Coding Plan The plan only triples your quota for Flash. Using Zhipu's own best-case math — maximum usage, off-peak hours, the official 0.8x API-equivalent rate — the plan works out to about 53% of pay-as-you-go API cost. Since that scenario is already extreme, he concludes the API is the better deal for almost everyone. 4️⃣ The architecture change behind the price From GLM-5 through 5.3, Zhipu used DSA — essentially an optimized full attention — which costs more than DeepSeek's CSA/HCA, Kimi's KDA, or Qwen's linear-global hybrid. That is why GLM used to be pricier than larger DeepSeek models. Flash is the first GLM to switch to linear attention plus an HCA-like compressed attention, trained with the HCA recipe as well. That brings it in line with mainstream domestic practice — and the price fell accordingly. The author expects a future GLM-5.5 can scale up without costing much more than 5.3. On top of that, Zhipu added native multimodality and leaned on domestic compute, which he reads as the reason for the generous free quotas during the anonymous test. 5️⃣ Zhipu is still the team to beat The author's closing line is unambiguous: Zhipu remains, in his words, the number-one Chinese model company. That is his judgment, not a benchmark result — but the cost argument underneath it is now easy to check yourself. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #DeepSeek# #LLM# #AIInference# #TokenEfficiency# #OpenWeights#
Show more
Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands at #4# on the Artificial Analysis Intelligence Index (54) behind only Anthropic, Google, and OpenAI (all 57) Key takeaways: ➤ Increase in performance on agentic tasks: @Kimi_Moonshot's Kimi K2.6 achieves an Elo of 1520 on our GDPval-AA evaluation, which is a marked improvement over Kimi K2.5’s Elo of 1309. GDPval-AA is our leading metric for general agentic performance, measuring the performance on knowledge work tasks such as preparing presentations and analysis. Models are given code execution and web browsing tools in an agentic loop via our open source reference agentic harness called Stirrup. This continues Kimi K2.6’s strength in tool use, maintaining a 96% score on τ²-Bench Telecom, placing it among other frontier models in this category. ➤ Low hallucination rate: Kimi K2.5 scores 6 on the AA-Omniscience Index, our knowledge evaluation measuring both accuracy and hallucination rate. This score is primarily driven by a comparatively low hallucination rate of 39% (reduced from Kimi K2.5’s 65%), indicating a greater capability to abstain rather than fabricate knowledge when the model is uncertain. Kimi K2.6’s low hallucination rate places it similarly to other models such as Claude Opus 4.7 (36%) and MiniMax-M2.7 (34%) ➤ High token usage: Kimi K2.6 demonstrates high token usage, but is in line with other frontier models in the same intelligence tier. To run the full Artificial Analysis Intelligence Index, Kimi K2.6 used ~160M reasoning tokens. This is slightly lower than Claude Sonnet 4.6 (~190M reasoning tokens) but much higher than GPT 5.4 (~110M reasoning tokens). ➤ Open weights: Kimi K2.6 is a Mixture-of-Experts (MoE) model with 1T total parameters and 32B active, same as the previous two generations of models Kimi K2 Thinking and Kimi K2.5. Kimi K2.6 again pushes the open weights frontier in intelligence. ➤ Third Party Access: Kimi K2.6 is accessible through Moonshot’s First Party API as well as third party API providers Novita, Baseten, Fireworks, and Parasail ➤ Multimodality: Kimi K2.6 supports Image and Video input and text output natively. The model’s max context length remains 256k. Further analysis in the threads below.
Show more
0
30
1.3K
130
Forward to community
The word ‘Multi-modality’ implies a focus on the integration of multiple semiotic modes or resources (such as gesture, image, language and music) in communication and social interaction… Full paper available (open access) here -
Show more