Register and share your invite link to earn from video plays and referrals.

Search results for Qwen4
Qwen4 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Qwen4
🧩 Qwen3.8-Flash-Next: A 6B-Active Preview of the Qwen4 Architecture On the same night Zhipu's GLM-5.3-Flash took over the timeline, @Alibaba_Qwen open-sourced Qwen3.8-Flash-Next — explicitly positioned as a preview of the Qwen4 architecture. Zhihu contributor Kitt在进化 argues that for people who actually run models locally, this is the more practical release of the night. The local-deployment groups he is in are, in his words, on fire. His take: the model is called 3.8, but the architecture is really Qwen4 in preview — and it borrows the best ideas from across the field. 1️⃣ Smaller, cheaper, and realistic for local deployment Flash-Next has 125B total parameters with only 6B active — far smaller than the 300B-class GLM-5.3-Flash. The API is priced at ¥1 input, ¥3 output, and ¥0.1 per cached million tokens, roughly matching DeepSeek-V4-Flash's off-peak rates. His comparison, per 1M tokens in RMB — input / output / cache hit: 🔹 Qwen3.8-Flash-Next: 1.0 / 3.0 / 0.1 🔹 GLM-5.3-Flash: 0.4 / 1.4 / 0.115 (limited-time promo rate) 🔹 DeepSeek-V4-Flash (off-peak): 1.5 / 4.5 / 0.05 2️⃣ A Qwen4 preview wearing a Qwen3.8 name The author points out this is the same play Zhipu made: GLM-5.3-Flash's architecture is also completely different from GLM-5.3. Shipping the new architecture as open weights early is deliberate pathfinding. Inference frameworks like vLLM and SGLang, plus the quantization toolchain, all need lead time to adapt before Qwen4 proper arrives. 3️⃣ An architecture that borrows from everyone The author reads Flash-Next as a synthesis of the field's best recent ideas. 🔹 Attention: Qwen's in-house QSA, built to balance throughput and speed on long context. 🔹 Knowledge: it absorbs DeepSeek's Engram line of work, packing prior knowledge into 51B of N-gram side parameters — high knowledge density at minimal compute cost. 🔹 Training: the Muon optimizer, popularized by Kimi, scheduled together with AdamW. 4️⃣ Half the active parameters, still overtaking At 6B active, Flash-Next is nearly half the size of DeepSeek-V4-Flash — yet Qwen's reported benchmarks show it overtaking that model on multiple coding and agent leaderboards. The author says he is not worried about real-world experience: the recent Qwen3.8-27B already proved itself in daily use, and this sits on the same foundation. 🔗 Full Reading: #Qwen# #Alibaba# #Qwen4# #OpenWeights# #LLM# #AIInference# #MoE#
Show more
Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight, Get your API KEY on QwenCloud! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4!👏 We can't wait to see what you build with Qwen3.8-Flash!
Show more
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. 💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe 👉 🧠 Qwen3.5-4B Q4_K_M GGUF ⚙️ Ryzen 5 7540U — 6C/12T 🧵 --threads 9 🧵 --threads-batch 12 ⚡ --prio 2 🔄 --poll 50 📦 --batch-size 2048 📦 --ubatch-size 512 🚀 --flash-attn on 🧠 KV cache: q4_0 / q4_0 🔧 --repack 💾 --mmap 👤 --parallel 1 🚫 --device none 🚫 --gpu-layers 0 🚫 KV/op GPU offload 🚫 MTP OFF Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Show more
Interesting gap in the new Qwen3.8 lineup. Qwen3.8-27B handled Cosmic Dodge game impressively one shot & polish. Qwen3.8-Flash-Next struggled with basic win/lose conditions even after multiple tries. Looks like Qwen4 still has work to do before it replaces 3.8.
Show more
Congrats to @Alibaba_Qwen on the release of Qwen3.8-Flash-Next, using the same architecture innovations as their upcoming Qwen4 model! Such innovations include: 🟠 51-billion-param N-gram Embedding to look up a table with very little extra computation, which means the embedding table can be offloaded to slower & less expensive tiers of DRAM 🟠 Gated Residual (GR): it seems like a lot of Chinese labs are now innovating on the res connections, like Kimi's AttentionRes and DeepSeek's mHC 🟠 Qwen Sparse Attention (QSA): lightning indexer to select context at micro-block granularity Glad to see great Chinese open innovations along with end-to-end model weights to show these innovations can compose well together!
Show more
Qwen is up. Come say hi😃
Qwen 3.8 Max 0902 is now live in Command Code. It's an upgraded Qwen 3.8 Max: · 2.4T params · 1M context · Post-trained on coding & cowork Available on GOAT, Pro, Max and API
QwenWork just launched - briefs in, finished docs, slides, web pages, images, audio, and video out. One platform. No tool switching.
QwenWork is now live — the all-in-one AI productivity platform built for global teams. It turns briefs into finished documents, slides, and web pages — and generates images, audio, and video in one place, so your team members don't jump between AI tools. Analyze data reports, process complex spreadsheets, and let AI deliver polished decks while you focus on the work that matters. Take your team and give it a try — let AI take care of the repetitive, time-consuming work ✨. 🔗: #QwenWork# #AI# #AIAgent#
Show more
Qwen 3.8 27B is still being seriously underrated i've been running it locally on my RTX 3090 with 24GB VRAM, and it's insane how close it gets to opus 4.8 in some areas not in raw frontend/design output -- but where qwen gets really interesting is reasoning efficiency
Show more