Register and share your invite link to earn from video plays and referrals.

Search results for qwen36
qwen36 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including qwen36
Qwen3.7-Max and Qwen3.8-Flash from @Alibaba_Qwen are now 40% off on serverless through September 30. Use Max for long-horizon agent work and Flash for high-volume, cost-sensitive workloads. Get started today!
Show more
Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high throughput and 180 output tokens/s/user at low latency, across tuned PD configurations on @nvidia GB300 NVL72. Workload: 8K input / 1K output. Drawing on lessons from trial and error, we walk through the tuning decisions step by step: budget KV cache, benchmark prefill and decode separately, then choose topologies and MTP settings for each serving target. Deployment configs are included so you can reproduce the results. Great work from the @NVIDIAAI contributors and the vLLM community! Explore the frontier:
Show more
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model 👇 ⚙️ model-specific kernels 🧠 hardware-aware memory planning 🚀 DFlash2 speculative decoding 💾 prompt-cache reuse 👥 continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: 🚀 Splash: 74 tok/s ⚡ oMLX: 38 tok/s 🐌 Ollama: 24 tok/s At 32K context: 🚀 Splash: 54 tok/s And with 4 concurrent requests: 🔥 170 aggregate tok/s vs 43 tok/s for oMLX Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max. And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime. Requirements 🍎 M3 or newer 💾 36GB minimum 👍 48GB+ recommended Completely different performance because the software stack is optimized around what it is running. ⚠️ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test. 🔗
Show more
qwen3.8-omni-flash is now out... Alibaba’s cheap & fast omni model.. one brain that takes text + image + audio + video and answers in text,, about 25% up vs the last Omni they also map it as the Gemini 3.8 Flash class swap for audio/video understanding
Show more
Qwen3.8-27B NVFP4 variants are very close in accuracy. The real differences are memory and speed. If you're VRAM-limited, minima-ai/mnma_qwen3.8_27b_nvfp4 is a good pick, but it doesn't include MTP for faster inference. NVIDIA's version has MTP, but in my long-context coding tests MTP-4 is only ~2× faster than no MTP, and still ~2.5–3× slower than Unsloth (RTX Pro 6000). A likely reason: NVIDIA quantizes lm_head to NVFP4, while Unsloth keeps it FP8. Since MTP shares the target model's lm_head, this can hurt prediction quality and acceptance rate. So my pick is Unsloth NVFP4. Now testing accuracy on long-horizon agentic coding.
Show more
Qwen3.8 27B: mini swe agent vs Claude Code vs Pi What I learned With enough iterations, boosting scores on agentic coding benchmarks is relatively easy. On DeepSWE 1.1, Qwen3.8 scores poorly with vanilla Pi. But a benchmaxxed Pi setup beats (Reward) the Qwen team’s published result using Claude Code. Also: experiments with thinking low show that it spends more tokens, more turns, and score lower than medium. Full details, including an ablation study and token-efficiency analysis:
Show more
Qwen3.8-27B Unsloth GGUF is now the #1# most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: Guide:
Show more
0
73
1.7K
124
Forward to community
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. 💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe 👉 🧠 Qwen3.5-4B Q4_K_M GGUF ⚙️ Ryzen 5 7540U — 6C/12T 🧵 --threads 9 🧵 --threads-batch 12 ⚡ --prio 2 🔄 --poll 50 📦 --batch-size 2048 📦 --ubatch-size 512 🚀 --flash-attn on 🧠 KV cache: q4_0 / q4_0 🔧 --repack 💾 --mmap 👤 --parallel 1 🚫 --device none 🚫 --gpu-layers 0 🚫 KV/op GPU offload 🚫 MTP OFF Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Show more
Qwen3.8-27B Uncensored Cyber model you can run 15 GB locally - 528 real tool calls - 262K context. - Aggressive abliteration to Cyber peel - Imatrix from real agent logs not wiki text. - Real tool-call calibration. -
Show more
Qwen3.8-Flash is now available in OpenCode Go 125B/6B · 1M context · multimodal
0
63
1.6K
64
Forward to community