Register and share your invite link to earn from video plays and referrals.

Search results for METAMUSE_TABOO
METAMUSE_TABOO community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including METAMUSE_TABOO
Meta Muse Spark 1.3 is now free on OpenCode
0
314
8.9K
442
Forward to community
Uncensored Meta Muse Glimmer 30B model run locally 10.7 GB on 5090. - Claim on the card: 0/300 measured refusals. - 131k context. - image understanding Agent out. - compact vision projector + DFlash drafter included - claimed 0 true refusals on a 450-prompt harmful suite -
Show more
You can now fine-tune Meta Muse Glimmer 30B for free! 🔥 Our free notebook also supports GRPO RL training. Unsloth trains Muse Glimmer 1.5× faster with 50% less VRAM vs FA2 setups. Train locally with 24GB VRAM. Guide: Notebooks:
Show more
Muse Image is live on AI Gateway. 𝚐𝚎𝚗𝚎𝚛𝚊𝚝𝚎𝙸𝚖𝚊𝚐𝚎({ 𝚖𝚘𝚍𝚎𝚕: '𝚖𝚎𝚝𝚊/𝚖𝚞𝚜𝚎-𝚒𝚖𝚊𝚐𝚎-𝟷.𝟶', 𝚙𝚛𝚘𝚖𝚙𝚝: '𝟾 𝚋𝚒𝚝 𝚙𝚊𝚕𝚊𝚌𝚎 𝚘𝚏 𝚏𝚒𝚗𝚎 𝚊𝚛𝚝𝚜', })
Show more
Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like architecture design. (“Glimmer” is probably a wordplay on “Spark,” the more likely capable model from which Glimmer was distilled. Muse Spark is only available through Meta’s Model API, though.) Architecture-wise, here are some of the main points: 1. "Only" a 131k context window, compared to Qwen3.6 and Gemma 4, which support 2x that natively; it's reasonable, but maybe on the shorter end in the age of agent harnesses 2. It's a dense model, not a mixture-of-experts. (So, it's fairer to compare it to Qwen3.6 27B than Qwen3.6 30B-A3B.) 3. Hybrid attention with grouped-query attention (GQA) and sliding window attention (SWA); the SWA:GQA pattern is a 3:1 local:global ratio. Other models like Gemma 4, which uses similar components, have a 5:1 ratio for comparison. 4. It adopts gated attention for both GQA and SWA; gated attention has become quite common in recent months. It basically applies a sigmoid gate to the attention output to decide how much of the attention information enters the residual connection. The interesting point is that it uses relatively standard GQA and SWA rather than hybrid attention mechanisms such as Nemotron or Qwen3.6. 5. A very extreme GQA ratio: 32 query heads and only 2 KV heads; for comparison, Gemma 4 31B uses 32 Q / 16 KV in the local heads and 32 Q / 4 KV in the global heads. This means that Meta Glimmer has a very small KV cache. Overall, the probably most similar architecture is Gemma 3 27B (including the Gemma-style pre/post RMSNorm placement) and Gemma 4 31B, but with some tweaks like SwiGLU instead of GeGLU activations, gated attention, and the more extreme GQA:SWA pattern mentioned before. What stands out is its extreme KV-cache efficiency. I.e., the KV CACHE / TOKEN ratios (in BF16) are: - Muse Glimmer: 52 KiB (lower is better) - Qwen3.6 27B: 64 KiB - Gemma 4 31B: 840 KiB Modeling-performance-wise, their own benchmarks show that it's mostly ahead of Qwen3.6. According to the independent composite benchmarks on the Artificial Analysis Intelligence Index, it's slightly behind Qwen3.6 (see figure below). So, a few days of using it will tell where it really ranks. Overall, it looks like a solid model, particularly for agentic workflows. What stands out most is its very low memory footprint and also pretty fast prefill and decode speed. It’s also just great to see Meta releasing open weights again :).
Show more
0
68
1.6K
230
Forward to community
Just Bloomberg and $META doing damage control after crashing the market with Meta Compute framing: Spokesperson: "Meta is still hungry for even more computing power. It is still moving forward with plans for expensive new data centers and recently inked major computing deals with $CRVW, Google, $ORCL, and others." Just dropped that in with the Meta Muse announcement, and evenn threw in the "expensive" framing with DCs to signal capex. But little late given we're likely seeing a lot of margin liquidation cascades and heavy losses from media framing earlier.
Show more
0
160
897
65
Forward to community