Register and share your invite link to earn from video plays and referrals.

Search results for LM活用術
LM活用術 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including LM活用術
LeisureMeta (LM) price has changed by +11.11%. . The highest price was 0.03300 baht and the lowest price was 0.02760 baht in the past 24 hours on Bitkub Exchange. . Check the latest prices at: . Bitkub Exchange reminds investors to thoroughly research before investing. . Cryptocurrency and digital tokens involve high risks; investors may lose all investment money and should study information carefully and make investments according to their own risk profile. . #Bitkub# #BitkubExchange# #Token# #priceupdate# #marketupdate# #LeisureMeta# #LM# _________________________________ Website: Facebook: Instagram: Tiktok: Call Center: 1518
Show more
『🐰』 camera : FUJIFILM X-H2 lens : XF33mm F1.4 R LM WR
i added native support for @NVIDIAAI's Nemotron Puzzle 75B to mlx-lm. it now runs natively on an M2 Max 64GB: ⚡️22 tok/s 💾45.5 GB peak memory usage 📚4-bit experts + 6-bit dense + BF16 head i also fixed an annoying numerical bug in mlx-lm. outputs were subtly wrong, cosine similarity was 0.8832 vs NVIDIA's reference (identical inputs). the culprit was one dtype cast in the Mamba layers happening in a different spot than NVIDIA's. once moved, the cosine similarity improved to a satisfying level (0.999...). related PR: weights:
Show more
Hy-MT2 keeps gaining momentum. Since its open-source release in May: → 700K+ downloads 🌟 → Hy-MT2-1.8B reached #1# on the Hugging Face trending, with 30B-A3B reaching #4# 🥇 → 70+ verified product and project integrations 💻 → Broader Hy-MT ecosystem support across Apple MLX-LM, Microsoft ONNX Runtime, NVIDIA NeMo, LLaMA-Factory, and more 👯 → Real-world adoption, including real-time multilingual translation of livestream comments on Bilibili 📺 And now, Hy-MT2-30B-A3B is officially available in GGUF format—addressing one of the community’s most-requested deployment needs and making local inference easier. Ready to run Hy-MT2-30B-A3B locally? Try the new GGUF release: Explore Hy-MT2: HuggingFace: Modelscope: Github: #TencentHy# #HyMT2# #OpenSource#
Show more
Can you make a reasoning model stop looping without just cutting it off? I used an MLX-LM logits processor to penalize rumination continuations inside : “Wait…”, “Actually…”, “double-check…”, “reconsider…” No finetune, no weight changes. Runtime steering. DeepSeek-R1-Distill-Qwen-14B-4bit · GSM8K n=100: baseline: 90% acc · 524 avg tokens · 250 marker hits penalty: 91% acc · 439 avg tokens · 7 marker hits hard cap near same budget: 48% acc · 410 avg tokens · 35 missing It changes how the model gets shorter: fewer self-doubt loops, not chopped reasoning. training-free, repo:
Show more
Built on PyTorch, Ray, SGLang, and NVIDIA Megatron-LM, Miles is an open source framework from RadixArk for large-scale LLM reinforcement learning post-training. Miles uses PyTorch for models, numerics, profiling, and extensibility; Ray for orchestration; SGLang for rollout generation; and Megatron-LM for distributed training. The framework supports asynchronous rollout and training, NCCL/RDMA weight synchronization, MoE-aware rollout/training alignment, low-precision recipes, LoRA, fault tolerance, observability, and extension points for custom algorithms and model architectures. 🔗 Read more in our latest blog from the Miles Team:
Show more
oMLX hits 47 tokens per second on a base M2 MacBook Pro by offloading context to the SSD. We explore how native MLX features achieve 3x faster generation than LM Studio in our latest test.
Show more
New free learning path on the Red Hat Developer Sandbox: compress, serve, and benchmark a model with @vllm_project, hands-on in Jupyter, no GPU needed. Here's what you'll actually do: ⚡ Quantize Qwen3 to W4A16 with LLM Compressor using GPTQ. Measure the result: 42% smaller, 8.2% perplexity increase. Learn to decide if that tradeoff fits your use case. 🚀 Connect to a running vLLM server and send requests via the OpenAI-compatible API. Watch 5 concurrent requests handled in real time. See prefix cache queries increment live via the Prometheus metrics endpoint. 📊 Run a GuideLLM benchmark: TTFT, inter-token latency, and E2E latency at p50, p95, and p99. Run Hellaswag with lm_eval. Cross-reference with the published model card to make a deployment decision backed by numbers. Less than an hour to complete. Free account. Built by @cedricclyburn and Michael Santos. 🙏
Show more
4 ways to run Hermes agent without paying a single API bill 1) free cloud tiers OpenRouter’s free endpoint includes Gemma 4, Llama 4 Maverick, and Hermes 3 405B Instruct — genuinely free, no card needed. 2) local models through Ollama Zero rate limits, zero API keys, nothing leaves your machine. 3) LM Studio If you want a GUI that tells you exactly which models your hardware can run well. 4) reuse subscriptions you already have Claude Pro, ChatGPT Pro, and SuperGrok all connect through OAuth with no extra key. Layer in a fallback chain and Hermes automatically switches providers the moment one hits its limit.
Show more