Register and share your invite link to earn from video plays and referrals.

WANGRUI
@wangruipro
水电能源领域大模型智能体系统落地架构设计
537 Following    270 Followers
@SecScottBessent @AndrewCurran_ Let’s talk about where anthropic and OpenAI got its training data then
0
14
3.1K
105
Forward to community
a big step forward
Introducing Unsloth for AMD 🚀 You can now train & run LLMs on your AMD hardware • We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs • Works on Windows, WSL, Linux • Train Qwen, Gemma on 3GB VRAM GitHub: Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference. Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models. 🔗Blog + Guide:
Show more
I feel like I could write anthropic's next blog post for them.
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model Key results: ➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1# position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation. ➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality. ➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params). ➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04) ➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores. ➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities Other model details: Context window: 1M Size: 2.8T total parameters Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens. Modality: Native multimodal input supports text and images, and the model remains text-only for output. Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Show more
We trained and released DSpark speculators for Kimi-K2.6 and Kimi-K2.7-Code on @huggingface, with native serving support in @vllm_project. Across six benchmarks in our batch-size-1 evaluation: Kimi-K2.6: 2.55× average throughput (+155%) Kimi-K2.7-Code: 2.36× average throughput (+136%)
Show more
DeepSeek released DSpark A speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification to accelerate LLM inference
SGLang now supports DSpark, enabling confidence-driven, variable-length verification for speculative decoding 🎉 DSpark addresses a key bottleneck under load: instead of verifying every draft token, it verifies only where the draft model is confident, so the gains hold even as batch size scales. We heavily optimized variable-length verification in SGLang. Across batch sizes 1 to 256, DSpark gives the best throughput/latency tradeoff on DeepSeek-V4-Flash, ahead of both MTP and non-spec. At high concurrency, dynamic scheduling provides up to ~20% higher throughput compared to a fixed budget, while maintaining high verification quality across workloads. With fused kernels and zero-overhead scheduling, DeepSeek-V4-Pro reaches 383.7 tok/s at B=1 on B300. DSpark is now available in SGLang with support for Qwen3 and DeepSeek-V4. Thanks @deepseek_ai for open-sourcing! Blog with full technical details and commands to run below 👇
Show more
spacex make it looks like daily mail-package delivery👍
Our 17th Transporter rideshare mission is targeted to launch tomorrow from California and will deliver 81 payloads to orbit →
open for everyone!
🚀 @deepseek_ai's DSpark speculative decoding now runs natively in vLLM! What it is: a semi-autoregressive drafter that proposes several tokens in parallel with non-causal sliding-window attention, then verifies them in a single pass. Output stays identical, decoding takes fewer steps. How vLLM runs it: it reuses the existing SparseMLA backends instead of custom attention kernels, captures the full draft backbone and sampling loop in one CUDA graph, and works with prefix caching and FP8 KV cache. Performance on DeepSeek-V4-Pro-DSpark (verified on NVIDIA 8×B300 GPUs): - ~250 tokens/s at batch size 1 - average acceptance length ~5 - 12-42% higher acceptance than MTP across draft depths Run with vLLM nightly today: vllm serve deepseek-ai/DeepSeek-V4-Pro-DSpark -tp 8 --trust-remote-code --kv-cache-dtype fp8 --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}' DSpark Core PR: Thanks @deepseek_ai for open-sourcing DSpark, and to @NVIDIAAI and the vLLM community for landing it! 🙏
Show more
FT: OpenAI has proposed giving Washington 5% of its $852B business to ease AI pressure. The idea borrows from Alaska’s oil fund, which shares resource wealth with residents. Here, the resource is not oil, but future income from advanced AI systems. OpenAI also wants other major AI companies to give similar 5% stakes. Anthropic, Google, Meta, and others have not agreed to join this plan. No deal exists yet. The mechanism would likely be: OpenAI gives shares to a government-linked fund, that fund holds them, and future IPO gains or dividends support public payouts. The hard part is legality. The legal route is unclear, and a deal may need Congress, especially if the government creates a formal public fund. The Intel deal made this idea less theoretical after taxpayers received a 9.9% stake. OpenAI has already proposed a public wealth fund giving citizens AI-linked financial upside. Shareholders matter a lot here. OpenAI Foundation owns 26%, Microsoft owns about 27%, and employees plus other investors own 47%. A new 5% stake could dilute everyone unless the shares come from an existing holder. So OpenAI’s board, Foundation, Microsoft, major investors, and maybe regulators would need to accept the structure. The cleanest path would be non-voting shares placed in a public wealth fund, so the government gets upside but not control. The messiest path would be voting shares, because then Washington becomes both regulator and part-owner.
Show more