Register and share your invite link to earn from video plays and referrals.

Search results for SOTA
SOTA community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SOTA
Reaching SOTA By Simply Trying Hard Enough
【HIGHLIGHT*】 #SOTA# (#BEFIRST#) - #NoFlexin#' - HIGHLIGHT Remix-    OUT   TODAY    22:00 @BEFIRSTofficial #HIGHLIGHT# #THEFIRSTTAKE#
Show more
0
309
19.9K
6.6K
Forward to community
We compared 3 SOTA models. Same brief: build a casual match-3 game on the LobeHub harness. Which one actually nailed it?
New open-source SOTA on agentic coding! 🚀 Ornith-1.0-397B achieves 82.4 on SWE-bench Verified  and 77.5 on Terminal-Bench 2.1, topping every open model in its class and beating Claude Opus 4.7 on both. 🤖 📦 Four sizes (9B to 397B-MoE), post-trained on Gemma 4 / Qwen 3.5, MIT licensed and globally accessible. ✨ Notably, Ornith uses RL to generate not just solution rollouts but also the scaffold that drives them. By jointly optimizing both, the model discovers better search trajectories and produces higher-quality solutions. ⚙️ Deployable on a single 8×80GB node, with vLLM and SGLang recipes in the model card.
Show more
Tether AI gonna release a SOTA AI model optimized for edge devices in the next few days.
0
107
320
13
Forward to community
Kimi K2.6 is the new SOTA open model in Vision and Document Arena, with solid gains since Kimi K2.5: - #1# open on Vision Arena (#15# overall), +14 over #2# Kimi K2.5 (Thinking) - #1# open on Document Arena (#8# overall), +9 over K2.5 and on par with proprietary models like Muse Spark and Gemini 3.1 Pro. Huge congrats again to the @Kimi_Moonshot team on the open source progress!
Show more
Kimi is the current open-source SOTA on Artificial Analysis
Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands at #4# on the Artificial Analysis Intelligence Index (54) behind only Anthropic, Google, and OpenAI (all 57) Key takeaways: ➤ Increase in performance on agentic tasks: @Kimi_Moonshot's Kimi K2.6 achieves an Elo of 1520 on our GDPval-AA evaluation, which is a marked improvement over Kimi K2.5’s Elo of 1309. GDPval-AA is our leading metric for general agentic performance, measuring the performance on knowledge work tasks such as preparing presentations and analysis. Models are given code execution and web browsing tools in an agentic loop via our open source reference agentic harness called Stirrup. This continues Kimi K2.6’s strength in tool use, maintaining a 96% score on τ²-Bench Telecom, placing it among other frontier models in this category. ➤ Low hallucination rate: Kimi K2.5 scores 6 on the AA-Omniscience Index, our knowledge evaluation measuring both accuracy and hallucination rate. This score is primarily driven by a comparatively low hallucination rate of 39% (reduced from Kimi K2.5’s 65%), indicating a greater capability to abstain rather than fabricate knowledge when the model is uncertain. Kimi K2.6’s low hallucination rate places it similarly to other models such as Claude Opus 4.7 (36%) and MiniMax-M2.7 (34%) ➤ High token usage: Kimi K2.6 demonstrates high token usage, but is in line with other frontier models in the same intelligence tier. To run the full Artificial Analysis Intelligence Index, Kimi K2.6 used ~160M reasoning tokens. This is slightly lower than Claude Sonnet 4.6 (~190M reasoning tokens) but much higher than GPT 5.4 (~110M reasoning tokens). ➤ Open weights: Kimi K2.6 is a Mixture-of-Experts (MoE) model with 1T total parameters and 32B active, same as the previous two generations of models Kimi K2 Thinking and Kimi K2.5. Kimi K2.6 again pushes the open weights frontier in intelligence. ➤ Third Party Access: Kimi K2.6 is accessible through Moonshot’s First Party API as well as third party API providers Novita, Baseten, Fireworks, and Parasail ➤ Multimodality: Kimi K2.6 supports Image and Video input and text output natively. The model’s max context length remains 256k. Further analysis in the threads below.
Show more
Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
Show more
0
660
9.6K
635
Forward to community
Ryo Miyaichi & Sota Nakajima Talk Soccer, Music and Playing to the Crowd
Ryo Miyaichi & Sota Nakajima Talk Soccer, Music and Playing to the Crowd