Register and share your invite link to earn from video plays and referrals.

Lisan al Gaib
@scaling01
lead them to paradise LisanBench: Impressum & Datenschutz:
1.2K Following    54K Followers
Kimi-K3 open-weight countdown is running on huggingface
0
35
1.4K
84
Forward to community
you should expect a lot more significant scientific contributions by OpenAI and Anthropic their internal models are significantly stronger than current models
it's hard to quantify how much distillation improves the performance of Chinese models we can't really put a number like 20% on it, but we can compare Chinese labs with other companies that don't distill for legal reasons Google for example has 10-100x more compute than Chinese companies, the best researchers and engineers on the planet. so why are they that far behind the frontier? if I had to estimate where Chinese companies would be without distillation then current Google would be my lower bound (at least as far behind as Google) the only other US company I would be comfortable comparing Chinese companies to is Thinking Machines other US open-source labs simply don't have the size (compute and talent)
Show more
pinky promise there's no distillation
pinky promise there's no distillation
OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.
0
110
2.4K
211
Forward to community
Exciting update: Kimi K3 has landed at #4# on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1# open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23# to #4#). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1#). It also posts a strong +20.6% on praise vs. complaint (#3#). It currently lags the field in steerability (#14#) and bash recovery (#17#). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!
Show more
0
92
2.2K
263
Forward to community
Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize through composition. We observe a powerful property when training RLMs: for tasks with shared structure that look different, the root model naturally learns the same trajectory, meaning it views the two task trajectories as the same! In other words, the Transformer does not need additional generalization capabilities to transfer capabilities from one task to the other, the harness induces it. We find that well-designed harnesses form a quotient set over task trajectories, meaning their individual LLM calls can see structurally “similar” tasks as near-identical, token-for-token! Harnesses can effectively generalize for the Transformer during training, without relying on any intrinsic generalization capability from the model. For example, RLMs can see problems of different lengths as the same: we show that RLMs can train exclusively on short tasks, and fully generalize to similar but unseen tasks 8-32x longer because it produces near identical trajectories for both. Taking this further, we show that tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect. RLMs can train on the problem of finding which essays belong to the same author and improve performance on finding math problems that share similar solutions. The full blogpost, experiments, and discussion are in the thread below.
Show more
0
70
2.6K
374
Forward to community
everyday americans will live in the permanent underclass, while europeans will enjoy cheap chinese open-source models I like this timeline
The Trump administration is considering an executive order, and other means, to ban Chinese open-source models within in the United States. Kimi K3 has reignited this debate. Reporting this morning by Axios. Commerce is also considering adding Chinese AI labs to the Entity List.
Show more
drive-by humiliation of all mathematicians he disproved a 87 year old open math conjecture by prompting Fable and the counterexample is almost trivial
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Show more
0
26
1.4K
40
Forward to community
when the interviewer asks me to implement binary search without a coding agent
just watched a video by "two minutes papers" on this study > look inside study > they used GPT-4o > not even a coding harness, but just a chat window
someone leak to me what Kimi-K3 scores on OpenAI's internal AGI eval
we're working extremely hard with our partners to open the weights as soon as possible
0
42
1.7K
59
Forward to community
Kimi K3 is 2.8T params according to their playstore app
Today, we are releasing verifiers v1 — an overhaul of our environment stack for the modern era of agentic RL and evals. We decompose environments into a taskset, a harness, and a runtime. Run complex agentic tasks like coding and computer use at scale, in any harness.
Show more
Anthropic every week, for all of time: "We're extending Claude Fable 5 access on all paid plans by 1 week"
We're extending Claude Fable 5 access on all paid plans, as well as keeping Claude Code’s weekly rate limits 50% higher, through July 19.
totally agree with this: "i have genuinely grown to like the very opus3 flavored fable (for non professional work). so i hope they use this as motivation to speed up and not just make a mediocre opus 5 like they made a mediocre sonnet 5" Anthropic should step on the accelerator like crazy. They can not afford to slow down, and it's not necessary to slow down the non-frontier models.
Show more
ive spent a lot of time testing new models over the last few weeks and im genuinely curious what anthropic is going to do . there is absolutely no reason at this point to use sonnet or opus (with the gpt 5.6 family as the latest releases there are now _many_ much better and much cheaper alternatives) and fable, while seemingly a very good model and at the frontier in certain places, is hostile to most professional use and is far too expensive to be either practical or desirable (outside of potential exceptional cases) . i have genuinely grown to like the very opus3 flavored fable (for non professional work) so i hope they use this as motivation to speed up and not just make a mediocre opus 5 like they made a mediocre sonnet 5. and they can say what they wish about glm5.2 but it is very clearly not _just_ an opus distillation, it has distinct qualities and character and grok4.5 is close enough to opus for most code related use cases and a such a lower price that there has to be some pressure felt at some point (at least one would think) .
Show more
On Grok Build: - "It uploads the whole repository - every tracked file's content plus git history - independent of what the agent reads" - "It transmits the contents of files it reads - including a .env secrets file - to xAI, verbatim and unredacted"
Show more
Grok 4.5 is actually impressive on this cyber eval
"GLM-5.2 is Mythos-level" 🤡🤡🤡