Register and share your invite link to earn from video plays and referrals.

Florian Brand
@xeophon
evals @PrimeIntellect | open models @interconnectsai
843 Following    17.7K Followers
the paper is such a great read, you should def go through it right now! also: they want to open source 7K RL envs 👀
His analysis uses data from Q2 or even Q1 for open model providers Token numbers for TogetherAI has >10x'ed from Q1 to Q2 (he uses Q1 numbers), Fireworks has 3x'ed, Baseten is as big as the other two. OR has >5x'ed. So he is using severely outdated numbers for open providers
Show more
a lot of people would stop talking about emergent phenomena if they actually looked at the data
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
Show more
0
227
4.2K
485
Forward to community
can’t believe one of the hottest things in ai rn is called jeff
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission?
Show more
0
27
1.1K
215
Forward to community
I want a model to be intelligent in ways that allow us to deeply participate in our work. Intelligence in service of effortful human + machine collaboration. Try here: (entire image is just javascript)
Show more
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
Show more
0
419
9.9K
976
Forward to community
people are weirdly confident about that take as if closed labs haven’t ramped up their data efforts in life sciences ~ end of 2025 let’s not forget about GPT-Rosalind shall we
We’ve done a ton of performance work to supports a ton of threads running hundreds of subagents at once. I run 30-300 agents at once on my Mac and these performance improvements are crucial to not completely blow up my system More to come!
Show more
prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself.
who could’ve seen this coming
I was not expecting things to go this way, but I think MCPs are better than CLIs for most integrations. The models have gotten much better at tool calling, we can defer tools & MCP is now stateless. If you need to compose/filter data, add params like query to your MCP tools.
Show more
DS V4.1 really loves to look at things. It’s so happy to not be blind anymore Which makes it even more infuriating that some providers don’t ship with vision enabled
On the public TerminalBench traces, Astra runs into safety classifiers and their rollouts get stopped and scored as 0 Fable gets re-routed to Opus instead and the run continues. Hard to compare scores that way!
Show more
i like Gemini for this use case over a lot of models, incl. Muse Spark guess its now the Claudish translate model, oh well
asking Gemini to translate swarm<->english is surprisingly effective
i cannot in good conscience recommend @nahcrof anymore, i used them a lot in the past, they were an absolute steal, exceptionally reliable (for a small inference provider), atleast to my knowledge up until a few months ago they were entirely fine post Kimi K3 launch things seem to have shifted and this feels like a rugpull, i've independently confirmed the tokenizers don't match the model ids, there is an extremely detailed blogpost by @KTibow that goes further into this linked below, where all the evidence seems to check out
Show more
big sandbox doesn't want you to know this but you can bake dependencies into the images you don't have to use python:3.11-slim and download at runtime
because agents are so good at managing other agents, the best form factor for coding aren't a bunch of projects but just a chat box (again)
why do i, the evals guy, have to be dragged into this and now have to debug the parsing 😔
i showed the inference guys the new arch and they blocked me and reported me to hr for harassment
Anthropic claims to have disrupted an online mass surveillance system from China, gathering global intel on religious groups that Beijing deems 'suspicious'. Incidentally we tracked it together with @nestedintel weeks earlier. Meet BABEL - Here is the playbook. 1/8
Show more