ICLR: first time?
Maybe it’s time for
@Steam to consider some quality control again.
700+ games released in ONE WEEK.
Small, weird & amateur games are part of what makes Steam great.
But with mass-produced, barely functional games, maybe "anything goes for $100" isn't sustainable forever.
Show more
No it's not, you need calibration to make it useful for decision making. Next time you go to ICML/ICLR/NeurIPS, do not walk away from those calibration posters.
LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface: discrete options in, fast probabilities out. Our own
@ekzhang1 made it better at the job with a $5, 10-minute run on Tinker.
Show more
ICLR 2027 submission: we benchmarked all frontier models and our model performs...
Anthropic & OpenAI today: no you don't.
ICLR deadline is 9/25, so i guess no excuse not including opus5.5 and gpt6-luna & sol in the paper.
Show more
wow, Opus 5.5 looks very promising! On my CRPG benchmark, Opus 5.5 takes only 40 steps to get out of the spawn house and onto the world map, much faster than Fable5.1. Let's see if it can speedrun past gpt6-astra. 4-hour gameplay livestream starting now:
Show more
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Model-specific inference engines are basically what everyone autoresearches and hill-climbs the hell out of for TPS after each Qwen release. In the early days I worked with MLX too, but these days I just trust oMLX to do the job.
Show more
Meet Husky: a Model-Specific Inference (MSI) engine up to 4.5× faster than Apple's MLX
Woof, Underdog's Pareto frontier model, now runs up to 730 tokens/sec on a MacBook
Finally local models are as fast & capable. Try it now in - your personal private AI
Show more
yeah this actually looks really bad for zhipu/glm given the recent situation, so far no response from their official account neither Zibo
@ZixuanLi_ or
@louszbd respond, probably their first crisis PR.
Show more
thoughts on Typesafe/Jev: I'm surprised that a general-purpose classifier (or discriminative model) can be just as interesting to the public as a general-purpose generative model. I've worked on search and discriminative models for years, so maybe I should have called this upfront, tbf I never thought about making them general-purpose beyond search. Typesafe also positions Jev really well as "System 1", a complement to the generative models we already have. Humble, unassuming, and sidestepping the frontier-lab warzone.
Jev's success could be disruptive to today's agentic systems in many ways, such as tool-calling and routing patterns. Over the last three years the consensus was to use a small generative model for routing, tool calling, MCP etc. Time to pause and rethink such architecture, because we'll probably see tool calling and routing move back to discriminative models. Whoever ships the next Jev-like open-weight base model will likely win the community.
Show more
timeline today
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Show more
Super excited about this new model! We don't ship a lot of generative models, so every time we release one I'm hyped. On M3 Ultra I managed to hit 600 tps aggregated at decoding with swift-metal, available in omni-macos:
Show more
Announcing jina-ocr-v1, our new visual document parser with 3.4B total parameters and 570M active parameters, with speculative decoding built in. Throw PDFs, scans, tables, charts, or invoices at it and get clean markdown back. Available on 🤗 & Jina Reader `x-respond-with` today
Show more
Anyone else notice union-alpha really likes android codebases for some reason? It brings up android stacks and repos out of nowhere.
1. union-alpha is not from xiaomi
2. bold & open to livestream the entire training pipeline
3. should send this gif to cfo
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run:
Show more
use union-alpha in pi:
$ pi install npm:pi-openrouter-realtime
then do:
/openrouter models-sync
not sure it’s even remotely related but back in 2024 I have this weird post about chain-of-thought classification using embedding model to achieve some kind of test time intelligence
Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Show more
people are sleeping on this release. I finally got a chance to test it today: qwen3.8-flash-next-q4 now runs at 50-70 tps on long-horizon tasks on M3 ultra, with no regression in quality or thinking. It used to be 20-30 tps. Very impressive work!
Show more
Still remember back in 2000 many software hardcopy have this shinny hologram sticker on the box to prove they r genuine.
Every time I see a small LM or encoder-only initiatives, my first guess is that they must from Europe. honestly that's fine, it's way more credible than some eu AI companies claiming frontier this and frontier that on LLMs to get eu/gov funding while they're really just post-training Chinese open weights.
The problem of SLM startups is scaling and fundraising: the SLM story is hard to sell in the EU because the VCs there lack the knowledge (or still dreaming about AI sovereignty), and also hard to sell in the US because VCs want to throw money at categories with huge upside and get that compounding effects. imo, the upside of SLMs is upperbounded to, say, dev tools? meaning the best SLM can probably be a pretty smart dev tools, and when compounding with some network effect it can be cool, but it's still nowhere near the LLM story and its upside.
Show more
We are releasing the most general and efficient encoder model ⚡
GLiFormer can extract virtually any structured JSON from text or PDFs up to 100× faster than generative LLMs for multi-task information extraction workloads.
It also stays fast on CPU. 🧵
Show more
For startup founders this a16z podcast episode is super super interesting and worth listening to multiple times. It reveals a lot about how SV VCs actually think and operate.
teaser video made with opus5, no video-specific harness/skills, I have a low bar and I think it’s good.
this blog post is way way too long to open on the phone...
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report:
Show more
Fable 5.1 obsession with bash has gone pathological: asking me to approve a 2-page bash cmd with one "rm" buried in the middle.