Register and share your invite link to earn from video plays and referrals.

Jeremy Howard
@jeremyphoward
🇦🇺 Co-founder: @AnswerDotAI/@FastDotAI ; Prev: Professor@UQ; @kaggle founding president; founder @fastmail/@enlitic/…
6.9K Following    328.4K Followers
At least n=2 but it’s not a lot! (synbio for two decades: dna synthesizers, sequencers, cell engineering, viral design and have been scaling transformers at GDM since 2018.) I agree with all of your points - thanks for making them. The biorisk pandemic stuff is straight out of the crackpipe from hucksters who clearly never worked in a lab. No appreciation of the timescales, biophysical complexity, supply chains, or economics.
Show more
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
Show more
0
573
14.4K
1.9K
Forward to community
"The models that are actually causing issues right now are all closed weight American models." -- @mitsuhiko
This is the type of post I might regret later, but I really had an urge to rant about the current discourse around pacing the frontier.
It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs. So today we're launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to. Let's make AI safer by making it more transparent!
Show more
0
285
3.9K
420
Forward to community
Today we're launching Base Labs, a research lab by Baseten. Our mandate is to make open-source as useful as possible, and our one rule is that we publish without exception, including what fails. Up until pretty recently I thought the way to get the world onto open models was to train them for one company at a time. @mudithj, @maxkirkby and I cofounded @parsedlabs on that bet, @baseten acquired us, and we spent the last year running their training team doing it for customers one by one. Every one of those engagements taught us something new about how models learn, forget, specialise and get cheaper, and almost none of it got spoken about. Sadly, in general that's the field's default in that the people who know the most about training have the least freedom to say it. To be clear I don't think the closed labs are the villains here. They get to new capabilities first, which buys the rest of us time to harden the world before that stuff is everywhere, and they're the ones paying to find out what's actually possible. But I'm fairly convinced the only real advantage they have is data and scale, and their incentives point squarely at the frontier. You can't do slow, public science on how these things learn when your job is the best model by end of quarter. Someone without that pressure has to, and there are very few of those someones around. Hence Base Labs. We have the broad remit of making open-source as useful as possible and our one rule is that we publish without exception, including what fails. The first problem is continual learning, which I have come to think is several problems wearing one name. We are also working on the open RL environments and data that open models need and currently can't get, because we have to aggregate data with the same ferocity everyone's been talking about aggregating compute. Plus a bunch of other stuff I'm genuinely excited about, eg a safety stack people can run on top of open deployments, and performance research so these things are cheap for everyone to serve. I still think open and closed coexist, and that's the good world. It's just that coexistence isn't free, someone has to actually do the work, and this is basically what keeps me up at night. We're hiring researchers, engineers and fellows. Come help distribute the mandate of heaven!
Show more
I'm in San Francisco this week. So is the Deputy Prime Minister, but we’re here for very different reasons. He's meeting with Anthropic and OpenAI, working on Australia’s data centre pipeline. I'm here as the co-founder of @_Firmable and a Partner at @GlitchCapital, meeting with Australian founders who moved here this year. The CGT changes came up in almost every conversation as part of the reason they left Australia to build their business in the US. What a mess. The AFR published two opinion pieces todaythat demonstrate this shambles. In the first, Treasurer Jim Chalmers set out how Australia can benefit from AI if we get it right. And that’s a big if. The headline number he was spruiking was $150 billion of data centre investment by 2030. Jim, data centres are the bottom of the stack. Concrete, power and imported GPUs. Enormous capex, very few jobs, and negligible compounding growth for our country. If that's only our AI strategy, we’re in trouble, as we're importing expensive chips, hosting someone else's product, and then buying the intelligence built on top of it at full retail. The bigger opportunity is in using and complementing the frontier models. Products built on top of frontier models that solve real business problems. That's where jobs are created and where the productivity Chalmers is chasing actually comes from. I’m seeing it at what we’re building at Firmable, the companies we invest in at Glitch and the founders I’m meeting in the Bay Area. In the second op-ed, @Shaun_Cartoon laid out what Chalmers is actually doing to the people who build that layer. Effective rates for founder and employee equity are doubling to 47%, which is why founders are leaving. Australia is courting the frontier labs on one hand, and taxing the hell out of the Australian founders and builders. You can't be a value-add exporter of AI when the value-adders are in San Francisco. If Australia is serious about backing innovation, restore the 50% CGT discount for all businesses. This alone will generate far greater productivity gains for our country than data centres will.
Show more
Papers on recurrent transformers: (ours): train full Tr, distill -> rec Tr train full and rec jointly, max. coherence train rec Tr w predict. obj. train full Tr, add rec latents
Show more
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index. Key takeaways ➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5 ➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34 ➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage ➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation Other model details: ➤ Context window: 1 million tokens, supporting image and text inputs as with Anthropic’s other recent launches ➤ Pricing: Fable 5.1 retains the $10/$50/$12.5 input, output, and cache write prices per million tokens from Fable 5, but cache hits have been reduced to $0.25 per million tokens, a 75% relative reduction from before that will materially reduce agentic workload costs
Show more
0
110
2.1K
192
Forward to community
RWKV-7 G1j (100% RNN) release 🙂 much better at agent/coding/STEM and everything. Demo: Weights:
my mom is using gemini on her android phone to write angry letters to companies/banks who did her wrong by law and they are all giving in, as gemini cites the law and i think this is beautiful.
0
62
5.5K
133
Forward to community
At 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and rerankers. Perhaps I'd even use listwise cross-encoders, they seem interesting.
How Anthropic's new results post would read without the PR: Claude orchestrated open-source protein design models, PXDesign, RFdiffusion, Genie, BoltzGen, from a 30k-token expert prompt and 12,500 H100-hours of compute, and designed binders against 14 of 15 targets. Hit rates of 22–35% against a 10–15% baseline, where some of those tools already report similar numbers on their own. The orchestration is genuinely impressive. But the open-source models did most of the lifting, and they came from the Baker lab, Columbia, MIT, ByteDance Seed, and most of them were already wet-lab validated before Claude touched them. Which also sets the ceiling. All these generators share a single PDB-shaped training distribution, so calling four of them doesn't diversify away the blind spot, since they fail together. The targets that worked are the well-studied ones. So the valid claim is that an agent can now drive this stack competently in the regime where the stack already works. Instead, we got this announcement:
Show more
0
75
2.3K
253
Forward to community
Regarding Sentence Transformers v6.0 update: A dense model compresses a whole text into one vector, then compares two vectors. A multi-vector model keeps one vector per token, scores every query token against every document token, takes the best match for each, and sums those.
Show more
LLMs increase the complexity of codebases. They duplicate methods, write overdefensive code against impossible edge cases, & overoptimize too early. Can further training fix this? Naur's “Programming as Theory Building” says no. -- @pol_avec 1/
Show more
I am a big fan of both Andy Hall and Alan Rozenshtein work and public commentary on AI. I am sure that they will continue to do important work as part of this new team at Anthropic, and sincerely congratulate them on what I am sure is an exciting next step for them personally. At the same time, I think both of them going into Anthropic is emblematic of a broader trend of the frontier AI companies grabbing up a truly remarkable number of previously independent voices and thinkers in AI. When massive events happen in AI - e.g. the DOW <> Ant conflict - the public, media, and elected officials will be looking for independent commentary and analysis from thinkers who are not at the frontier labs. The number of folks they can go to for that sort of independent commentary seems like it is going down precipitously. This seems like a real problem, and worth thinking more systemically about how we can create sufficiently attractive positions outside of companies - e.g. at independent nonprofits, in government - to attract the sorts of folks like Andy and Alan who are going into Anthropic. Anyway, each of these decisions individually is reasonable and not worth overly fixating on, and I do believe that both of them will do good work as part of this team (and have access to novel information that will help them do research) but zooming out I am worried that we are losing independent voices at the exact time we need more of them not fewer.
Show more
The amazingly fast (only 30M!) and extremely accurate @answerdotai ColBERT model is now supported in Sentence Transformers to make it easy to create and query embedding indexes locally directly in Python! 🥳
Show more
🚨I've just released Sentence Transformers v6.0! MultiVectorEncoder joins the family: ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation, alongside dense, sparse & reranker models. Big thread 🧵
Show more
Similar story. My company was one of the handful of cos that could go public in the end of the dotcom IPO window bc we had $170m in bookings. Accel went on a jihad to replace me as CEO bc i was 32 and had never run a public co. They literally couldnt find a single reason other than my lack of age and experience and i had the junior partner who needed to prove he had the mojo to change out one of ‘his’ ceo’s. I regret that i eventually caved since it led to such a lower $$ outcome with the grey haired HP ceo who had no ability to navigate the multiple reinventions needed. Thats why my advice to founders is 1) keep control at all costs and 2) take passive investors (like yuri milner/dst) over the VC builders who ultimately want to put their own stamp on your co.
Show more
0
52
1.8K
82
Forward to community