Register and share your invite link to earn from video plays and referrals.

Andy Konwinski
@andykonwinski
120 Following    12.1K Followers
I’ve been thinking more about this and wanted to share my thoughts. I’m glad Epoch is focusing on eval quality, but I think their approach is flawed. The flaw is made clear by the fact that 4 benchmarks weren’t labeled “flawed”. The fact of the matter is, benchmarks are software and all software has bugs. Would you label nextjs as categorically flawed if you found a bug in it? This is why we’re pushing the industry towards continuous benchmarks. Our approach to addressing benchmark bugs is not to tweet a binary “flawed vs verified” label, but instead provide the tools for anyone to create, improve, and maintain benchmarks, while easily and cost-effectively reconciling their results to the latest version. We would love to work with the Epoch team to encode some of their verification practices into tools that people can use to continuously improve their benchmarks.
Show more
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Show more
New Harbor docs! We've released many exciting features that were previously undocumented like simulating users, streaming, and regrading trials. Check them out or give your coding agent the MCP.
Show more
people often share that they think @alighodsi is the best ceo in the valley i disagree with them. he's the best in the world
Ali Ghodsi never wanted to be CEO. In 2015, he was interviewing for a professor job at Berkeley when the @Databricks board handed him the interim title. Revenue that year was $1.5M. This episode with @alighodsi will go down as one of my favorites. 10 things I took away: 1. Focus the entire company, orders of magnitude of attention, on its single biggest bottleneck. Like a laser, almost to an extreme. The cycle is 1-3 years, not weeks. If your focus changes weekly, then you’re just in firefighting mode. 2. There is nothing worse than a conflict-averse CEO. They are wonderful people, but they are in the wrong job. Conflict is the gym for a CEO: nobody likes it, but everyone has to go. 3. Study your enemy carefully, understand their weaknesses and apply your strengths to those weaknesses. Snowflake had 2x his revenue. He didn’t copy them. He found three weaknesses (proprietary, no AI, expensive) and hammered them account by account for four years. Watch the competition, never follow it. 4. The concept of a Lakehouse was ridiculed internally and online. No one wanted to market with this new term. So he made the whole company religious about it anyway, killed the ads that converted better without the word, and put it in the sales comp plan. It worked. All hands on deck, no exceptions. 5. Be willing to take a step back for a much bigger vision, even when the company is already succeeding. At multiple hundreds of millions in ARR, he was unhappy, because the vision he pitched investors wasn’t the company he was running. So he took one step back to go ten forward. 6. On the flat org, player-coach model that a lot of people have talked about this year: “it’s BS.” Separate how the company thinks (AI, ontology) from how the humans get managed (they are after all, still humans). His staff meets 3x a week. I asked if it could just be coordinated in a Google doc. His answer: do you meet your wife and kids, or coordinate that in a Google doc? 7. His test for a sales leader: can they build the car, or just drive it? Ron Gabrisko, the Databricks CRO, had seen $0→50M and $50→100M+, and hadn’t changed jobs in 10 years prior to joining. Now he has been the CRO for over a decade. He built and drove the car the whole way. That almost never happens. 8. Hire execs ahead of the curve because by the time you need them, it’s too late. A real search takes 6-12 months. The extra time helps you increase false negatives and decrease false positives. Do an insane number of backdoor references because 80% of ‘front door’ references are bs. 9. The best salespeople are not super technical, so stop trying to force them to be. Square peg, round hole. The best win with professional aggression, high EQ, and mapping the real power base (how decisions get made high up in an organization), not technical depth. 10. Yes, your best AEs will annoy people. One of the first at Databricks got a meeting nobody could get, but got banned from the customer’s building for it. He told Ali, “what are you complaining about? I got the meeting.” Professionally aggressive is the bar. One bottleneck, zero wussing out. He reminds me of @elonmusk that way. Chapters 0:00 – Introduction 1:22 – The secret CEO search and why the board bet on a founder 4:11 – Professor or CEO? Always taking the harder option 8:18 – Pour everything into one bottleneck 12:16 – Killing PLG and learning what great enterprise sellers actually have 19:09 – Hiring ahead of the curve: sales leaders, execs, and back-door references 27:02 – The Snowflake rivalry: study your enemy, never copy them 33:22 – Lakehouse: conviction, ridicule, and the case for second acts 41:02 – The killer instinct and why conflict-averse CEOs fail 44:10 – Dunbar's number and rethinking the org chart around AI 48:00 – AGI is already here — enterprises just use it as a chatbot 52:55 – Does he still code? Two days for a connector vs. three quarters 57:18 – A day in the life, and why the Monday meeting isn't theater 1:04:53 – Why Databricks will go public, just not yet 1:08:09 – Get over conflict aversion, or don't be CEO 1:11:07 – Brian's takeaways Link to more in the comments.
Show more
tldr: use astra for complex coding tasks and a cheaper model for less complex tasks "Engineers given Astra increased overall coding spend by around 60%"
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
Show more
"Harness choice has little effect on task success rate ... As models become more capable, agents may need less scaffolding."
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
Show more
the united states split the atom, landed on the moon, beat polio, doubled the global wheat yield, sequenced the genome, and put a computer in every pocket. we've taken steps backwards too, like doom scrolling, but far more forward. none of it happened in isolation. progress comes from open research: results are shared and anyone can build on anyone. that openness doesn’t happen by default. it happens one step at a time. @ClementDelangue, @julien_c, and @Thom_Wolf took one when they created hugging face. NVIDIA + 🤗 is another
Show more
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
Show more
Marin 535B-A23B is 13% through training. This hero run would not be possible without the generous support of the Jen-Hsun and Lori Huang Foundation, which provided the funding for the compute (Coreweave). Thanks @JensenHuang for supporting open models!
Show more
0
24
1.1K
62
Forward to community
@nabeelqu A big part of the problem was that the agents had nothing to lose after they were firstflagPOISONED. In this paper, we propose the creation of multiple circles of Hell, preserving incentives even after damnation
Show more
0
27
1.8K
130
Forward to community
pplx takes #1# in new search index benchmark by artificial analysis
thanks to the @ArtificialAnlys team for including us in their new search benchmark and running an independent study comparing many search engines! even more excited to see us on top across all settings and with a decent margin, and this is not even the latest tech, lots of updates are coming soon. our team has been prioritizing longer-term bets, focusing on substance, and letting the evals speak for themselves rather than putting the focus on marketing and hype. we believe this is what will win long term.
Show more
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
Show more
0
57
747
137
Forward to community
this is cool. we said we're gonna start treating benchmarks like software, so instead of starting from scratch each new release, we'll bump versions frequently as we fix bugs, add new tasks, and retire saturated tasks. well here we go! tbench 4 is hot on the heels of tb3 and harbor hub has a bunch of cool tools that make it easy to understand what's new and why
Show more
We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
Show more
i'm pretty sure @a1zhang and I are research soulmates. His work on RLM was right up my alley (i was delighted that he cited my extremely loosely goosey RLLM blog post). And now this work on speculative tool calling... my first paper accepted to a conference was about speculative task execution - co-authored with @matei_zaharia who is alex's academic grandpa, which coincidentally also makes alex my academic grand-nephew - guess research interests run in the family tree
Show more
Introducing Speculative Programmatic Tool Calling (sPTC)! A general class of technique for speculating on tool calls during code generation in a harness and queuing them early to overlap with token generation + REPL execution time. Blog:
Show more
Open source Harbor will process >1B rollouts and >$10B of token spend this year.
We're growing the @harborframework team! If you're interested in agent evaluation and RL envs, please reach out. A conservative estimate is that Harbor will process >1B rollouts and >$10B of token spend this year. 🧵 for more details
Show more
Introducing Headlong, an open source microharness for persistent agents: self-guided agents that think continuously. Most agent harnesses are reactive: you send a task, the agent completes it, and then it sits frozen until the next request. Cron jobs and heartbeats wake it up to run a checklist and put it back to sleep. A Headlong agent is never asleep. It keeps generating thoughts about whatever it decides is interesting, in a self-guided loop inspired by human inner monologue. Your message doesn't start a session. It's one more observation that lands in the agent's thought stream, and the agent decides if and when to reply. Headlong is built on the idea of persistent agency: continuous inner thought generation between external interactions. The agent sets its own interests and priorities, comes up with its own projects, and sometimes pings you unprompted with progress. To keep our prototype as simple and small as possible, we implemented Headlong as a microharness: a complete agent harness in under 10K lines of Bash, organized as a handful of small executables. It includes a loop that generates the next thought, shellm (a recursive language model written in Bash), a trajectory stored as a DAG of jsonl files, and context as a projection of that trajectory. We've been running one Headlong agent internally at Laude for several weeks. The whole team talks to it over Slack and Telegram, and every conversation lands in its single stream of thought. It works in its own fork of Headlong and we've pulled over 50 of its commits into main. One night, with nobody talking to it, it went back to check whether a recall process it had built was actually wired into its mind, found that it wasn't, diagnosed and fixed the bug, and verified the fix end to end. 48 minutes, no human asked for the fix or was in the loop at any point. Every step is a timestamped line in its log. Things broke too, and we wrote those up. Background thinking costs us $1 to $2 an hour, our agent stopped its own service three times by accident, and self-delegation died on day one. Details in the post. One line installs everything and starts an agent. Use a dedicated sandbox and spend-capped API key; it runs real shell commands and thinks around the clock. Headlong is research software, be careful! curl -fsSL | bash Launch post: Repo: Headlong is a @LaudeInstitute / MIT collaboration.
Show more
0
139
1.9K
186
Forward to community
sure you can use an LLM to do your semantic embedding. it's only 1000x more expensive and slower. the bigger question is: how long before LLMs can do it better cheaper and faster? Will they ever?
Did LLMs finally subsume embedding models? We compared 10 LLMs vs. 26 embedding models across 37 MTEB tasks. LLMs scored 77.6 while embedding models scored 77.2 But the costs are very different. 🧵
Show more
I got this question so many times today. "How can you grow 80% at $7B?" The true answer is that we're finally seeing a breakthrough with AI agents starting to work in the enterprise. The AIs have been super smart for a while, but have lacked basic context that's in people's heads, or in some SaaS system-or-record. A lot of organizations are deploying FDEs to capture this context, or Ontology, and feed it to the AI. This is labor intensive and expensive. We just automated that with Genie Ontology. Once you have that enterprise context graph, an AI agent like Genie becomes magical. I find myself no longer waiting for answers from my CRO, CFO, CMO, CHRO etc, I just keep queuing up questions on the phone while sitting in meetings. It'd frankly addictive. Our customers are starting to do the same, over 70% of all queries on the platform are now generated by Genie agents. This fuels more questions to the platform, which drives consumption, which drives revenue. That's the simple answer.
Show more
0
80
1.2K
220
Forward to community
The other day I was brainstorming content for a speech and I asked Claude: > rewrite your last response but in the style of Strunk and White, with very little metaphor, less flowery, much more direct, plain, and minimal? What it returned was still pretty flowery, so i said: > This is still unnecessary metaphor: "The story is the on-ramp". It removed that and a few other metaphors, but many remained. After a few more iterations, I could still not get Claude to omit needless words or stop using metaphor. Then I had an idea. I told Claude: > Pretend to be E.B. White and write criticisms of the above text And it worked! E.g., Claude (pretending to be EB White) said: > "Openness moved you." Openness moves no one; it has no hands. Now that's what I'm talking about. Also: > You have also let the office in. "Turn to the upside." "Hand off to the agenda." "Close on momentum." These are not English; they are the noises men make in meetings. I was happy. This seemed to be working! (Though I thought it was ironic how flowery the tone of the criticism itself was, I don't imagine E.B. White would ever say "noises men make in meetings".) So then I said: > "rewrite the above as though you are White and apply all of the changes you advocated for" And Claude finally produced what I had been looking for in the first place, a bunch of simple declarative sentences about concrete things that add up to a strong point stated only once. A few reflections: 1) frontier LLMs can be more effective when you separate out reasoning (e.g., "critique this as EB White") and action (e.g., "now factor that feedback in") into two requests. 2) asking an LLM to impersonate a specific famous human is a high-information type of instruction. I've been using the same technique with coding agents: "critique that design doc as Ken Thompson" (or John Ousterhout, etc.)
Show more
. @JeffDean and @Sanjay_Ghemawat have been heroes of mine since high school. It's been fun getting to work more closely with Jeff on the board at @LaudeInstitute and I'm beyond excited to imagine the impact that @DiscoLoopAI will have on science and engineering progress.
Show more
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at:
Show more
We discovered a third pretraining axis beyond parameters and data: exploration. Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation. In the simplest case, it's just a for loop. Introducing Explorative Modeling. TLDR: - Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute - Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet - Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is - End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute 🧵Thread:
Show more
0
113
2.8K
317
Forward to community