Register and share your invite link to earn from video plays and referrals.

Jerry Liu
@jerryjliu0
Parsing the world's hardest PDFs @llama_index. cofounder/CEO Careers: Enterprise:
1.6K Following    81.2K Followers
The best "raw" frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more expensive while flatlining on visual recognition across complex documents. This has been the case for every frontier model including the latest OpenAI/Anthropic models - see the diagram below for GPT (since then we've also benchmarked 5.6) In the meantime, hybrid approaches like LlamaParse that blend specialized VLMs with a text engine offer better performance; our own LlamaParse accuracy has increased 15% over tables and charts. If you have document OCR needs and are thinking about using a frontier model, you might as well come check out LlamaParse! We have a full eval harness through ParseBench that you can configure over your own docs:
Show more
Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively using /loop in Codex/Claude Code * You can build long-running autonomous agent loops through multi-agent handoffs, event triggers, or….just stacks of cron jobs (?) * Almost everyone believes that no one will be reviewing code in 1-2 years. * The more interesting question is whether we’d be reviewing *anything* in 1-2 years. * AI is still a bit of a skill issue. Humans are responsible for maximizing AI output and reducing slopification. * Will human intellect provide alpha as models get better, or will the playing field be leveled? Most people think it will be leveled a bit, but there is a need for humans to provide alignment, guardrails, judgment, creativity. * The minimum amount of context you need for AI could be just the codebase with some documentation. Any research/plan files are for one-off tasks and not meant to be maintained. Having a self-organizing wiki is nice but adds complexity. If you missed this one, we’ll be hosting more dinners like this on a regular cadence! If you have thoughts on what we should talk about, let us know :) (e.g. continual learning, RL envs, competitive differentiation vs. Anthropic, etc.)
Show more
i used to do this a lot. but then i felt like i was getting progressively dumber at writing. i think there's some value in forcing your own brain to sharpen your thinking, instead of offloading it to AI what i found helpful is to manually type out my own understanding in distilled bullet points after reading the LLM output
Show more
congrats on the release!
We released Sonic-3.5 and Ink-2, the #1# streaming models for text to speech and speech to text you can use in your voice agents today. New architectures enable new frontiers for speed and quality. We're now the only provider to have #1# models for both speaking and listening.
Show more
Every enterprise organization receives and generates a massive volume of contracts. Each contract oftentimes follows a non-standardized template. The difficulty extends beyond simply digitalizing the document (OCR) to actually semantically interpreting the meaning of each term and clause. There are references between sections and to amendments. It is non-trivial to use LLMs to make sense of this information both at high accuracy and also in a cost-effective manner (e.g. you can't run Fable/Opus on every page). We see a massive opportunity to provide well-tuned document extraction workflows that can both digitalize and reason across contracts to power production systems at scale. This is courtesy of our Extract feature in LlamaParse. Check out our blog: If you have Extract use cases, come chat:
Show more
Contracts are where business commitments live, but most organizations still manage them manually, searching PDFs for renewal dates, chasing down payment terms, and hoping nothing slips through the cracks. The problem isn't just volume. Legacy OCR treats contracts like flat text, unable to interpret what it reads: the same payment term appears under three different headings, renewal conditions are buried mid-paragraph, and termination clauses span multiple amendments. We wrote up how LlamaParse solves this by: ✅️ Preserving document hierarchy ✅️ Using semantic reasoning to identify key fields regardless of how they're drafted ✅️ Mapping everything into validated, schema-aligned output Ultimately, it transforms contracts from flat PDF text into structured data your downstream systems can actually use. Read the full breakdown here:
Show more
Contracts are where business commitments live, but most organizations still manage them manually, searching PDFs for renewal dates, chasing down payment terms, and hoping nothing slips through the cracks. The problem isn't just volume. Legacy OCR treats contracts like flat text, unable to interpret what it reads: the same payment term appears under three different headings, renewal conditions are buried mid-paragraph, and termination clauses span multiple amendments. We wrote up how LlamaParse solves this by: ✅️ Preserving document hierarchy ✅️ Using semantic reasoning to identify key fields regardless of how they're drafted ✅️ Mapping everything into validated, schema-aligned output Ultimately, it transforms contracts from flat PDF text into structured data your downstream systems can actually use. Read the full breakdown here:
Show more
This is an insane release from OpenRouter, and not just because it's perfect timing. It shows that frontier models alone do not own all the points on the cost-accuracy Pareto curve for knowledge work tasks; in fact they may not be on the Pareto curve at all. The Pareto curve may be defined by a mixture of models, which any independent third-party (e.g. an AI startup) has access to but the model labs do not. It's also surprising because this feature seems extremely horizontal and is not even well-tuned for a specific task. You can prompt the Fusion API with anything. This just means that for any given workflow subset, there's even greater alpha to exploit, by hillclimbing a task-specific benchmark. The more specific the workflow, the more hillclimbing you can do. This should be pretty obvious with a practical example - if you're trying to automate invoice reconciliation at scale, you can be orders of magnitude cheaper and more reliable than "raw" Claude by tuning an agentic workflow with a mixture of models for document extraction, line-item validation, and contract matching. That alpha is what's exploitable by any company out there that's not a frontier lab.
Show more
How is the US going to deal with all the Mythos level open-weight models in 6-12 months time? Enact a full AI ban?
wrapped up the week with a great technical conversation with @jerryjliu0, CEO & co-founder of @llama_index! Filmed 4 podcasts this week. Now time to start preparing for my next guests :) Lemme know if you wanna be on our @composio podcast btwww
Show more
My mayor Muslim My bagel’s Jewish The US government took away my AI Knicks in five
Think about all the times in life you had a lead and blew it That is every single game for the Spurs in the NBA finals
Imagining a future where you have to show your US passport every day to turn on your AI 💀
Had a lot of fun talking about retrieval in the agent of agents at the Vector Space meetup in Berlin on Thursday!🚀 Together with @qdrant_engine, @deepset_ai, @cognee_ and @n8n_io, we discussed a broad range of topics, from evals to architecture decisions to hot takes... But the panel was just the beginning: the room was packed with so many brilliant people and I had many genuinely interesting conversations about how to run agents and retrieval systems in production🤖 Huge thanks to the Qdrant team for organizing this!🙌 PS: if you're doing events around agents, retrieval and how to harness/evaluate them, esp in Germany🇩🇪 and EU🇪🇺, feel free to reach out, always happy to give my contribution👩‍💻
Show more
My mayor Muslim My bagel’s Jewish The US government took away my AI Knicks in five
Fuck I had free time this weekend too 😭
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Claude models is not affected. We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible. Read our full statement:
Show more
I just fell to my knees in a Sweetgreen
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Claude models is not affected. We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible. Read our full statement:
Show more
Up until yesterday, our entire MTS team has operated under the philosophy of tokenmaxxing as much as possible on Claude Max plans. With Fable, this may no longer be possible: - One of our team members hit his limit 3 times yesterday and used the equivalent of $1.5k in 10 hours - Half of our team has hit quota limits on eng work This era of tokenmaxxing may need to be restrained - or at least have clear guardrails defined. We are concerned about running Fable at API-based billing. If every engineer starts spending tokens at levels equivalent to headcount costs, our burn rate will meaningfully increase. Just as startups are starting to bake model routing into their core product, we will have to start thinking about model routing in our core engineering usage.
Show more