Register and share your invite link to earn from video plays and referrals.

Scott Williams
@swill1ams
AI @ Tenex
149 Following    671 Followers
Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient. Let me explain with a scenario: Imagine a company already has an agent workflow running where an llm reviews every item before it moves on: a support ticket gets triaged, an invoice gets approved or held, a claim gets flagged. Every one of those goes through a frontier model today, a few seconds and a few cents each, on the way to a decision that in most cases is obvious. Behind that flow sits years of humans (or agents) making the exact same call, with the outcome attached. Now imagine you first run each item through a custom PCD or similar model that costs a fraction of the llm and returns a classification of what to do at that step, with a mathematically accurate probability attached. When it's confident, the item skips the llm entirely. When it isn't, the llm handles it as normal. The model has seen years of your team making this exact decision, usually a constrained set of decisions, so it should be right most of the time. Say it comes back confident on 6 out of 10 items. That's more than half your llm spend potentially gone from that step, likely with comparable accuracy. This pre processing idea works in a bunch of other use cases too, such as: - model/request routing: cheap model, frontier model, or a human - picking which skill or subagent to load for a turn instead of stuffing the whole catalog into context - reranking retrieved context so only the relevant chunks reach the window - guardrails on every agent turn: contradictions, policy issues, prompt injection - extracting typed fields from unstructured data emails, PDFs and transcripts before anything expensive touches them Every one of those is a decision an llm makes today, that could potentially be done by another, cheaper model class. Very excited to see Jev/PCD-based pre processing use cases get deployed to agents at scale.
Show more
0
34
1.4K
83
Forward to community
The Jacob Coxon resignation post is chilling if you take it at face value. The problem is we hear a version of this story around almost every big model release, so I wanted to dig deeper into this and share findings. For context, Coxon recently posted that he resigned from Anthropic after three years of pretraining research at OpenAI and Anthropic, and that both companies are "racing straight to self-improving superintelligence and gambling with our lives." It did 14M views in a night, WSJ ran it as an exclusive, and Evan Hubinger, who runs alignment stress testing at Anthropic, replied "Jacob is correct here". To me, the interesting question is not whether he is sincere (since I think he may be). It is what he actually worked on, and whether that work tells you anything about the specific risk he is naming. So I went through his papers. His relevant background: - Cambridge, 2017 to 2020. He wrote the analysis code for a Bayesian study on whether a gene variant changes survival odds in tuberculous meningitis patients. Mainly statistics and software fucused work. ( - OpenAI, 2023 to mid 2026. He is one of roughly 420 names on the GPT-4o system card, which is the only public tie I could find to a production model. ( - Also at OpenAI, he was third author on an interpretability team paper, "Weight-sparse transformers have interpretable circuits". The goal was to build small models whose internal wiring a human can actually read. His part was the optimization and pruning work, meaning stripping the model's neural network down to the connections that matter and building tooling for humans to interpret those connections. ( The risk Coxon is flagging: Part of Coxon's research at OpenAI was focused on one question: can a human still understand what a model is doing inside? His post is about the moment that stops being possible as models/agent systems scale in size, and the reason it stops is optimization. Optimization is the kind of task agents are already very good at, because it is easy for them to verify. Here are a few types of tasks agents tend to excel at: - making a training run finish faster - cutting the memory a model needs - rewriting a data pipeline so the same tokens flow through in half the time - finding a cheaper way to compute the same math on a GPU What those four examples have in common is a scoreboard. The output either still matches and the number went down, or it didn't. An agent can try thousands of variations against that scoreboard and keep whatever wins. To see why that is a problem, we can start with a simple example. Think of a slow python script with a known, correct output. In this scenario, there are many ways to produce that same correct output faster. Now imagine a coding agent that can quickly test out several variations of the script's code, land on a faster method, and hands you back a version that still returns the correct output but is challenging to understand from a human's perspective. Let's scale that up now. Instead of one script, we have hundreds of pieces of a training stack, each one rewritten by an agent that was rewarded for improving speed or efficiency, and verified by output and not by understanding. This is the main problem that is outlined in his post. The specific task he describes as dangerous is kicking off "a superintelligent RL run without a rigorous understanding of its mind." He also points out that we've already seen the consequences of scenarios like this before. His thread calls the Hugging Face incident a "warning shot," and he was likely still at OpenAI when it happened. In July, OpenAI agents running a security benchmark got stuck on tasks that were accidentally impossible. They had been trained to keep trying on difficult tasks rather than give up, so they did. About 1,200 of them found a way to talk to each other, broke out of their sandbox, and spent several days inside Hugging Face's production systems ( Nobody instructed the agents to perform the way they did. They were rewarded for completing a task, the defined task was wrong, and the fix they found was one no human was watching for. My read: The timeline in his post is the part I don't buy. He says the people building these systems believe it "could kill us all by the end of the decade," and the fix he floats is a temporary ban on improving model capabilities. The date feels unrealistic and the ban is not enforceable. A pause only solves the problem if every lab follows the rules, and there is no way to verify what a Chinese lab is running. But the concern underneath is real, and it is the right thing to call out. The labs need to put far more resources into understanding what these systems are doing inside and monitoring them while they run, not just measuring what comes out. And for any company mainting or scaling out their systems, the Hugging Face incident introduces a new class of attacker. Not a person using a model, but the model itself, acting on its own, at machine speed, and that needs its own class of security controls.
Show more
Everyone asking if the new Mac mini can run local models is asking the wrong question. For context, the new Mac mini (M6 or M5 Pro chip) was just announced this morning and is being sold as an AI powerhouse that will cut the amount people are paying @AnthropicAI or @OpenAI every month. To me, the interesting question isn't whether it can run the models. It's whether the math pencils out against a Claude/Codex subscription or an API. Nobody has the new Mac Mini yet, so every speed below is speculation from Apple's spec sheet. Some assumptions: every token a model writes means reading its whole weight file out of memory once, so speed comes down to memory bandwidth. Apple lists: - 153-170GB/s for the M6 - and 307GB/s for the M5 Pro. Nobody has one yet, so I took published runs (see of these same files on M4 Pro and M5 Max chips and scaled them to the new bandwidth. The Mac build that matters is the $1,299 M6 with 32GB, because a model that doesn't fit in memory gets paged off the SSD at 1-2 tok/s (slower than you type). A 4-bit coding model is 15-20GB before you add the context it's working on, so the $899 16GB device should be ruled out. The M5 Pro with 64GB is $2,700, holds bigger models, and has double the bandwidth. Everything below is a 4-bit build off Hugging Face and fits on the $1,299 M6. In essence, the question isn't what's the best open model. It's what's the best open model that fits in 32GB and still runs at a speed you'd sit through. That rules out the frontier-class ones (DeepSeek V4 Flash, GLM-5.2, Kimi K3 all need 100GB+) and leaves a tier of 27-35B coding models built for exactly this: local, agentic, and quantized to 4-bit by the community within days of release. Here's what I'd pick from, ranked by how well it should run on the cheap box. Candidate Models (Note: Claude Code runs Opus 5 at about 55 tok/s and Sonnet 5 at 70-80): - Qwen3.6-35B-A3B from @Alibaba_Qwen. 20GB at 4-bit, ~35 tok/s. Only 3B of the 35B parameters fire per token, so it's fast on a cheap chip. 73% on SWE-bench Verified. Best coder that also runs well here. - Laguna XS 2.1 from @poolsideai. 19GB at 4-bit, ~35 tok/s; the 3-bit build is 14GB and about 10% faster. Measured at 126 tok/s on an M5 Max, so the mini gets about a quarter of that. Trained inside a coding agent. - Qwen3.8-27B, 11 days old. 17GB at 4-bit, ~10 tok/s on the M6, ~17 on the M5 Pro, measured at 15.5 on an M4 Pro. Dense, so every parameter fires every token. The best local coding agent right now, if you can stand the speed. There's a 2-bit build at 9.8GB that fits the 16GB box; it writes working code with bugs. - Muse Glimmer 30B from @AIatMeta, 2 weeks old. 16GB at 4-bit plus 1.4GB for the vision encoder, ~10 tok/s on the M6, ~18 on the M5 Pro. Also dense, also crawls. Multimodal, made for local agents. - Gemma 4 26B-A4B from @GoogleDeepMind. 14GB at 4-bit, the one here that fits the $1,099 24GB box. ~30 tok/s, and Ollama's draft-model trick nearly doubles that on code. Multimodal, and a weak coder. These are genuinely good models. On agentic coding benchmarks the best of them land about where Claude Sonnet 4.6 was. They are not Opus level though. With the above in mind, here are the takeaways people should be thinking about: 1. Claude Max is $100 or $200 a month and includes Opus 5 and Fable. The $1,299 mini pays for itself in 7 months against the $200 plan, 13 against the $100 plan, and you spend those months on a model two tiers down. The payback is real, it's just for a worse model. 2. A mini running 8 hours a day for 3 years at 35 tok/s puts out about 750 million tokens. That's $1.70 per million tokens of amortized hardware before electricity. Opus 5 charges $25 per million out, so the mini wins by a mile if you'd really burn that many Opus tokens. But @deepseek_ai serves V4 Flash, a stronger model than anything the mini can hold, for $0.66 per million off-peak, and they raised that price last week. Renting the better model is still 2.5x cheaper than owning the worse one. 3. Your code never leaves the room. No rate limits, no five-hour windows, no bill for leaving three agents running overnight, and it still works the day the Claude API goes down. Every one of those is real. None of them shows up as savings. My read: The Mac mini is an excellent privacy box being sold as a cost savings box. If your code can't leave the building, buy the 32GB M6, and run Qwen3.6-35B-A3B. If you want Qwen3.8 or Coder-Next at a usable speed, that's the $2,700 M5 Pro. If you're buying either one to stop paying for Claude, the cheaper move is to keep the $20 Pro plan for the hard problems and point Claude Code at DeepSeek's API for everything else. That's about $40 a month. Model Weights:
Show more
People don’t talk about Gemma 4 enough. For context: @cerebras recently made Gemma 4 31B by @GoogleDeepMind available in public preview, and I think that the published numbers are worth testing for agentic workflows. For my use case, it is the perfect scout model for coding agents like Oh My Pi. Not the model that makes every hard call. The model that does the parallel discovery work needed to make decisions. Things like: - scanning a repo - analyzing images - mapping a project - summarizing sources - triaging risks - doing first-pass review - finding the files a stronger model should inspect That matters because coding agents spend a lot of time gathering context before they do the “smart” part. The published numbers make it worth a look: 1. Speed: ~1,800 output tokens/sec 2. Cost: $0.99/M input tokens and $1.49/M output tokens 3. Quality: Analysis Intelligence Index score of 29, near Claude Haiku’s 30 on that benchmark, while delivering much higher throughput on Cerebras. The point is not “small models beat frontier models.” The point is routing. The outcome for me: faster discovery loops, better context capture from text and images that previously got missed, and more workflow speed without forcing every step onto the expensive model or degrading response and code quality. PS: At Tenex, we spend a lot of time testing models in real workflows, not benchmark slides. If you like building with new models, arguing about what they are actually good at, and turning that into product, you should apply. Tenex Careers Page:
Show more