Register and share your invite link to earn from video plays and referrals.

Nate Berkopec
@nateberkopec
Ruby and Rails applications, made fast. 日本語: @nateberkopec_ja
118 Following    32.9K Followers
I agree with @thorstenball, I think unit tests are dead in the water. The ones the models write are terrible, at best just doubling total LOC. Inverting the testing approach - heavy e2e/black box/golden master, reaching for lower levels only if necessary, works better for me
Show more
With the numbers from a day of benchmarking @typesafeai's Jev against our production judges... What Jev is good for Picking a label from a fixed list, when the list is defined by you and the evidence is in the input. That is the shape where it matched or beat the models we run today, at a fraction of the cost and latency. - Classifying sales-call transcripts into one of 32 call types: 35 of 36 correct on a synthetic set with 20 adversarial cases, versus 34 of 36 for our current small-model classifier. Zero label changes across three repeat runs. About 100 times cheaper per call. - Typing the relationship between two knowledge-graph nodes from a fixed vocabulary: passed all four of our hard quality gates on a 47-case gold corpus, one case behind production, at roughly 65 times lower cost and 15 times lower latency. - Categorizing Slack channels into four buckets: 16 of 16 on the held-out set, tied with a frontier model. - Yes/no questions about evidence that is visible in the input: in our tests, "does this item contain a visible defect" and "is this candidate a real customer signal" both worked. It is also fast and stable. Median latency was 100 to 200 milliseconds per call. Across roughly 3,700 calls we saw nine server errors and no throttling. Confident answers were identical run to run; only low-confidence answers flipped. The confidence score is the most useful thing about it. When it was wrong, its confidence was usually low. A frontier model told us it was 95 percent confident on both of its misses. What Jev is not good for Judging quality, or anything where the right answer depends on the model's own sense of what counts as good. - Judging whether an AI agent's output is good enough to publish, against a written rubric: 45 of 50 on our labeled fixtures versus 48 of 50 for a frontier model alone. On a held-out set weighted toward data-heavy outputs it dropped to 18 of 27 with nine false rejects, while production got 24 of 27 with none. Six rewrites of the question changed nothing. On 71 real production rejections it would have let 27 through. - Judging whether two decisions contradict each other: on 36 human-labeled pairs it found 2 of the 12 real contradictions. Ten of them scored a probability of 0.13 or lower. - Code review. A staged Jev-only reviewer we tested on 14 real merged PRs, with human review findings as ground truth, found none of the 13 human-identified problems on the PRs it could process and produced 24 low-severity "compatibility" and "test gap" notes across PRs that humans approved cleanly. The pattern across all of it: Jev works when a definition outside the model fixes the answer, such as a taxonomy or a checkable condition. It fails when the answer is a matter of judgment, because its prior about what is trivial, safe, or contradictory does not match ours and no phrasing moved it.
Show more
Increasingly feels like the architecture for working with LLMs is going to involve a subagent that does the work, and then a personality hire main agent whose job it is to speak to you in plain english like the little dummy you are.
Show more
Astra: "Caveats: A photon could have hit the RAM bank while calculating this answer, and so our data could be incorrect. Butterfly wings may be flapping in India. Mercury may be in retrograde, I have not verified astrological observations yet."
Show more
"AI is shipping code people don't understandd" bro nobody understands public cloud IAM and yet it runs the entire economy we're already here
github could halve the load on their infrastructure tomorrow if they fixed this stupid "react" app having to make refresh constantly to actually see the latest data
this is actually literally insane because it doesn’t appear to have any trade-offs. it’s just better engineering. the linked blog goes into detail about the implementation: mapping to CTIDs, various bitmaps, AVX, doing all this while they keep all the postgres constraints, etc.
Show more
If you're on the fence about SFRuby Conf they just gave me a $50 off coupon to share! Keynote from Ryan at Intercom/Fin. AFAICT Intercom is maybe the most effective software factory at scale in existence today so eager to hear what he has to share.
Show more
I think agent memory is like a kind of shitty glue over the work you _should_ be doing to build a software factory manually. Building new skills, loops, bots, etc. It _kind of_ works but ultimately falls down because agents aren't as good as you are.
Show more
Current status: believing that the dots will connect.
One of our clients reduced their total time spend handling requests by 37% this week. Tidy, tidy, tidy. Always nice to be able to delete 1/3 of your instance hosts!!!
For the last 3 months I've been telling all my clients to move everything to MCPs and I saw the light when @jsharkey showed me an internal "mcp proxy" they had built that was similar, and saw big co's like Ramp adopt the same "mcp of mcps" approach
Show more
total MCP victory, some quick misc thoughts about why MCP is so much better than CLIs: - indexable tool catalog letting agents scale to unlimited tools - no requirements to be running a full sandbox - consistent auth across all MCPs rather than each CLI inventing its own auth - multi account support for all MCPs unlike CLIs - implementations like code mode let the model know what will be returned allowing for super efficient token usage unlike CLIs the reasons it took this long for MCPs to finally have their moment is mostly due to bad MCP implementations in clients: - you had to restart your whole client to use an mcp (no hot reloading) - agents weren't as familiar with debugging mcps as they were CLIs, so it was a lot easier for people to get set up using them but over the past year, things like codex and claude plugins have all been using MCP under the hood, i.e computer use is an MCP, i believe claude artifacts are an MCP app, just the silent steady adoption there is still an element of MCPs that is 'this MCP could've been an OpenAPI spec' but that'll go away as things like triggers get more adoption at the end of the day, what's important to realize is while yes there are these difference between CLIs / MCPs / etc they're all just different ways of doing tool calling, and you can do some combination of lazy loading, searchable tools, and filtering to build efficient harnesses
Show more
Interesting to see Yegge say he never successfully built anything with Gas Town. In I mentioned not finding these ultra vibed orchestrators useful b/c reliability (w.r.t. completing tasks). Turns out the author of the most famous one had the same issue.
Show more
Watching an LLM just rip all-out on an autoresearch project is fascinating. Just thinking about all the human hours it would've taken to get this font file reduced by 5%... now available to you for the low cost of $7.82.
Show more
Prompt I use a lot: "how can I pokayoke this so this kind of error never happens again" pokayoke is using simple, low-complexity methods to make a particular kind of mistake impossible to make. classic one would be connector shape: make it impossible to insert "the wrong way".
Show more
True a year ago and still true now: Agents think that the way you solve 99% of problems is Just Add More Code. Your job is to fight against that.
Here's all of the models currently in the "frontier" discussion for interactive, human in the loop coding sessions. @ArtificialAnlys data. Takeaways: 1. GLM-5.3-Flash is cheap but extremely slow (high output tokens per task) 2. Sol dominates the time per task frontier 3. Gemini 3.7 flash is very similar to Sol. 4. Opus is smart and fast, but costs far more. 5. Grok is good but a step down from Opus and Gemini. 6. OSS models require ~10-20x as many output tokens per task, which means they're extremely slow. 7. For interactive HITL work, you should pick a sub at a frontier AI lab and use it. 8. Google/Spacex/OpenAI/Anthropic are close enough that you can squint and say they're the same 9. OSS models are capable and cheap or equivalent per-task but extremely, painfully slow because they require so many output tokens to succeed.
Show more
These two charts from @ArtificialAnlys explain why GPT-5.6 Sol (med) is the king of all models for interactive coding agent sessions right now. Sol(m) is 5x faster per task than Fable, 10x cheaper. It's even faster per task than Flash 3.7, despite having ~6x less tok/sec.
Show more
I’m so sick of AI marketing copy but it seems like that’s another bit of slop we will all just learn to accept
>self improving Looks inside >two human review gates
Claude 发文推荐 Warp 这套「自进化 Agent」做法,我觉得很值得看。 Agent 每次干完活,人类正常给反馈。另一个 Improver Agent 会定期把这些反馈捞出来,看它哪里反复犯错,再对原来的 Skill 提一个小修改。 修改直接走 Git PR,人 Review、Merge 以后,下一次 Agent 就会带着这次经验继续工作。 Warp 已经把这套机制用到了 Code Review、写 Spec 和 GitHub Issue Triage。 我觉得比较有趣的一点是,他们没有把学习理解成不停往 Prompt 里塞规则。 反馈最好告诉 Agent「为什么错」,Skill 也要保持小,只沉淀真正能复用的原则。 这种自进化其实很朴素: Agent 干活 → 人纠正 → Agent 总结 → 修改自己的 Skill → 下一次少犯一次。 很多 Agent 现在缺的可能就是这个循环。 每次 Session 都积累了大量反馈,但下一次又像第一次来上班一样。 文章
Show more