Register and share your invite link to earn from video plays and referrals.

Isaac Flath
@isaac_flath
Documents in AI Products
15 Following    3K Followers
"In practice, this all means that evals and structured experiments are even more important than they are for regular LLM projects." 💯
I've been using Jev by @typesafeai Here's the six things i've tried and am confident I'll still use Jev for 60 days from now. There's many more experiments, ideas, and things I think I will use it for. It's a big deal (more on why in next post). But I am only sharing things that I am 99% sure will lead to stuff I will still be using Jev for in 60 days. That means I started with small, boring, but useful, stuff. - Fact-checking my scripts - Ranking my news feed - Finding the right text in PDFs - Checking citations - Grouping my review notes - Figuring out why agents fail (eval over traces)
Show more
watermarks-remover for AI content reaches ~11k stars on github. People really hate the watermarks. And i fully understand them.
0
159
7.3K
497
Forward to community
For AI watermarking text (like anthropic), how does it work in the case of partial things? Example 1: I have a transcript of a talk. I want to use AI to clean it up to be readable. Remove clutter, filler words, stutters. And re-organize a bit to move all the all the audience questions to the end. Then I notice the speaker rambled a bit and said the same thing 3 times, so I consolidate it down to 1 sentence and fix the minor transitions needed for that. But basically trying to be as faithful as possible (even in word choice) to the speaker/transcript, just make is a pleasant read. Example 2: What if I write a post by hand. It's almost perfect. But I have 3 examples in the post and I want to put them in a different order. So I ask AI to put example #3# first. I might need to remove some terminology explanation from the original #1# and put in the new #1# so the reader gets all terms explained when they first encounter them. And a few minor transitions might need editing. But mostly just moving things around. Might some parts get watermarked, like the transitions or consolidation bits if there's a few different wordings that'd be fine? How much of the piece needs to be watermarked for it to identify the general post as watermarked? Is it dependent on how confident the model is in the tokens it chose?
Show more
Amp orbs are actually pretty useful. Here's an example. I started a thread in the cli working on my blog. Had it use a portal to host the site to show me the changes. I'm working on a technical post and wanted it to render jupyter notebooks nicer so it did that and I saw it work in the site portal. I moved to to the web UI because it was just kinda nice to have my website preview and coding agent both in chrome in the same app. I went back to work on my post, so I told it to make another portal for jupyter. This was nice because I don't have jupyter on this machine and I don't really like telling agents to just go install stuff on my local computer, but in an orb sandbox it's fine. It's on same machine so save of notebook reloads the preview site. And it all just worked. I could have done all this on localhost like I usually do. But it was nice because it was just a bit less friction than normal
Show more
The model "just needs encouragement" is a UX issue. It's not something to be proud of
@HamelHusain I can't believe that Anthropic's take is that word choice doesn't matter in writing....
Good engineering is about observability. Speculative decoding accelerates LLM inference, but you're running it blind. Every rejected draft token is wasted compute, but are you looking at the drafts? This weekend I wrote specspecs to solve that!
Show more
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
Show more
0
399
8K
834
Forward to community
Been trying out Buzz for the past day, rotating using it with using Codex Desktop and Herdr. Wasn't sure at first because it felt like just another multi-agent UI. But it's really growing on me and I see the power of replacing Slack with it
Show more
I dunno what's going to happen to text mediums. There's so much AI slop that's awful to read through. And so we're fighting back with automated AI detectors. Which also motivates automated de-slopifiers. And because being accused of using AI to write feels like a big insult, even people who don't use any AI are starting to change how they write just to sound less AI like. Most I know that write a lot avoid some writing patterns to sound less like AI. And with all that, i'm not sure what we should be doing. I don't really mind AI writing patterns in writing like em-dashes, or a bit of extra rhythm, or some extra contrast much. As long as it's clear enough to understand, But it is also undeniable that the presence of those signals often means there will be other more serious issues with the writing (like the core concept and ideas not being fully based on real experiences, factual errors, and those ideas being presented with more or less certainty than they should be).
Show more
I’m happy to share what we’ve been working on recently at Collaborative Computing Inc. Our first product is live - is a shared visual desktop for teams of multiple agents and people - ^^^ try it now
Show more
Final lesson in the series: How to turn eval results into a better model ✨ with @willccbb and @xeophon If you've put effort into evals, you should also consider if customizing your own model is right for you. I can't think of any better people to walk us through this, link to sign up is here: 10 am PT today, recording sent to those who sign up
Show more
More stuff in the AI enabled notebook space! I am particularly excited about this one, because I really really loved R studio and this is made by the same group. I've been following it for a while and very excited to see it stable and released 🤩
Show more
The Positron Notebook Editor is out of beta and ready for all your notebook workflows! I am super proud of the work the team did to build this notebook experience from scratch. Something we decided was necessary to have a fantastic notebook experience for data science.
Show more
This new AI Notebook IDE looks really promising from the folks at @posit_pbc (the folks who made the beloved RStudio). It's super polished and has features like allowing you to control which cells the AI sees. Elastic License 2.0
Show more
I used to hate getting voice notes from people, because I can read so much faster. But I just had a couple people send them and it honestly felt refreshingly human just having a casual honest recording of thoughts sent to me. I think i'm a convert.
Show more
Had some truly awesome multiplayer threads (shipped) and talking interactions (not yet shipped new stuff) with Amp tonight, definitely felt like the future Will stream again soon to show you all
Show more
everyone tells you to look at your data but no one really prepares you for what you’ll see