Just spoke to one of the big data labeling businesses.
Few interesting insights:
- They predict the majority of their revenue will come from Fortune 1000 enterprises, not labs in a few years
- They believe every company will want to own their intelligence, but owning intelligence does not necessarily mean using open source models
- A company’s evals will become their main proprietary IP given the improvement in agent performance after properly setting up & running internal eval environments
- Most enterprises haven’t graduated from coding agents and it’s largely due to not having the proper eval infrastructure to make non-Eng agents performant
Show more
Unhinged use of Jev, but I’m here for it…
what happens when a Pokénerd gets access to Jev
Being at a legacy company must feel like a clusterf*ck of emotions right now.
Optimism: collection of moats earned over time (brand/trust, proprietary data, network effects) that disruptors haven't yet accumulated. it's a huge unfair advantage.
Frustration: roll out of AI tools feels sluggish, turning a cruise ship on a dime feels painful, and there's not nearly enough urgency to become AI-native before disruptors eat our lunch.
Fear: it feels like AI transformation will take us 5-10 years, but do we have the luxury of that much time before we lose our market position?
Show more
37 mistakes companies make with AI transformation:
1) Not investing in your data foundation/not having a data “clean-up” strategy. Often people expect that with tools, everything gets solved.
2) Starting with “we need AI” instead of a real problem (this is true for every tech cycle ever).
3) Underresourced AI center of excellence that serves every part of the organization. Backlog builds up, employees get disenfranchised, shadow AI explodes.
4) Trying to automate the same workflow vs rethinking from scratch. Building AI add-ons to existing processes rather than rethinking processes from the ground up.
5) Thinking too big and flashy. Not considering the implications day-to-day and the value of quick, unsexy wins.
6) Over-engineering. Sometimes you dont need a full agentic system and traditional software works just fine.
7) Obsessing over cost before proving feasibility of a use case (i.e using a smaller model first before validating technical feasibility with larger models).
8) Encouraging/pushing employees to use AI without real depth. Widespread rollout with limited education/lack of training for employees.
9) Telling your people that AI won’t impact jobs.
10) Overprotecting data + spend to the point of limited experimentation from your workforce. IT/Security blocking this or slow rolling it out (which is fair but bad for the speed in which this is moving). Culture doesn’t encourage AI use.
11) Not having places to go to ask questions / knowledge share. Whether that be a skills library, shared repo, or internal AI office hours.
12) Failing to solve the last mile. Everyone’s so focused on models, but successful applied AI is a complex last mile problem: governance, data, observability, context management, people, process, etc.
13) Shipping it and call it done. Lack of discipline to go beyond the shiny demo and ensure sustained adoption that meaningfully empowers teams.
14) Slop is tolerated.
15) No governed way to build for non-technical people. No Citizen SDLC to empower SMEs to build and share production apps.
16) Assuming AI transformation is the responsibility of one person within the org.
17) Run like an IT project. No senior exec actually owns injecting AI across the business, therefore initiatives stall and leave no lasting impact. There is no clear owner.
18) CEO is not a driving force. Leadership enforcement without the leaders actually knowing how or what to enforce.
19) Not getting the buy in of the “bad guys.” Bring Legal, Finance, and IT along for the ride early.
20) Not investing in / underestimating change management. Easy to get the folks who are excited on board, but it's a long process to make others feel comfortable.
21) Not measuring baselines before any adoption. What are the metrics pre-AI tool to post AI tool? No baseline = no roi story, and thinking that all AI usage is positive ROI without measuring usage/tying it to real outcomes fails the same way.
22) Inventing new KPIs for AI instead of focusing on having AI accelerate existing functional KPIs.
23) Reducing AI to headcount and being overly stringent on ROI too early into programs.
24) Being driven by FOMO and not having the patience to treat AI transformation as the multi-year migration it actually is.
25) Being married to past purchasing mistakes and not choosing the best technology at the moment.
26) Not anticipating the complexity of getting systems to work nicely together (a kind of scope creep as the reality blows up work required).
27) Not being agile enough to change course when the landscape changes drastically.
28) Locking in to a single provider ecosystem.
29) Not providing employees access to the underlying systems needed to make AI useful to take action, not just chat.
30) Underestimating how much of an impact AI can actually have. It is both a cooler and scarier time than ever before to be an incumbent.
31) Outsourcing thinking to AI - everyone can prompt, the differentiation is how you wield the tool to multiply the work you're doing. If you have good judgement you can do a lot more. If you don't, you end up wasting a lot of tokens spinning your wheels.
32) One functional department thinking they should own AI transformation. It treats AI as a vertical solution vs. horizontal capability that’s more than just technology.
33) Executing on AI initiatives before anchoring your work in a clear strategy that’s tied to business goals, a map of key processes, understanding of your technology and data reality, and clarity around how to meet your people where they are.
34) Not solving data permissioning and RBAC considerations before rolling out agentic tools firmwide.
35) Not giving people dedicated time to experiment or carving out time in their roles for it.
36) Not understanding how a business function ACTUALLY works before trying to apply AI. In someone’s head, the process for generating some end state dashboard is simple: systems generate the data, it gets consistently transformed and warehoused, then read into the dashboard that the VP sees. In reality, it’s a complete mess.
37) Neglecting internal evals to constantly test and evaluate how new models/harnesses perform company tasks on a $ per successful task basis.
What's missing?
Show more
POV: your boss asks you to build something by next week. You've never built software before. At least not something other people have to use.
This video walks you through it, start to finish.
Here's the path:
- The stack: Claude Code builds it, GitHub backs it up, Vercel puts it live
- Build vs. buy: only build the thing that makes you better than everyone else
- Write the plan first. It's the GPS for your agent.
- Map every person who will use it, and what each one needs
- Design it in Claude Design before a single line of code
- Build the thing
By the end, you've shipped a real, working app.
Know someone who wants to build their own software? Send this to them.
Timestamps:
0:00 The ask
0:35 The stack
3:03 Build vs. buy
4:10 Why you plan first
6:00 Set up your project and GitHub
7:22 Build the plan with Claude
9:20 Feeling overwhelmed? That's normal
10:03 Map your users
11:19 Design it in Claude Design
13:21 Review and tweak the design
15:18 The finished app
This might be the nuttiest video I've made to date. Yay or nay? Let me know.
Show more
37 mistakes companies make with AI transformation:
1) Not investing in your data foundation/not having a data “clean-up” strategy. Often people expect that with tools, everything gets solved.
2) Starting with “we need AI” instead of a real problem (this is true for every tech cycle ever).
3) Underresourced AI center of excellence that serves every part of the organization. Backlog builds up, employees get disenfranchised, shadow AI explodes.
4) Trying to automate the same workflow vs rethinking from scratch. Building AI add-ons to existing processes rather than rethinking processes from the ground up.
5) Thinking too big and flashy. Not considering the implications day-to-day and the value of quick, unsexy wins.
6) Over-engineering. Sometimes you dont need a full agentic system and traditional software works just fine.
7) Obsessing over cost before proving feasibility of a use case (i.e using a smaller model first before validating technical feasibility with larger models).
8) Encouraging/pushing employees to use AI without real depth. Widespread rollout with limited education/lack of training for employees.
9) Telling your people that AI won’t impact jobs.
10) Overprotecting data + spend to the point of limited experimentation from your workforce. IT/Security blocking this or slow rolling it out (which is fair but bad for the speed in which this is moving). Culture doesn’t encourage AI use.
11) Not having places to go to ask questions / knowledge share. Whether that be a skills library, shared repo, or internal AI office hours.
12) Failing to solve the last mile. Everyone’s so focused on models, but successful applied AI is a complex last mile problem: governance, data, observability, context management, people, process, etc.
13) Shipping it and call it done. Lack of discipline to go beyond the shiny demo and ensure sustained adoption that meaningfully empowers teams.
14) Slop is tolerated.
15) No governed way to build for non-technical people. No Citizen SDLC to empower SMEs to build and share production apps.
16) Assuming AI transformation is the responsibility of one person within the org.
17) Run like an IT project. No senior exec actually owns injecting AI across the business, therefore initiatives stall and leave no lasting impact. There is no clear owner.
18) CEO is not a driving force. Leadership enforcement without the leaders actually knowing how or what to enforce.
19) Not getting the buy in of the “bad guys.” Bring Legal, Finance, and IT along for the ride early.
20) Not investing in / underestimating change management. Easy to get the folks who are excited on board, but it's a long process to make others feel comfortable.
21) Not measuring baselines before any adoption. What are the metrics pre-AI tool to post AI tool? No baseline = no roi story, and thinking that all AI usage is positive ROI without measuring usage/tying it to real outcomes fails the same way.
22) Inventing new KPIs for AI instead of focusing on having AI accelerate existing functional KPIs.
23) Reducing AI to headcount and being overly stringent on ROI too early into programs.
24) Being driven by FOMO and not having the patience to treat AI transformation as the multi-year migration it actually is.
25) Being married to past purchasing mistakes and not choosing the best technology at the moment.
26) Not anticipating the complexity of getting systems to work nicely together (a kind of scope creep as the reality blows up work required).
27) Not being agile enough to change course when the landscape changes drastically.
28) Locking in to a single provider ecosystem.
29) Not providing employees access to the underlying systems needed to make AI useful to take action, not just chat.
30) Underestimating how much of an impact AI can actually have. It is both a cooler and scarier time than ever before to be an incumbent.
31) Outsourcing thinking to AI - everyone can prompt, the differentiation is how you wield the tool to multiply the work you're doing. If you have good judgement you can do a lot more. If you don't, you end up wasting a lot of tokens spinning your wheels.
32) One functional department thinking they should own AI transformation. It treats AI as a vertical solution vs. horizontal capability that’s more than just technology.
33) Executing on AI initiatives before anchoring your work in a clear strategy that’s tied to business goals, a map of key processes, understanding of your technology and data reality, and clarity around how to meet your people where they are.
34) Not solving data permissioning and RBAC considerations before rolling out agentic tools firmwide.
35) Not giving people dedicated time to experiment or carving out time in their roles for it.
36) Not understanding how a business function ACTUALLY works before trying to apply AI. In someone’s head, the process for generating some end state dashboard is simple: systems generate the data, it gets consistently transformed and warehoused, then read into the dashboard that the VP sees. In reality, it’s a complete mess.
37) Neglecting internal evals to constantly test and evaluate how new models/harnesses perform company tasks on a $ per successful task basis.
What's missing?
Show more
Reaction to this has been incredible.
Thanks so much for the support and will try to onboard customers as quickly but thoughtfully as possible.
Im so excited to announce Tenex Agents!
Off-the-shelf, enterprise-grade AI agents focused on tackling recurring work across functions like Finance, Sales, Marketing, HR, and IT.
Everyone & their mother is building/launching agents. So how are ours different?
• < 2 weeks to deployment
• Built for responsible oversight (permissioning, approvals, audit trails)
• Optimized performance layer (model routing, evals, context management)
• Custom configured for your processes & systems of record
• Tied to agreed upon outcomes
Our agents JUST WORK because they’re a labor of love, 18 months in the making.
We started Tenex to help companies turn frontier AI research into real business results. Since then, we’ve built hundreds of custom solutions for some of the biggest companies in the world.
Along the way, clients keep asking: “What are you seeing on the ground, and what should we be doing next?”
We’ve studied the market and spoken with hundreds of executives. Three priorities kept coming up:
1) Faster time to value. See meaningful results quickly and cheaply.
2) A clearer starting point. Know which workflows to tackle and how to measure success.
3) A fit with the business. Work within existing systems, with the right access controls and human oversight.
That’s when Tenex Agents said “hold my beer.”
We’ve been quietly turned hair on fire customer processes problems into a suite of composable and battle-tested agents.
They give you a ready starting point, with Tenex handling the implementation:
• Up and running quickly. Most deployments targeted for under two weeks.
• Configured for your business. Adapted to your processes and connected to your systems of record.
• Built for oversight. Defined permissions, human approvals, and an audit trail.
• Measured against an agreed outcome. A clear definition of success for the first workflow.
We’re slowly rolling out Tenex Agents, so please fill out this interest form if you’d like to learn more:
Show more
In one day, I went from complete fucking noob to... a Jev power user? (Read this, and you can too!)
I've been playing with Jev, a new AI model from
@typesafeai. I'm not a developer, so I wanted to find out what I could actually build with it.
I made four increasingly crafty/interesting demos:
An email sorter
2,000 test emails sorted into buckets in 4.4 seconds, for about 4 cents. It flags uncertain cases for a human to review.
A customer-review playground
Would this person buy again? What should we improve? How did they feel about the product? Ask all three and get answers your app can use.
“Should I send this Message?”
Got anger issues? Give it your draft Slack or email... whether you're about to rage quit, or saying something nice. It picks send now, soften first, sleep on it, or say it in person.
“Am I bombing?”
A pitch-coach experiment. The browser uses Face API to turn camera and voice cues into numbers; Jev uses those to suggest feedback. I simulated nervous and confident pitches to show the feedback changing.
Each one uses the same basic idea: give Jev the situation and the possible answers. It picks an answer, tells the app how sure it is, and the app uses that result.
You can start with one judgment you're tired of making over and over!
Here's my Jev Test™ for new and existing projects:
(1) Do I keep asking the same question?
(2) Can I name the possible answers?
(3) Would deciding faster or more often actually help?
If something in your work passes all three, that's a place to start.
Full 10-minute walkthrough attached, including when I'd still use Claude or ChatGPT.
Watch on YouTube:
What decision at work would you build one of these for?
(Shoutout to
@tenex_labs' finest --
@seejayhess and
@_raghavdixit_ -- for the knowledge you imparted for this one!!)
Show more
Im so excited to announce Tenex Agents!
Off-the-shelf, enterprise-grade AI agents focused on tackling recurring work across functions like Finance, Sales, Marketing, HR, and IT.
Everyone & their mother is building/launching agents. So how are ours different?
• < 2 weeks to deployment
• Built for responsible oversight (permissioning, approvals, audit trails)
• Optimized performance layer (model routing, evals, context management)
• Custom configured for your processes & systems of record
• Tied to agreed upon outcomes
Our agents JUST WORK because they’re a labor of love, 18 months in the making.
We started Tenex to help companies turn frontier AI research into real business results. Since then, we’ve built hundreds of custom solutions for some of the biggest companies in the world.
Along the way, clients keep asking: “What are you seeing on the ground, and what should we be doing next?”
We’ve studied the market and spoken with hundreds of executives. Three priorities kept coming up:
1) Faster time to value. See meaningful results quickly and cheaply.
2) A clearer starting point. Know which workflows to tackle and how to measure success.
3) A fit with the business. Work within existing systems, with the right access controls and human oversight.
That’s when Tenex Agents said “hold my beer.”
We’ve been quietly turned hair on fire customer processes problems into a suite of composable and battle-tested agents.
They give you a ready starting point, with Tenex handling the implementation:
• Up and running quickly. Most deployments targeted for under two weeks.
• Configured for your business. Adapted to your processes and connected to your systems of record.
• Built for oversight. Defined permissions, human approvals, and an audit trail.
• Measured against an agreed outcome. A clear definition of success for the first workflow.
We’re slowly rolling out Tenex Agents, so please fill out this interest form if you’d like to learn more:
Show more
This guy build a copy of Call of Duty with one prompt & went stupid viral (20 million views).
Now while building video games is cool, what's even cooler is that his process, called The Gauntlet Loop, can be applied to any type of professional work.
Here's how
@mattshumer_'s process works:
Problem this solves:
Agents stop at “good enough.” They do the ask once and declare done, especially in visual/creative work.
Even if you force more iterations, models judge their own work.
Like a student grading their own exam, they give themselves 100. That self-grading is the ceiling.
Step 1: Set a real, inspectable bar
- “most recent Call of Duty” level, not “a great FPS”). If yours isn’t better, you’re not done.
Step 2: Split the job across specialist sub-agents instead of one overloaded agent.
Step 3: Loop each piece until it looks like a real game/ real work, not “pretty good for AI.”
- Version 1 will be garbage.
Step 4: Blind critics compare your output to a reference and keep going until they pick yours (or you stop).
Why the critic must be blind
- In Claude Code / Codex-style harnesses, sub-agents often fork the main agent’s context, so a “critic” still remembers the builder’s rationale and rubber-stamps it.
- Fix: spawn critics with totally fresh context. They only see the artifacts (e.g. two images), pick which looks better, and don’t know which is yours.
- Until the critic picks yours, keep looping. That’s the whole loop.
How Claude of Duty actually ran
- Lead agent decomposes → each piece gets builder + critic → before/after per subsystem → fold back into the game → repeat waves.
- He did not hand-feed a big pile of Call of Duty reference images. He set the prompt and the agent went out and found comparison material itself.
References when the thing doesn’t exist yet
For something novel (e.g. a futuristic weapon), you can:
- Use a different game/object as a quality-level comp (critic judges “which looks better overall,” not identity match), or
- Generate target stills with an image model until you like them, then feed those as the bar.
Same loop outside games
- Writing (he called this the harder example): after each iteration, A/B paragraphs (or page/chapter) against strong recent comps. Avoid famous dead authors the model already “knows” as themselves (IP + identity). Use contemporary comparable writing.
- Websites: feed sites you think are world-class; don’t stop until blind critics consistently prefer yours.
- How prescriptive on “what good means”: optional. Experimental runs he lets loose. Real work he gets more prescriptive and steers mid-loop (“like this direction, but pare back”).
- Growth: Something Big newsletter is already using the loop for growth strategies and conversion copy; he said subscriber CAC results are far above industry standard.
Cost and when to stop
The loop can run forever.
- Demos / toys: don’t max it. Too expensive for the value.
- Real work that matters: willing to spend a couple hundred dollars / hit subscription limits. If the artifact is valuable, $200 is cheap relative to what that demo would have cost a few years ago.
- He stopped Claude of Duty while it was still improving, because of cost and “already wow,” not because the bar was fully beaten. He bets a couple more days would get much closer to real CoD level.
Full episode:
Show more
This guy build a copy of Call of Duty with one prompt & went stupid viral (20 million views).
Now while building video games is cool, what's even cooler is that his process, called The Gauntlet Loop, can be applied to any type of professional work.
Here's how
@mattshumer_'s process works:
Problem this solves:
Agents stop at “good enough.” They do the ask once and declare done, especially in visual/creative work.
Even if you force more iterations, models judge their own work.
Like a student grading their own exam, they give themselves 100. That self-grading is the ceiling.
Step 1: Set a real, inspectable bar
- “most recent Call of Duty” level, not “a great FPS”). If yours isn’t better, you’re not done.
Step 2: Split the job across specialist sub-agents instead of one overloaded agent.
Step 3: Loop each piece until it looks like a real game/ real work, not “pretty good for AI.”
- Version 1 will be garbage.
Step 4: Blind critics compare your output to a reference and keep going until they pick yours (or you stop).
Why the critic must be blind
- In Claude Code / Codex-style harnesses, sub-agents often fork the main agent’s context, so a “critic” still remembers the builder’s rationale and rubber-stamps it.
- Fix: spawn critics with totally fresh context. They only see the artifacts (e.g. two images), pick which looks better, and don’t know which is yours.
- Until the critic picks yours, keep looping. That’s the whole loop.
How Claude of Duty actually ran
- Lead agent decomposes → each piece gets builder + critic → before/after per subsystem → fold back into the game → repeat waves.
- He did not hand-feed a big pile of Call of Duty reference images. He set the prompt and the agent went out and found comparison material itself.
References when the thing doesn’t exist yet
For something novel (e.g. a futuristic weapon), you can:
- Use a different game/object as a quality-level comp (critic judges “which looks better overall,” not identity match), or
- Generate target stills with an image model until you like them, then feed those as the bar.
Same loop outside games
- Writing (he called this the harder example): after each iteration, A/B paragraphs (or page/chapter) against strong recent comps. Avoid famous dead authors the model already “knows” as themselves (IP + identity). Use contemporary comparable writing.
- Websites: feed sites you think are world-class; don’t stop until blind critics consistently prefer yours.
- How prescriptive on “what good means”: optional. Experimental runs he lets loose. Real work he gets more prescriptive and steers mid-loop (“like this direction, but pare back”).
- Growth: Something Big newsletter is already using the loop for growth strategies and conversion copy; he said subscriber CAC results are far above industry standard.
Cost and when to stop
The loop can run forever.
- Demos / toys: don’t max it. Too expensive for the value.
- Real work that matters: willing to spend a couple hundred dollars / hit subscription limits. If the artifact is valuable, $200 is cheap relative to what that demo would have cost a few years ago.
- He stopped Claude of Duty while it was still improving, because of cost and “already wow,” not because the bar was fully beaten. He bets a couple more days would get much closer to real CoD level.
Full episode:
Show more
This week I delivered a 60-minute workshop to CFOs, CHROs, and CLOs teaching them how to go from desired outcome (i.e. offer acceptance >60%) to designed AI pilot in 6 steps.
Here are the 6 steps:
*If you want the deck & worksheet I used for the workshop, check the link below.
1) Identify the outcome
Transformation isn't about AI. It's about driving outcomes and using tools (like AI) to solve problems that stand in the way. Before deciding what work/product to transform, you need to decide the outcome.
To do that, fill in the blanks:
Improve [measure] from [baseline] to [target] by [date].
Examples: overdue invoices 18% to 12%. Month-end close 10 days to 6. Cash-forecast updates 2 days to 2 hours. Offer acceptance 40% to 60%.
2) Find the workflow
Every business is plagued with inefficiency, even the most AI-native ones. Make a list of all of the workflows within your function/company that sit close enough to your outcome to impact it.
Shorten the list to the workflows that pass the PAIN IN THE ASS TEST, and then circle one workflow that you believe, if made maximally efficient, would have the greatest impact on driving your outcome.
One exercise for brainstorming workflows is forcing yourself to answer two questions:
If you imagine work 12 months from now...
1) What is one specific way work happens differently?
2) What measurable result does that change create?
3) Design the new work
First, create a process diagram of the workflow you chose as it stands today. It should look like 5-7 steps connected by lines.
Second, draw a new process diagram of the workflow in its future, most efficient form. Next to each step put a label for what drives the work. Three labels to choose from: AI-led, AI-assisted, Human-only.
Example: Hiring outreach process
Current: Define hiring criteria --> Search for prospects
--> Research fit and contact history --> Write outreach messages --> Approve, send, and log in ATS --> Copy activity into a separate tracker --> Handle replies and hand off
New: Define criteria and permitted sources (Human) --> Find prospects using approved sources (AI-led) --> Review AI research; choose contacts (AI-assisted) --> Draft outreach from approved context (AI-led) --> Approve (Human) --> send, and log in ATS (AI-led) --> Handle replies and hand off (AI-assisted)
4) Assess readiness
Once you've reimagined the work, you still can't press go. There's a list of 8 items that must be checked (GREEN/YELLOW/RED) before proceeding. You don't necessarily need all items to be GREEN ahead of a pilot, but you definitely do ahead of production.
• Value at stake - You can name the specific benefit to your business (you've built out the ROI case) and a plausible path from this workflow to it
• Process clarity and measurability - You know the (new/old) steps, start and finish, exceptions, and how to establish a baseline
• Data and context readiness - The required context exists, is fit for the test, and can be accessed within agreed boundaries
• Deployment capability - Someone (internally/externally) can configure, connect, test, support, and roll back the pilot
• User adoption - Intended users help design it, understand their role, and will try it in real work
• Change capacity - The team can spare time for training, review, feedback, and workflow changes
• Ownership and decision rights - A business owner owns the result. People know who approves changes, resolves issues, and stops the test
• Legal risk and controls - Required approvals, access limits, human review, logging, and stop rules are in place for the pilot.
5) Design & run pilot
Every pilot should be documented & pitched internally with the following considerations included:
• Scope: Who's included, what's the work they're doing, for how long?
- Example: Two recruiters. One engineering role family. Four weeks, after permissions and controls are cleared.
• Success: What's the business result you're looking to drive and what are early signals you need to see?
- Example: Compare with the manual baseline: ≥25% less net prospecting time, no drop in prospect relevance, and saved time used for candidate care. Track offer acceptance longer.
• Economics: Which benefit and cost assumptions need testing?
- Example: Measure setup and running costs, including recruiter review time. Use observed time savings and prospect quality to update the business-case assumptions. Track actual agency fees avoided and additional delivery contribution over a longer period.
• Guardrails: What must not happen?
- Example: No unauthorized outreach or material opt-out breach.
• Decision: What must be true to expand, revise, or stop the pilot?
- Example: At week 4: do the early results justify continuing a bounded test? Continue if early gates pass; revise if fixes are viable; stop for a material breach or no credible path to value. Full ROI is not yet proven.
6) Scale scope & autonomy
Based on performance of pilot & state of readiness criteria (from step 4), progressive productionizing of the workflow, rollout to the org, and autonomy for AI occurs.
Link to the workshop & worksheet:
Show more
I was the biggest skeptic of AI video editing. Until recently.
So I asked one of the most talented creators I know (
@cinkotweets) to break down exactly how AI can support the entire video process.
The guy cooked. Here are the high-level notes:
The tl;dr
Right now AI mostly solves two genres. The video essay and the tutorial. Both have a straightforward outline, a script you stick to, and a one-shot record that AI can chop against. Other formats still lean heavily on craft.
Easy Mode: research and ideas without losing your voice
The gist: Use AI to find opportunities and pull source material from tools and team discourse. You still decide what is worth saying. Don't outsource final sign-off.
Pro tips:
- Marketplace is huge. Type YouTube and you will see VidIQ, TubeBuddy, and a pile of others. If your company shows a “request” button, go bug the admin. Do not wait on the queue.
- Hook Ahrefs into Claude (and Notion) and ask something like: use the Ahrefs connector to find AI content opportunities this week, and tell me if videos already sit in the top five.
- Tribal Knowledge is the other half. It is a skill that scrapes every connector he has (Slack, Notion, meetings) and surfaces real conversations around a topic so the idea comes from work talk.
- Turn that into a personalized daily brief for the kind of content you make. Anthony built a content planning app with Codex. Every day it drops a brief, proposes video topics, and lets him queue the ones he likes. Notion is the shared backend so the org can see the same database without living in his app. Even on Opus, a daily brief runs about 8 cents.
Hard Mode: script and first edit with AI assistants
The gist: AI helps you outline and cut. You keep the words, the takes, and the judgment. Editing is not just the blade button. It is taste. If you are an editor, you have an advantage. If you have vision but not technical chops, you can still leverage these tools.
Pro tips:
- Create the outline as a human. Anthony still believes you should not let AI one-shot structure. His format is three columns: dialogue, visuals (what is on screen while you talk), and sound effects.
- Then use the Scriptwriter skill. Point it at a topic (he demoed “what is an MCP”) and it interviews you through the whole creative process so you speak like yourself. Short-form or long. You can ramble, go out of order, contradict yourself. It sorts after.
- Beat by beat it spits your words back as bullets. You react raw in the moment. Those reactions become the real script. It also prompts visuals and SFX while you talk (day-to-night, crickets, rooster) so creativity stays in the loop. Hooks are hard. You can ask it to interview you into a better hook without handing it the whole voice.
*Film it yourself*
- Split editing into phases. Baseline “radio cut” (dialogue paced). Then VFX / archival / B-roll. Then audio (dialogue mix, SFX, music).
- Codex + Final Cut: pointed it at the footage and the script, asked for a baseline cut. First pass had half-second gaps of silence. One follow-up (“every clip is consistently half a second off”) and it produced a fluid cut with no awkward pauses, dropped straight into Final Cut. Cost: two prompts burned a big chunk of a 5-hour context window (100% down to 14%), but against weekly Codex usage (decks, research, more content) he still had about 53% left.
- Codex in Final Cut won the radio cut. Descript was faster on the clock and worse on taste cleanup.
God Mode: polish stack plus a measurement loop
Definition: hand off the laborious extras (B-roll, motion, music, reporting). Keep the taste calls. Close the loop so you know if the work is working.
Pro tips:
- B-roll / archival / VFX: Anthony spun an “archival finder” agent. He gave it a vibe (internety, clicky, fun) and a reference (simple, warm, approachable using shows, TV, movies). It came back with taste.
- It also downloaded clips, filed them, dropped them on the timeline, and added a paper texture background he had asked it to design. All via the Final Cut connector plus computer use.
- Motion graphics: Remotion (free, open source) plus an agent. Dump everything in your head. Ask for a plan and structure first. He asked for an Apple liquid glass ultra-clean feel. Plan, approve, tweaks, then it popped into the Final Cut timeline in about 9 minutes 10 seconds. That used to be a brief to a VFX designer (timing, seconds, design, wait for turnaround).
- Music and SFX: Epidemic Sound via MCP (Artlist and free-sound options exist too). Same pattern: send the URL if you need to install, give it a creator reference (he used Zoe), iterate because you are picky about music. It dropped tracks into the timeline and used Epidemic’s crop/trim so the music was cut for the edit. Mix sat under vocals instead of overpowering them. Not just music. Mouse-click SFX timed to B-roll.
Full episode:
Show more
If you want to understand how Jev works, highly recommended read 👇
There's a lot of confusion about what Jev actually is, and when it's worth using over what you already have.
Here is my attempt at reverse engineering its philosophy, its architecture, and where it falls down.
Show more
This week I delivered a 60-minute workshop to CFOs, CHROs, and CLOs teaching them how to go from desired outcome (i.e. offer acceptance >60%) to designed AI pilot in 6 steps.
Here are the 6 steps:
*If you want the deck & worksheet I used for the workshop, check the link below.
1) Identify the outcome
Transformation isn't about AI. It's about driving outcomes and using tools (like AI) to solve problems that stand in the way. Before deciding what work/product to transform, you need to decide the outcome.
To do that, fill in the blanks:
Improve [measure] from [baseline] to [target] by [date].
Examples: overdue invoices 18% to 12%. Month-end close 10 days to 6. Cash-forecast updates 2 days to 2 hours. Offer acceptance 40% to 60%.
2) Find the workflow
Every business is plagued with inefficiency, even the most AI-native ones. Make a list of all of the workflows within your function/company that sit close enough to your outcome to impact it.
Shorten the list to the workflows that pass the PAIN IN THE ASS TEST, and then circle one workflow that you believe, if made maximally efficient, would have the greatest impact on driving your outcome.
One exercise for brainstorming workflows is forcing yourself to answer two questions:
If you imagine work 12 months from now...
1) What is one specific way work happens differently?
2) What measurable result does that change create?
3) Design the new work
First, create a process diagram of the workflow you chose as it stands today. It should look like 5-7 steps connected by lines.
Second, draw a new process diagram of the workflow in its future, most efficient form. Next to each step put a label for what drives the work. Three labels to choose from: AI-led, AI-assisted, Human-only.
Example: Hiring outreach process
Current: Define hiring criteria --> Search for prospects
--> Research fit and contact history --> Write outreach messages --> Approve, send, and log in ATS --> Copy activity into a separate tracker --> Handle replies and hand off
New: Define criteria and permitted sources (Human) --> Find prospects using approved sources (AI-led) --> Review AI research; choose contacts (AI-assisted) --> Draft outreach from approved context (AI-led) --> Approve (Human) --> send, and log in ATS (AI-led) --> Handle replies and hand off (AI-assisted)
4) Assess readiness
Once you've reimagined the work, you still can't press go. There's a list of 8 items that must be checked (GREEN/YELLOW/RED) before proceeding. You don't necessarily need all items to be GREEN ahead of a pilot, but you definitely do ahead of production.
• Value at stake - You can name the specific benefit to your business (you've built out the ROI case) and a plausible path from this workflow to it
• Process clarity and measurability - You know the (new/old) steps, start and finish, exceptions, and how to establish a baseline
• Data and context readiness - The required context exists, is fit for the test, and can be accessed within agreed boundaries
• Deployment capability - Someone (internally/externally) can configure, connect, test, support, and roll back the pilot
• User adoption - Intended users help design it, understand their role, and will try it in real work
• Change capacity - The team can spare time for training, review, feedback, and workflow changes
• Ownership and decision rights - A business owner owns the result. People know who approves changes, resolves issues, and stops the test
• Legal risk and controls - Required approvals, access limits, human review, logging, and stop rules are in place for the pilot.
5) Design & run pilot
Every pilot should be documented & pitched internally with the following considerations included:
• Scope: Who's included, what's the work they're doing, for how long?
- Example: Two recruiters. One engineering role family. Four weeks, after permissions and controls are cleared.
• Success: What's the business result you're looking to drive and what are early signals you need to see?
- Example: Compare with the manual baseline: ≥25% less net prospecting time, no drop in prospect relevance, and saved time used for candidate care. Track offer acceptance longer.
• Economics: Which benefit and cost assumptions need testing?
- Example: Measure setup and running costs, including recruiter review time. Use observed time savings and prospect quality to update the business-case assumptions. Track actual agency fees avoided and additional delivery contribution over a longer period.
• Guardrails: What must not happen?
- Example: No unauthorized outreach or material opt-out breach.
• Decision: What must be true to expand, revise, or stop the pilot?
- Example: At week 4: do the early results justify continuing a bounded test? Continue if early gates pass; revise if fixes are viable; stop for a material breach or no credible path to value. Full ROI is not yet proven.
6) Scale scope & autonomy
Based on performance of pilot & state of readiness criteria (from step 4), progressive productionizing of the workflow, rollout to the org, and autonomy for AI occurs.
Link to the workshop & worksheet:
Show more
I was the biggest skeptic of AI video editing. Until recently.
So I asked one of the most talented creators I know (
@cinkotweets) to break down exactly how AI can support the entire video process.
The guy cooked. Here are the high-level notes:
The tl;dr
Right now AI mostly solves two genres. The video essay and the tutorial. Both have a straightforward outline, a script you stick to, and a one-shot record that AI can chop against. Other formats still lean heavily on craft.
Easy Mode: research and ideas without losing your voice
The gist: Use AI to find opportunities and pull source material from tools and team discourse. You still decide what is worth saying. Don't outsource final sign-off.
Pro tips:
- Marketplace is huge. Type YouTube and you will see VidIQ, TubeBuddy, and a pile of others. If your company shows a “request” button, go bug the admin. Do not wait on the queue.
- Hook Ahrefs into Claude (and Notion) and ask something like: use the Ahrefs connector to find AI content opportunities this week, and tell me if videos already sit in the top five.
- Tribal Knowledge is the other half. It is a skill that scrapes every connector he has (Slack, Notion, meetings) and surfaces real conversations around a topic so the idea comes from work talk.
- Turn that into a personalized daily brief for the kind of content you make. Anthony built a content planning app with Codex. Every day it drops a brief, proposes video topics, and lets him queue the ones he likes. Notion is the shared backend so the org can see the same database without living in his app. Even on Opus, a daily brief runs about 8 cents.
Hard Mode: script and first edit with AI assistants
The gist: AI helps you outline and cut. You keep the words, the takes, and the judgment. Editing is not just the blade button. It is taste. If you are an editor, you have an advantage. If you have vision but not technical chops, you can still leverage these tools.
Pro tips:
- Create the outline as a human. Anthony still believes you should not let AI one-shot structure. His format is three columns: dialogue, visuals (what is on screen while you talk), and sound effects.
- Then use the Scriptwriter skill. Point it at a topic (he demoed “what is an MCP”) and it interviews you through the whole creative process so you speak like yourself. Short-form or long. You can ramble, go out of order, contradict yourself. It sorts after.
- Beat by beat it spits your words back as bullets. You react raw in the moment. Those reactions become the real script. It also prompts visuals and SFX while you talk (day-to-night, crickets, rooster) so creativity stays in the loop. Hooks are hard. You can ask it to interview you into a better hook without handing it the whole voice.
*Film it yourself*
- Split editing into phases. Baseline “radio cut” (dialogue paced). Then VFX / archival / B-roll. Then audio (dialogue mix, SFX, music).
- Codex + Final Cut: pointed it at the footage and the script, asked for a baseline cut. First pass had half-second gaps of silence. One follow-up (“every clip is consistently half a second off”) and it produced a fluid cut with no awkward pauses, dropped straight into Final Cut. Cost: two prompts burned a big chunk of a 5-hour context window (100% down to 14%), but against weekly Codex usage (decks, research, more content) he still had about 53% left.
- Codex in Final Cut won the radio cut. Descript was faster on the clock and worse on taste cleanup.
God Mode: polish stack plus a measurement loop
Definition: hand off the laborious extras (B-roll, motion, music, reporting). Keep the taste calls. Close the loop so you know if the work is working.
Pro tips:
- B-roll / archival / VFX: Anthony spun an “archival finder” agent. He gave it a vibe (internety, clicky, fun) and a reference (simple, warm, approachable using shows, TV, movies). It came back with taste.
- It also downloaded clips, filed them, dropped them on the timeline, and added a paper texture background he had asked it to design. All via the Final Cut connector plus computer use.
- Motion graphics: Remotion (free, open source) plus an agent. Dump everything in your head. Ask for a plan and structure first. He asked for an Apple liquid glass ultra-clean feel. Plan, approve, tweaks, then it popped into the Final Cut timeline in about 9 minutes 10 seconds. That used to be a brief to a VFX designer (timing, seconds, design, wait for turnaround).
- Music and SFX: Epidemic Sound via MCP (Artlist and free-sound options exist too). Same pattern: send the URL if you need to install, give it a creator reference (he used Zoe), iterate because you are picky about music. It dropped tracks into the timeline and used Epidemic’s crop/trim so the music was cut for the edit. Mix sat under vocals instead of overpowering them. Not just music. Mouse-click SFX timed to B-roll.
Full episode:
Show more
Hide your kids. Hide your wife. AI slop is taking over.
Thankfully, my guy
@cinkotweets gives you an actionable playbook for fighting AI slop in your company. Give it a watch.
My brother-in-law landed a job interview for two reasons:
1. He used AI to optimize his resume keywords for the ATS scanner
2. He dropped a bullet under Skills that read: "bicycle repair, eccentric ideating, frolicking amongst foliage and critters, and patriotism"
Guess which one they asked him about.
AI has a personality, but it’s just the average of everything it knows, and everybody's has the same one. Your quirks, your references, the weird asides you put in in parentheses — your ✨authenticity✨ — that’s the last remaining “moat” to stand out in anything you produce.
That's one of 5 things I learned making this video essay.
I spent a day yapping with my team at Tenex (people who are hired FOR their AI skills how they avoid making slop with it) to try and figure out: Why do people drift from using AI for grunt work to completely outsourcing their thinking to it?
The point here isn't as simple as "AI Bad" -- it's about how to put your thinking cap back on!
Tap in and tell me if there's a pointer I missed.
Show more
Profound just raised a $180M Series D at a $1.8B valuation, co-led by Sequoia and Kleiner Perkins.
The company has been a textbook example of what it takes to build a huge business with sustained advantages in a post-AI world:
Proprietary Data: 100M+ AI answers analyzed every month, across 18 countries, which created unfair differentiation within the AEO space.
Trusted distribution: by shipping laser-fast, coining the phrase Marketing Engineer, and expanding from a hyper-specific marketing product (AEO) into a full-stack AI platform for marketers, the exact right audience (marketers) deeply trust the brand.
Domain-specific harness: created the AI Marketer (Aim), the first AI harness designed for marketing teams. It’s grounded in your AEO data, the tools your team already uses, and your business’ complete context.
The result:
@profound has landed 1,000+ enterprise brands and 18% of the Fortune 500 since launch in 2024.
Congrats to
@thejamescad and team on the new round and valuation.
Find out how you can build the marketing team of the future here:
#
ProfoundPartner# #
Sponsored#
Show more
I cant help but think sometimes that AI ruined coding
Remember the feeling of locking in for hours and hours. The context of the entire codebase all loaded in to your brain. Typing code letter by letter. Bug after bug after bug until it finally worked?
Exactly zero new engineers will ever feel that level of flow
I wouldn't trade this for the world, but I do miss those days
Show more
Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient.
Let me explain with a scenario:
Imagine a company already has an agent workflow running where an llm reviews every item before it moves on: a support ticket gets triaged, an invoice gets approved or held, a claim gets flagged.
Every one of those goes through a frontier model today, a few seconds and a few cents each, on the way to a decision that in most cases is obvious. Behind that flow sits years of humans (or agents) making the exact same call, with the outcome attached.
Now imagine you first run each item through a custom PCD or similar model that costs a fraction of the llm and returns a classification of what to do at that step, with a mathematically accurate probability attached.
When it's confident, the item skips the llm entirely.
When it isn't, the llm handles it as normal.
The model has seen years of your team making this exact decision, usually a constrained set of decisions, so it should be right most of the time. Say it comes back confident on 6 out of 10 items. That's more than half your llm spend potentially gone from that step, likely with comparable accuracy.
This pre processing idea works in a bunch of other use cases too, such as:
- model/request routing: cheap model, frontier model, or a human
- picking which skill or subagent to load for a turn instead of stuffing the whole catalog into context
- reranking retrieved context so only the relevant chunks reach the window
- guardrails on every agent turn: contradictions, policy issues, prompt injection
- extracting typed fields from unstructured data emails, PDFs and transcripts before anything expensive touches them
Every one of those is a decision an llm makes today, that could potentially be done by another, cheaper model class. Very excited to see Jev/PCD-based pre processing use cases get deployed to agents at scale.
Show more