A retired plumber in Nebraska beat the CIA at predicting foreign elections in 2013. A 22 year old college dropout beat every pollster in America in 2024. A 9 billion dollar prediction market called Polymarket is now telling you whether the AI bubble bursts this year, whether the US enters a recession, and whether Anthropic overtakes OpenAI before December.
They are all using the same 4-step technique a Berkeley professor proved actually works in a 20 year experiment that should have ended the careers of half the experts on television.
His name is Philip Tetlock.
In the late 1980s he started collecting predictions from 284 experts. Political scientists. Economists. CIA-adjacent analysts. People paid their entire careers to forecast geopolitics. Over 20 years he gathered 28,000 predictions. Then he scored them.
The result destroyed an entire profession. The average expert was barely better than chance. The famous ones, the ones with the most media appearances and the loudest voices, were the worst of the group. The more confident the voice, the worse the score.
Buried in his data was something nobody else had emphasized. Fewer than 2% of forecasters were dramatically better than the rest. Year after year. Across domains they had no training in.
In 2011 the US intelligence community gave him his chance to prove it at scale. Still bruised from missing Iraq, IARPA ran a 4 year tournament on 500 geopolitical questions. Will North Korea launch a missile. Will Russia invade. The intelligence analysts had classified intercepts. Tetlock's team had retired plumbers and ballroom dancers.
Tetlock's team won by 35 to 72 percent against the other academic teams. His top forecasters scored 30 percent better than the CIA reading classified data.
A retired pipe installer in Nebraska was outpredicting the intelligence community using only the newspaper.
The technique sits in his book in plain language and almost nobody applies it.
It starts with a method invented by Enrico Fermi during the Manhattan Project.
Fermi handed students problems that looked impossible. How many piano tuners are there in Chicago. He did not want a guess. He wanted them to break the question into smaller questions they could actually estimate. Each sub-estimate was rough. Multiplied together, they landed remarkably close to the truth.
Superforecasters Fermi-ize everything. They never try to predict a complex event directly. They shatter it into smaller questions where base rates are knowable, then reassemble the pieces.
The second move is the outside view. Most people, asked whether a startup will survive or a war will end by a date, dive into the specific details. Story details feel useful. They are not. Superforecasters first ask how often events of this general type happen across history. The story comes last, not first.
The third move is what most people refuse to do. Superforecasters update their predictions constantly, in tiny increments. Not dramatic reversals. Small honest nudges. They move like a Bayesian. Most people move like a teenager defending a position.
The fourth move is the one that hits closest. Superforecasters express predictions as actual numbers. Not "likely" or "probably." 62 percent. 18 percent. Vague language is unfalsifiable. A number forces accountability, and accountability is the engine of accuracy.
Polymarket is the same idea scaled to a hundred thousand strangers with real money on the line.
In the final weeks of the 2024 US presidential election, every major pollster called the race a coin flip. FiveThirtyEight had Harris at 50 to 49. Nate Silver had her at 48.6 to 47.6. Polymarket had Trump at 58 to 42 the morning of the election. By midnight, while networks still refused to call swing states, Polymarket was at 97 percent.
In late 2025 the New York Stock Exchange invested 2 billion dollars in the platform. Polymarket is now valued at 9 billion. The largest stock exchange in the world is integrating prediction prices directly into the data feed traders use to make decisions.
The plumber did not have classified intercepts. The college dropout did not have a polling model. Polymarket does not have an algorithm nobody else can see.
They all have the same thing. A method that forces you to write a number on paper, attach a date, and let the world watch you be wrong.
In 2026 the gap between people who price the future and people who narrate it is going to be the most expensive gap in the world to be on the wrong side of.
Show more
NEW RESEARCH: Robot hands can now play the piano!
It involves
@amberxie_,
@HaozhiQ, and
@DorsaSadigh.
Called HandelBot, it is a bimanual system that plays real songs on a real piano using two Tesollo DG-5F dexterous hands, one on a Franka Panda arm and one on an FR3 arm, using only three fingers per hand (index, middle, ring; 9 residual action dimensions per hand).
A policy trained entirely in simulation with RL is adapted to the real piano in two stages:
1. a structured refinement step that repeatedly executes the trajectory, compares the intended key against the key actually pressed, and nudges each finger's lateral joint to correct it.
2. then residual RL that learns fine corrective actions on top.
The arm wrist poses are scripted from sheet music. The only real-world feedback, and the reward, is the keyboard's own MIDI output.
There are no cameras and no tactile sensing.
It replaces collecting large real demonstration sets for millimeter-precise contact tasks.
The DG-5F fingertip (about 2.2 cm) is wider than a piano key (about 2.16 cm), so a single fingertip physically straddles two keys.
Which is why sub-millimeter lateral alignment is the whole game and why the thumb and pinky are dropped, only three fingers per hand fit, and songs must be re-fingered and the hands separated by octaves to avoid arm collision.
The real-world adaptation uses only the keyboard's MIDI output: as the per-finger error signal for the lateral refinement and as the sole RL reward (key-press reward, no fingering or energy terms).
No camera, no touch. The task's own output is a dense, objective reward for free, the cleanest instance of the instrument being the reward function.
For this task, 30 to 60 minutes of real data beats 40 million sim steps deployed directly.
RL trained from scratch on hardware (mean F1 55.8) beats the open-loop sim policy (41.8) on 4 of 5 songs, therefore sim pretraining's value is as an exploration-reducing initialization.
The scope is narrow and honest: five short pieces (16 to 33 s), three fingers per hand, no chords beyond three notes, no thumb, and the hardest piece (Fur Elise) tops out at 66 to 71 F1 with large left-hand jumps as named failures.
Refinement can only fix mis-located presses, not missed ones, and its heuristics are piano-specific (the authors suggest a VLM to generalize the coarse step).
Show more
Requiem for building in public
Music video workflow + Prompts:
One song, a simple story, singers across 4 locations, a full edit, captions burned in. Models used: Suno for the track, Seedance 2.5 for the story footage, MiniMax H3 for the lip synced singing, faster-whisper for word timing, Claude to edit, ffmpeg to burn captions.
THE SONG (Suno)
Short lyrics, fast beat, vocals on second 0. Long slow AI songs fall apart because the model has nothing to hide behind. Fire small batches, 2 clips at a time, and listen before firing more. Write every new attempt from zero.
Stacking "less this, no that" onto the last try feeds the model your confusion and hands it back. One adjective moves everything. I put "soft" in a prompt once and the whole vocal switched to a woman.
Remix prompt (paste into Suno style box):
aggressive male rap, hard boom bap drums with fast energy, dark piano loop, deep male voice on every line including the hook, punchy mix, vocals start immediately at 0:00, no instrumental intro
Lyrics:
[Hook]
It's just this thing I feel
When I wanna steal
It's just this thing I feel
When I wanna steal
[Verse 1]
Yo, I see you on X, all over my feed
You're building in public, I'm watching you build
Your MRR chart looks like a hockey stick
I screenshot it sometimes, that's normal right
[Hook]
It's just this thing I feel
When I wanna steal
[Verse 2]
I learned a lot from you
I think I deserve it too
So I copied everything from you
Same landing page, same pricing, same font
And now you blocked me
What happened bro
I was your biggest fan
[Hook]
It's just this thing I feel
When I wanna steal
Tip: short lyrics, fast beat, vocals at 0:00, fresh prompt every round.
THE STORY
Think old MTV. The video is the movie this song is the soundtrack of. Keep the plot dead simple, something you can follow with the sound off.
Mine: a broke founder copies a guy, dreams he is rich, wakes up, sees he got blocked, spits his cereal at the screen. That is all of it.
Tip: if you cannot explain the story with zero words, cut it down.
THE STORY FOOTAGE (Seedance 2.5)
Seedance made the apartment story as one 30 second clip from reference images.
Two things kill Seedance:
Too many object interactions in one shot, and timestamps like "0 to 4 seconds" which it reads as a time lapse and speeds through.
plain shots labelled "Shot 1, Shot 2" at natural speed will do the trick
For the hard beat at the end, a guy waking up, eating cereal, seeing a screen, then spitting milk on the lens, text alone will not hold it. I
built a 3 panel storyboard image and fed it as a reference.
In the prompt you tag that image at the exact moment it happens, tell it the board reads left to right, describe it once, and move on.
The rule that saved it: chronological order, tag each image where it belongs in time, say everything a single time, never repeat a thing.
Repeat one detail twice and the model fixates on it and breaks the shot.
Tip: for anything complex, hand it a storyboard picture and describe it once, in order.
THE SINGING (MiniMax H3)
H3 is the model that lip syncs to your actual track and keeps it. Seedance cannot, it regenerates its own audio. In H3 you attach your audio slice, set it to copy, and the mouth follows your real song.
H3 caps around 15 seconds a clip and the song is 43. So I cut the song into 4 windows of about 11 seconds and generated a shot for each window.
Then I did it across 4 locations, subway, warehouse, empty office, street. That is a 4 by 4 grid, 16 clips. I added 8 more where the whole crew sings and dances. Around 24 short singing clips to cover a 43 second song. You are building a bank of clips to cut from.
Small H3 rules that matter: the audio slice must be a touch shorter than the clip, name every speaker, and compress the slice so there are no silent gaps for the model to fill with invented sound.
Tip: chop the song into sub 15 second windows, shoot each shot per window, build a clip bank.
THE EDIT (Claude)
This is where most people lose hours. I made editing fast by doing the prep once. Every clip gets normalized to the same size and frame rate up front.
After that each edit is a single ffmpeg pass with no re-encoding loops.
The base layer is the song.
Every clip's own audio is thrown out. To keep mouths in sync I gave the editor the math: each clip knows which second of the song its first frame belongs to, so to place it at song second S you trim it to start at S minus that offset.
I also handed over word level timing from a whisper pass so cuts could land on real lyric moments.
The rules I gave: nothing stays on screen too long, pace every cut to the lyric and the beat and what is on screen, never put two shots from the same location back to back, keep it heavy on story B-roll, never reuse a frame. If the cut feels like a metronome you failed. If it feels random you failed.
Then the actual move. I did not ask for one perfect edit. I gave 7 agents the same rules and the same clip bank and told each to cut the whole thing its own way with a different emphasis. 6 came out flat. 1 landed around 90 percent. I finished that one by hand in CapCut.
Tip: give strict rules plus the timing data, generate many full edits, keep the best and finish it yourself.
CAPTIONS (ffmpeg)
Burned straight from a styled subtitle file with ffmpeg. Seconds, not the long render a motion tool costs.
The words come from the real lyrics, the timing comes from a whisper pass on the audio, and it highlights the word being sung. Big, thick, one pop color on the active word.
Tip: real lyrics for the words, whisper for the timing, burn with ffmpeg.
The AI did not make this video. I directed it, generated in volume, and kept the best takes. That is the whole game right now.
Show more