This prediction-market arbitrage playbook is f*cking quant-grade
A quant just released the full L2 order book dataset and mathematical framework that hedge funds use to extract riskless alpha from prediction platform spreads → I compiled it into a walkthrough:
Their release, a 55GB tick-by-tick L2 dataset with 850 million state updates across Polymarket and Kalshi, solves the fundamental data gap in prediction market arbitrage: retail traders analyze historical transaction prints, quant desks analyze the bid-ask queues that produce them
Why retail arbitrage fails:
→ Reads charts of execution prices → misses the liquidity distribution at each level
→ Enters the moment a gap appears → gets picked off by desks solving optimal stopping problems
→ Trades directional views on the underlying event → the desks are event-agnostic
Two methods deployed on the data:
[1] Cointegration + Ornstein-Uhlenbeck Mean Reversion: models the spread between platforms as a stationary I(0) process, validates cointegration via Augmented Dickey-Fuller or Johansen tests, then fits dS_t = θ(μ - S_t)dt + σdW_t to extract the exact mean-reversion rate θ via MLE on 850M rows
[2] Optimal Entry/Exit Thresholds: instead of "enter when the gap appears," solves the stopping problem for x_open and x_close that maximize expected return per unit time, net of transaction costs c across both venues
[3] Order Book Imbalance + Micro-Price Prediction: at millisecond horizons, computes I_t = (V_b - V_a)/(V_b + V_a) and derives the micro-price P_micro = P_mid + I_t(Δspread/2), the true instantaneous value before a transaction prints
[4] Cross-Venue Predictive Signaling: rolling Markov chain transition matrix measures P(imbalance shift on Platform A → ask book clears on Platform B within δ = 200ms), because Polymarket runs hybrid/on-chain and Kalshi runs centralized clearing at different velocities
The Institutional Edge Scorecard:
Dataset depth: L2 snapshots up to 20 levels deep, sampled at 100ms intervals
Storage: Apache Parquet, cross-platform synced timestamps, 39GB compressed via zstd
Signal window: δ = 200ms cross-venue lag between imbalance shift and price adjustment
Tradeability filter: asset qualifies only when mean-reversion time τ = 1/θ is faster than platform execution latency
Hypothesis Verification & Practical Validation:
The quant hypothesized that spreads between competing prediction platforms are dominated by mean-reverting microstructure rather than divergent views on the underlying event. Practical validation confirmed the finding: platforms hosting identical real-world outcomes must converge to identical terminal values ($1.00 or $0.00), so every intermediate divergence is either a mean-reverting spread (OU) or a latency-driven imbalance signal (OBI), both of which are extractable with the released dataset before retail order routers register the shift
Alpha in prediction markets is not about reading polls or macro reports. The unfair advantage is operational: OU calibration over nominal price spreads, micro-price dominance over mid-price, latency capitalization over UI refresh rates
Read the complete breakdown in the article below ↓