⚡ Make aggregations over billions of records up to 100x faster by adding one line at the top of your query, with statistically rigorous confidence intervals built in. That's approximate queries in Elasticsearch ES|QL.
Title: Approximate queries in Elasticsearch ES|QL: 100x faster on billions of records, with built-in confidence intervals
URL:
📝 Overview
Elasticsearch 9.4 adds approximate query execution to ES|QL. Just prepend SET approximation = true; to an existing query and automatic sampling and extrapolation kick in, with no query rewrites needed.
❓ Challenges Solved
Exact aggregations over billions of documents are costly because compute scales linearly with row count. That hampered interactive exploration and real-time dashboards on large indices.
💡 Methodology & Proposed Approach
・Sampling happens at the Lucene layer, reading only the sampled documents, so I/O and compute savings are proportional to the sampling rate
・The query runs on the sample, then results are automatically scaled up to represent the full dataset
・Confidence intervals are computed rigorously via a bootstrap over sub-partitions of the sample
・Each result carries a certified flag indicating whether formal statistical guarantees hold
🎯 Use Cases
An agent can sweep billions of documents in sub-second time, narrow down candidates, and zoom into exact queries only where needed. It also speeds up dashboards and pattern detection over massive logs.
📊 Results
・On ClickBench, an average of 23x with confidence intervals, peaks around 100x per query, and up to ~300x without interval computation
・Since sampling cost stays constant, speedup grows with dataset size
・Supported aggregations include COUNT, SUM, AVG, MEDIAN, PERCENTILE, and STD_DEV, with rows and confidence_level tuning accuracy versus speed
#
Elasticsearch# #
DataAnalytics#