But latency alone doesn’t cut it. A sound effect has to follow the prompt, preserve crisp transients and believable acoustics, and avoid unrelated content.
Across the systems evaluated, Pika SFX scored highest for content usefulness and production quality—the measures closest to whether an output is ready to use in a creative workflow.
Show more
Our serving infrastructure lowers latency and improves throughput across both online and batch embedding workloads.
Combining Ivy, Tulip, and ROSE results in faster search at a reduced cost compared to off-the-shelf solutions.
Show more
Hyperliquid Opens Low-Latency Data Nodes to Infrastructure Providers at Under $1,000 a Month
Hyperliquid Foundation has opened its low-latency on-chain data nodes to qualified infrastructure providers, allowing them to offer access at standardized pricing, currently indicated at under $1,000 per month. Previously, direct access required staking 10,000 HYPE and qualifying for Tier 1 maker rebates, defined as more than 0.5% of 14-day weighted maker volume.
Show more
Most agent latency comes down to one of these:
- Model latency
- Waiting on external APIs
- Retries after failed tool calls
- Waiting on a human to approve something
Where does your agent actually lose most of its wall-clock time?
Show more
Most providers show you latency on a good day. BUT we rebuilt our node infra for the bad ones. New node-data transfer: 15x faster. A 13-node Base fleet, rebuilt in ~35 min with zero failures.
ngl, we're fast even when hardware fails. Read the deep dive:
Show more
Multimodal reasoning has a latency problem. More video frames leads to more waiting.
We built Damage Scout with Gemma 4 on Cerebras, running at over 2,300 toks/s, to show what fast multimodal inference unlocks.
Damage Scout samples frames from a rental car walkaround, sends them to Gemma 4, gets back structured findings and box coordinates, then renders an annotated damage report in under 6 seconds.
Same task. Same frames. A complete different experience powered by Cerebras ⚡️
Show more
We’ve reduced p95 latency by at least 25% across Realtime voice models through improved caching.
The bandwidth latency tradeoff:
Solana blocks are split into FEC sets, each of these sets is broadcast through rotor/turbine. Turbine splits FEC set into 32 pieces and then adds 32 more shreds containing erasure coding to form 64 shreds then sends each of these shreds out to the validator set through a specific randomly chosen turbine path.
How long does it take for each FEC set to get from the leader to a specific validator over turbine?
Because of erasure coding, the validator does not need every shred. It only needs enough distinct shreds to decode.
In the simple 32-of-64 case, the leader sends 64 shreds, but a validator only needs any 32 of them. So the slowest 32 paths do not matter for decoding.
We can model this mathematically: for validator i, define Lᵢ as the latency distribution induced by:
leader → random root → validator i
where the root is sampled according to stake. (this is technically rotor not turbine but its just a simplification, you can do the same trick for turbine but the equations are messier).
Each shred samples one relay path from Lᵢ.
So in the 32-of-64 case, validator i observes
X₁,…,X₆₄ ∼ Lᵢ
These are the arrival times of the 64 shreds.
But the relevant arrival time is
Bᵢ = X₍₃₂₎
the 32nd order statistic or the time when the 32nd fastest shred arrived, completing the FEC set. We can write the cumulative distribution of X₍₃₂₎ as:
Pr[Bᵢ ≤ t] = ∑ⱼ₌₃₂⁶⁴ (64 choose j) Lᵢ(t)ʲ(1−Lᵢ(t))⁶⁴⁻ʲ
More generally, if a slice has m data shreds and p coding shreds, then validator i sees
X₁,…,Xₘ₊ₚ ∼ Lᵢ
and can decode once m have arrived Bᵢ = X₍ₘ₎.
This turns Turbine design into a quantile/bandwidth tradeoff.
Let n = m+p and m/n → q. Then
Bᵢ ≈ Lᵢ⁻¹(q)
and by the CLT for order statistics,
√n · (X₍ₘ₎ − Lᵢ⁻¹(q)) ⇒ N(0, q(1−q)/fᵢ(Lᵢ⁻¹(q))²)
TLDR More coding shreds -> block arrives faster!
With more large validators moving to larger NICs and XDP activated, should we crank up the fan out to make the leader handoff faster if it costs us some theoretical max throughput?
Show more
SpaceX Starlink has lower latency than fiber for intercontinental distances. This is because light travel 50% faster through a vacuum, hopping between satellites before returning to Earth. It’s also doing less detours.
Sf to Mumbai
Fiber : 250 ms
Starlink: 120 ms
Show more
Realtime Ethereum delivers the biggest gains where latency already has the highest cost.
For a DEX, the swap that fills off-price.
For a wallet, the payment stuck "pending."
For a lender, the liquidation a block too late.
Better experiences, better loops, better apps.
Show more