We've been using similar concepts when building Vercel Data Pipeline
- Processing over 5GB/s
- Supports multi fanout
- Deduplication
- At-least once delivery
I think I need to write a blogpost
I'm pretty sure the only reason to use Kafka for telemetry data in 2026 is muscle memory.
A simple, open source, S3-backed pipeline can shuttle 1Gb/s of logs to @ClickHouseDB for $200/mo with 2.8s p50 latency. Fast enough for us.