We've migrated our data pipeline to use Vercel Queues.
Each queue item is just a lightweight marker that says, "There's a new chunk of data available." The actual data lives in S3.
Using Queues, we can now fan out processing into independent consumer groups, for example:
→ Insert the data into ClickHouse
→ Convert it to Parquet and write it back to S3
→ Sample it and send it elsewhere
Each consumer group is isolated, so if one falls behind or experiences issues, the others continue processing uninterrupted.
We've been using similar concepts when building Vercel Data Pipeline
- Processing over 5GB/s
- Supports multi fanout
- Deduplication
- At-least once delivery
I think I need to write a blogpost