Register and share your invite link to earn from video plays and referrals.

Search results for DataInfrastructure
DataInfrastructure community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DataInfrastructure
🏗 Migrating tens of thousands of jobs that ingest petabytes a day, without ever stopping data delivery. Meta's playbook of shadow then reverse-shadow then cleanup deprecated the legacy system 100%. Title: Migrating Data Ingestion Systems at Meta Scale URL: 📝 Overview Meta incrementally scrapes several petabytes of social-graph data daily from one of the world's largest MySQL deployments into its data warehouse. This post explains how they migrated tens of thousands of those ingestion jobs to a new self-managed service without disrupting analytics, reporting, and ML pipelines. ❓ Challenges Solved The legacy system was customer-owned pipelines, fine at small scale but unstable at hyperscale. They had to meet increasingly strict data landing-time requirements while migrating without interrupting data delivery across the organization. 💡 Methodology & Proposed Approach They migrate through a three-phase lifecycle. ・Shadow: in pre-production, consume production data while writing to isolated tables, continuously monitoring row-count and checksum mismatches against production jobs ・Reverse shadow: promote shadow jobs to production tables and send the original production jobs to shadow, keep comparing outputs for quality signals, and roll back instantly if needed ・Cleanup: deprecate old jobs after confirming consistency ・Each job is verified on four axes (zero differences, landing latency, resource usage, custom criteria), and CDC maintains full-dump, delta, and target tables 🎯 Use Cases It informs migrating large data-ingestion platforms, phasing CDC pipeline cutovers, and designing zero-downtime system replacements. 📊 Outcomes ・100% of the workload was migrated and the legacy system fully deprecated ・Job status signals streamed continuously to Scuba, and a migration tool monitored each job and auto-promoted/demoted between stages to manage thousands of concurrent migrations ・Bad partitions were flagged in metadata to prevent propagation to downstream jobs and trigger alerts ・To handle capacity limits, they reused old-system snapshot partitions as initial snapshots to cut full-dump load, and the resulting data-quality analysis tool is still used in release validation after the migration #DataEngineering# #DataInfrastructure#
Show more
Looking to build the data infrastructure powering onchain finance? We're hiring a Senior Platform Engineer to strengthen the cloud and production systems behind Chronicle's data infrastructure for tokenized assets and onchain markets. If you have deep experience with AWS, Kubernetes, distributed systems, and building reliable production environments, we'd love to talk. 🚀 Apply here:
Show more
Looking to build the data infrastructure powering onchain finance? We’re hiring a Head of Operations at Chronicle. This is a hands-on leadership role owning the day-to-day operations of the business across finance, people, compliance, and internal systems. 📍 Remote, Europe Apply here:
Show more
Looking to build the data infrastructure powering onchain finance? We're hiring a Head of Platform Engineering to lead the systems behind Chronicle's data infrastructure for tokenized assets and onchain markets. If you've built and scaled high-performance distributed systems, we'd love to talk. 🚀 Apply here:
Show more
LATEST: ⚡️ Data infrastructure firm Inveniam Capital Partners plans to acquire Mantra, expanding its push into RWA tokenization and AI infrastructure.
Diamond partner Bright Data is what 70%+ of the world's leading AI labs use to train their models. From large-scale video data for robotics training to reliable agentic web access in production - it's the web data infrastructure the AI industry runs on. SuperAI Singapore, 10-11 June.
Show more
Truth should be verifiable. Data should be transparent. Trust should be decentralized. Introducing xtruth ⚖️ The oracle layer built for XLayer. Secure, fair, and real-time on-chain data infrastructure for the next generation of Web3 applications. #XLayer# #Web3# #Oracle# #DeFi#
Show more
Based on a true data infrastructure story AI builders are giving Shelby ⭐⭐⭐⭐⭐
Trustworthy AI can’t exist without trustworthy data. And right now, the data behind AI is becoming harder to verify. Models are training on synthetic content, scraped datasets, and feedback loops that are difficult to trace. Human review still happens, but it is often anonymous, fragmented, and disconnected from any lasting record of who contributed, what they verified, or how reliable their work was. But by verifying contributors, tracking performance, and recording each validation step, AI data can become accountable. Experts can build reputation over time. High-quality work can be routed to higher-value tasks. Enterprises can see the human judgment behind the datasets their systems rely on. It’s why we’re building Perle Labs: expert-validated, human-verified, on-chain auditable data infrastructure for AI systems that need to be trusted in the real world.
Show more
Truly determining AI upper limit is data modeling capability. This insight gives us more confidence in PAN Project's future data infrastructure! New ideas. New connections. Back to accelerate crypto × AI. 🚀 #AI# #Crypto#
Show more