Register and share your invite link to earn from video plays and referrals.

Search results for DataInfrastructure
DataInfrastructure community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DataInfrastructure
🏗 Migrating tens of thousands of jobs that ingest petabytes a day, without ever stopping data delivery. Meta's playbook of shadow then reverse-shadow then cleanup deprecated the legacy system 100%. Title: Migrating Data Ingestion Systems at Meta Scale URL: 📝 Overview Meta incrementally scrapes several petabytes of social-graph data daily from one of the world's largest MySQL deployments into its data warehouse. This post explains how they migrated tens of thousands of those ingestion jobs to a new self-managed service without disrupting analytics, reporting, and ML pipelines. ❓ Challenges Solved The legacy system was customer-owned pipelines, fine at small scale but unstable at hyperscale. They had to meet increasingly strict data landing-time requirements while migrating without interrupting data delivery across the organization. 💡 Methodology & Proposed Approach They migrate through a three-phase lifecycle. ・Shadow: in pre-production, consume production data while writing to isolated tables, continuously monitoring row-count and checksum mismatches against production jobs ・Reverse shadow: promote shadow jobs to production tables and send the original production jobs to shadow, keep comparing outputs for quality signals, and roll back instantly if needed ・Cleanup: deprecate old jobs after confirming consistency ・Each job is verified on four axes (zero differences, landing latency, resource usage, custom criteria), and CDC maintains full-dump, delta, and target tables 🎯 Use Cases It informs migrating large data-ingestion platforms, phasing CDC pipeline cutovers, and designing zero-downtime system replacements. 📊 Outcomes ・100% of the workload was migrated and the legacy system fully deprecated ・Job status signals streamed continuously to Scuba, and a migration tool monitored each job and auto-promoted/demoted between stages to manage thousands of concurrent migrations ・Bad partitions were flagged in metadata to prevent propagation to downstream jobs and trigger alerts ・To handle capacity limits, they reused old-system snapshot partitions as initial snapshots to cut full-dump load, and the resulting data-quality analysis tool is still used in release validation after the migration #DataEngineering# #DataInfrastructure#
Show more
The Data Outpost talk agenda is live. Nov. 4 is hands-on-keyboard: a full day of workshops for people building with data & AI. Nov. 5 is talks and panels with founders, engineers, product leaders, researchers, operators, and data architects making machine intelligence useful inside actual companies. We’re talking agents, interfaces and workflows, memory, trust, and the data infrastructure underneath all of it. 25+ speakers. 2 days. $499. 📍 San Francisco 🗓️ Nov 4–5 🔗
Show more
. @Silicon_Data and @computeexchange were both built after the ChatGPT moment. But I still wouldn’t call either company truly AI-native—yet. Being founded in the AI era doesn’t automatically make an organization AI-native. Giving every employee access to ChatGPT certainly doesn’t. I’ve been thinking about the organizational structures of both companies, and the exercise has made me realize that AI-native organizations will not all look the same. @Silicon_Data is organized around building the independent reference layer for the compute economy: data infrastructure, indices, benchmarking, research, product commercialization and market adoption. @computeexchange is organized around creating liquidity: sourcing, verification, pricing, matching, contracting and settlement. Agents can transform both companies—but differently. At @Silicon_Data, agents can accelerate data analysis, research, product development, content production and customer intelligence. At @computeexchange, they can automate inventory normalization, provider onboarding, RFQs, matching and transaction workflows. This has also changed how I think about organizational design. Traditional companies are built around people, roles and reporting lines. Knowledge is distributed across individual brains, inboxes, documents, Slack channels and meetings. In that sense, a human organization is web-based: every person is a node, and work moves through the relationships connecting those nodes. An agent organization may be fundamentally different. It is Brain-based. Instead of every agent holding a fragmented version of the company, agents can operate from a centralized institutional Brain containing shared knowledge, history, decisions, priorities, permissions and real-time operating context. Each Brain sits a task-ownership system. Instead of asking, “Whose job is this?” the organization asks: What needs to be accomplished? What context and authority does it require? Should a human, an agent or a human-agent team own it? What constitutes completion? Who remains accountable? Humans continue to operate through networks of relationships, judgment, negotiation and trust. Agents operate through centralized knowledge, shared context and structured task ownership. The task layer tells you what needs to happen, who—or what—owns it, and whether it has actually been completed. To me, becoming AI-native means continuously redesigning this boundary between people, agents, knowledge and work. We are still experimenting. I’ll share what works, what fails, and how the two organizations evolve.
Show more
API Developer Resources Robust data infrastructure demands clear documentation. 🔹 Fetch live market prices and historical OHLCV 🔹 Build custom price widgets and data screeners 🔹 Deploy exact code scripts via Python and Node 7/8
Show more
Looking to build the data infrastructure powering onchain finance? We're hiring a Senior Platform Engineer to strengthen the cloud and production systems behind Chronicle's data infrastructure for tokenized assets and onchain markets. If you have deep experience with AWS, Kubernetes, distributed systems, and building reliable production environments, we'd love to talk. 🚀 Apply here:
Show more
Looking to build the data infrastructure powering onchain finance? We’re hiring a Head of Operations at Chronicle. This is a hands-on leadership role owning the day-to-day operations of the business across finance, people, compliance, and internal systems. 📍 Remote, Europe Apply here:
Show more
Looking to build the data infrastructure powering onchain finance? We're hiring a Head of Platform Engineering to lead the systems behind Chronicle's data infrastructure for tokenized assets and onchain markets. If you've built and scaled high-performance distributed systems, we'd love to talk. 🚀 Apply here:
Show more
LATEST: ⚡️ Data infrastructure firm Inveniam Capital Partners plans to acquire Mantra, expanding its push into RWA tokenization and AI infrastructure.
Diamond partner Bright Data is what 70%+ of the world's leading AI labs use to train their models. From large-scale video data for robotics training to reliable agentic web access in production - it's the web data infrastructure the AI industry runs on. SuperAI Singapore, 10-11 June.
Show more
Truth should be verifiable. Data should be transparent. Trust should be decentralized. Introducing xtruth ⚖️ The oracle layer built for XLayer. Secure, fair, and real-time on-chain data infrastructure for the next generation of Web3 applications. #XLayer# #Web3# #Oracle# #DeFi#
Show more