Register and share your invite link to earn from video plays and referrals.

Ali Ansari
@aliansarinik
2.8K Following    28.9K Followers
Today we’re launching micro1’s PII transformation model, flow-transform 1.0, delivering frontier-level performance across detection, identity synthesis, and transformation of personally identifiable information. On PrivacyBench, our model reaches 96.0% F1, outperforming every detection baseline we tested, including Tonic Textual, Claude Opus 4.8, Sonnet 4.6, Microsoft Presidio, Haiku 4.5 and GLiNER2. Some of the most valuable training data for frontier AI models lives inside fully functioning companies. It captures years of real work across decisions, communications, tools, handoffs, exceptions and the relationships connecting them. The problem is that this data is also full of PII. Traditional redaction makes the data safe, but it also destroys the very workflows and relationships frontier models need to learn from. flow-transform 1.0 solves this by turning enterprise operational data into high-fidelity training data for frontier models by replacing real-world identities without flattening the reality the data captures.
Show more
0
90
656
110
Forward to community
in the past 24 hours alone, micro1 has paid out $5.8M to 9 businesses for their enterprise data. that’s an average of over $600k per business. frontier AI needs training data that captures the complexity of real-world work. real businesses hold decades of decisions, exceptions and learnings that make the next breakthroughs possible. realism is a new dimension of scale & arguably the most important ingredient for model training data. if your company is interested in a data partnership, reach out to us.
Show more
Why we bid $12.5 million for Spirit Airlines’s data: As you may have seen in the news, we are aiming to acquire the Spirit operational data. We'd like to transparently & directly explain why. The future of AI is a bet on two things: the messiness of the real world and the brilliance of the humans working inside it. The combination of these two things results in a new dimension of scale called realism. Realism is the most significant thing that has ever happened to model training. It’s what lets RL environments and tasks match the conditions a model is deployed into. The limit is the real world itself. But the real world is also dynamic. It grows with a company and shifts as operations evolve. There is not a binary scaling of entering the real world. Realism keeps scaling as long as real work keeps happening. As a data lab, we want this data to benefit the entire AI ecosystem, while ensuring labs that we partner with for this data, agree to strict privacy obligations. This includes keeping the data away from labs outside of U.S. and U.S-allied countries. Our mission is to push the frontier in novel ways, with a human & privacy first approach. To do that, we've bid on a portion of Spirit’s Airline’s non-sensitive operational data. Effectively scaling realism requires de-identification as a core competency. Very few organizations can do this. We believe the few that can have an obligation to act transparently. That’s why we’ve made binding commitments. We will prohibit re-associating the data and will not profile any former employee. We’re paying for an independent data ombudsman to review the data and how we use it. Furthermore, our offer aims to create net new opportunities for as many as we can from the 17,000+ former Spirit employees. We’re providing former Spirit employees preferential access to paid human data roles on our platform, along with free access to our AI training and reskilling curriculum. These commitments bind us to spending at least $1M. That is just the minimum. It’s entirely possible we spend orders of magnitude more than $1M working with and training former Spirit employees to help define the frontier. We think this transaction will set the terms for how data like this changes hands from here. Those terms should be worth copying.
Show more
close to 20,000 people have signed up in 6 hours. we are onboarding about 2,000 over next 12h. biggest robotics training project is being set up as we speak.
we’re hiring 10,000 robotics trainers in the next 7 days. $50–$90/hour, accepting applicants globally. you’ll review and label videos of robots performing tasks to help them improve. no prior AI experience required. an entirely new category of work is emerging around teaching robots how to interact with the world. application link in the comments below.
Show more
0
464
7.2K
499
Forward to community
fun scroll for our experts: :)
Excited to share that micro1 has been selected to support the U.S. Department of Energy’s Genesis Mission. Genesis is a Manhattan Project-scale national effort to accelerate scientific discovery with AI and build the energy foundation needed to power America’s AI leadership. micro1 will support the mission by developing the data and evaluations that enable frontier AI to tackle some of the country’s most difficult scientific and engineering challenges. We’re proud to play a role in an effort this important to the future of American science, energy, and AI. More on our work with the DOE linked in the comments. 🇺🇸
Show more
The future of AI belongs to the humans behind it. There’s a common fear that as AI gets better, people get pushed out of the picture. We believe the opposite is happening. AI is creating entirely new categories of work, and an entirely new economy around the people whose knowledge, judgment, and experience are helping these systems improve. Today, tens of thousands of experts are actively contributing to AI training projects through micro1. Over time, we believe this will grow to tens of millions+. And if humans are going to play such a critical role in building the future of AI, the companies they work with should raise the standard for how they’re supported. Today, I’m proud to announce the micro1 Expert Support Program. We’re building a new set of protections, resources, and support for our expert community, starting with: -Expert Bill of Rights: a clear set of commitments outlining what experts can expect when working with micro1. -Rest Credits: paid time away from projects when experts need or want a break. -Expert Emergency Fund: financial support for experts facing emergencies. -Mental Health & Coaching: new resources to support expert wellbeing and growth. -Confidential Support Line: a confidential way to ask questions, seek support, or raise concerns related to pipelines. Within the next two weeks, every expert currently active with micro1 will receive an email with the full program details and launch dates. AI is going to keep getting more capable. The opportunity in front of us is to make sure the people helping build it benefit from that progress too.
Show more
and over $2M in referral payouts for companies sent over last few weeks. 🚀
more than 1,000 companies have signed up to get paid for their anonymized data to train models just in the last few weeks. a massive TAM opportunity for the entire economy has emerged almost overnight.
Show more
more than 1,000 companies have signed up to get paid for their anonymized data to train models just in the last few weeks. a massive TAM opportunity for the entire economy has emerged almost overnight.
Show more
in the last 11 days, we've committed more than $20,000,000 to license real operational data to seed our RL environments. scaling on the realism dimension is just beginning.
Regarding the last topic @DavidSacks : With respect, the data being used to train frontier models in the U.S. is not a commodity. It is American intelligence, paired with anonymized operational data from U.S. enterprises. Put simply, we have top doctors, scientists, lawyers, and physicists in the U.S. using their knowledge to create detailed rubrics that train these models. Each individual data point and its corresponding rubric can take an expert anywhere from 10 to 30 hours to create. This is not preference labeling or drawing bounding boxes. It is highly complex, structured human judgment from leading experts here in the West—people who deeply understand and directly contribute to the latest American innovations in their respective fields. That expertise is then converted through highly specific data structures and RL environments (developed collaboratively by U.S. AI labs and data labs) from raw human intelligence into high-signal rewards that improve frontier models. On top of that, any AI advancement, even something that begins as a simple chatbot designed to improve operations within a defense agency, can create a major competitive advantage in adversarial situations and may have dual-use applications. Lastly, many datasets today are seeded with anonymized, real-world operational data from U.S. companies to build highly realistic environments. When those datasets are sold to China, we are not simply exporting “labeling.” We are exporting proprietary American intelligence, structured for machine learning and delivered directly to China at scale. It is very easy to categorize this work as “labeling” and ignore what it actually represents. However, as @altcap suggested: “Then these things will get a lot more scrutiny than they’re getting today. I think the only reason they pass muster today is because we’re still leading the race.” If China catches up to U.S. labs, this will become much harder to ignore. In retrospect, the role of data as the root cause will become very clear — and by that point, it may be too late. The race is tight. I suggest looking into this now.
Show more
POD UP!🚨 Fifth Bestie @altcap Joins the Show! Brad fills in for @chamath and the Besties discuss: -- Google's AI Shakeup: Brain Drain or Strategy? $GOOG -- SpaceX's Massive Quarter: Terafab, Capex, EWS, $1T Projection $SPCX -- Airtable Sells for a 90% Discount, Signs of SaaSpocalypse? $BSP -- US Data is Fueling Chinese AI (0:00) Bestie intros! Brad Gerstner fills in for Chamath (2:16) Major shakeups at Google: AI brain drain or better strategy? (20:39) SpaceX's big quarter: Terafab, AI Capex, $1T revenue projection? (45:44) All-In Summit Speaker Announcements! (48:01) Airtable sells for a 90% discount: SaaSpocalypse? (1:05:56) Chinese AI labs are buying US training data to catch up
Show more
Shameful act to optimize for short term revenue increases and serve an adversarial nation in the most important race of our lifetime. The only way models improve is through data. if you have the recipe on what data pushes the frontier and send that to China, you are doing a disservice to U.S. AI Labs as well as the U.S. government. Beyond that, a large portion of datasets today require anonymized real operational data from U.S. companies. Selling such environments to Chinese labs means exporting U.S. company data directly to China at a large scale. This must stop.
Show more
The AI arms race isn’t just being fought over Nvidia chips. Chinese labs are matching OpenAI and Anthropic by purchasing the exact same data from US vendors like @mercor and Surge AI (who also work with the US federal govt)
Show more
human brilliance is needed more than ever. our mission is to ensure that's always true. m1 in 🇬🇧
there are two dimensions that data demand is scaling on tremendously: horizon and realism of tasks/data points. environments built on top of anonymized real operational data is step function change in scaling on both of these dimensions. and as we do this, companies that choose to contribute can make data a very meaningful portion of their revenue while accelerating their path towards becoming more AI native. thanks for covering Stephanie!
Show more
ICYMI: I chatted w/ the CEO of a small Michigan-based HVAC company about the data labeling work they're doing for micro1—and why these data startups are increasingly buying data from SMBs.
Show more
The more we think about robotics, the more it seems the field is moving toward increasingly specialized datasets. Better foundation models increase the value of task/environment-specific datasets. As base models become more capable, annotation becomes more about data structuring. The challenge shifts toward adapting data to the environments, objectives, and edge cases a model will encounter. Over the past several months, our robotics team has been running a large number of experiments around the question of how do you maximize the training signal from the same raw data? One conclusion we've become increasingly convinced of is that there isn't a universal annotation pipeline for robotics. Different tasks require different combinations of models, verification, and human expertise to produce the highest-quality datasets. The great @AndrewLeeMaas and Mitali Potnis from our team have put together a paper walking through the ideas and experiments behind this approach. The full paper showcasing examples of how we have applied these annotations is in the comments.
Show more
the immense investment in AI has not translated into adoption at the scale it appears to have. across comparable early periods, AI investment been growing approximately 40%/year, versus 25% for electricity and 16% for IT. however, only 19.8% of U.S. businesses use AI in at least one function, roughly 5 points below both benchmarks. meanwhile the real production deployment gap is even deeper at top companies: recent MIT study shows that only 11% of the S&P500 has "deeply" integrated AI. capability is advancing faster than trust and real implementation. there's one fundamental reason for this: enterprise investment in evaluations is less than 0.1% of where it needs to be. once large enterprises start investing in continuous evaluation loops, they will precisely define what "good" means for their use case. within that context, they will measure the intelligence of any given system. when you measure something, you can improve it and watch performance increase against those measurements. this continuous loop is how enterprises "own their own intelligence." it's irrelevant whether you're building on a closed or open model. in most cases, closed models will perform much better for your use case. what matters is whether you're precisely defining what good means and consistently measuring against it. if you are, then you're owning your own intelligence in a largely model-agnostic way.
Show more
some human data companies work with foreign adversaries. and the results show today in Kimi K3. we believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with. people come to America to build extraordinary things for humanity. we must protect that brilliance through American AI dominance.
Show more
the demand for intelligence improvement will be in the same order of magnitude as intelligence itself (and potentially even more). and the rapid increase in RL environments / data spend is the very beginning of this playing out. a world where you can predictably buy more "IQ" points for your AI system is a world where every single enterprise dedicates a large portion of their budget buying such units of intelligence improvements. the trillion dollar market & the largest job sector ever is just now getting started.
Show more