Register and share your invite link to earn from video plays and referrals.

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal
11.8K Following    17.1K Followers
A little bit technical, but will help you understand a potential approach to continual learning...
Growth of AI data labelling industry in China: If your mental model is distillation is the major route for getting data for China based labs, where do you think SeeDance 2 of ByteDance got their data? They certainly did not distill SORA 2. SeeDance 2 is much superior to Sora 2. ByteDance has large internal data labelling operation curating data for their models. SeeDance is closed source, but the point to note is they curate their own data and are world leader in that space. China has both: AI data labelling companies like Datatang and teams at large companies doing it in-house like at ByteDance, Alibaba, Baidu, Tencent etc. Furthermore, there are large crowd sourcing platforms like Baidu Crowd and Alibaba Cloud PAI-iTAG. In fact, data labelling is one of the fastest growing job roles in China. Quote from Global Times: "According to the Digital China Development Report (2025) released by the National Data Administration in May, seven data annotation bases have been established in the country, with a total workforce of 95,000 data annotation professionals as of the end of 2025" Contrary to perception - data is not so difficult to get/curate. It requires lot of diligence though and may cost a lot. While in the west Mercor, Turing etc. have rebranded themselves as "Intelligence Clouds" they have roots in IT body shopping/providing elite IT human resources. China has multiple such companies. These companies provide the human labor for AI tasks that works on hourly basis. For example: 1. Proginn (程序员客栈 - Chuangyeyuan Kezhan) 2. Yuanwuxian (猿森林 / 猿五线) & Jiedan Platforms They provide highly technical, engineering-grade human resources for the most complex stages of the AI testing and alignment lifecycle Also, companies all over the world are accessible to labs from China (closed, as well as, open source). There are no trade restrictions on buying and selling data, as far as, I know. This is a crowded space and many providers are looking for buyers. Let us review task types and how to get data for them: Verifiable domains: For post training particularly, there are already vast amount of problem statements and solutions in millions of text books. Also, synthetically generating data (e.g. introducing bugs in programs) is an effective strategy. This works for verifiable domains like Code. Math problem statements can also be generated systematically and their solutions can be verified automatically. Knowledge work domains: As for tasks like GDPEval etc. (e.g. spreadsheet making, making presentations etc.) you do need experts curating problem statements and ideal solutions; but it is not a rocket science. Complex problems in science and engineering, you need to hire bunch of professors and phds (e.g. SpaceX hires Olympiad winners globally) to make AI models fail and discover their weaknesses and design tasks to train them to remove the weaknesses. China has more of them that the entire developed world combined. They also have very diverse expertise due to China's vast industrial base. RL environments can be somewhat complex to make: you have build mock web services, mock apps etc. But, with strong models these days such environments can be designed with ease. Collecting computer use data and egocentric data is very very human labor intensive. There too China has a huge advantage, with large number of youth available, it is easy and not that expensive. The reason you don't hear a lot about Chinese data labelling companies is the same why you don't hear a lot about Chinese IT service companies: they have limited themselves to the mainland where they have ample business. E.g. Chinasoft International has 80K employees, but I bet most haven't heard about them. Also, many large Chinese tech companies have their own data labelling units or they directly manage their contractors. One of the reason one should care about this topic is because, safe open source AI is important for a more just & equitable world.
Show more