China unveils massive plan to build AI training datasets as global data shortage looms

1 hour ago 2



China’s National Data Administration just made its most explicit admission yet: the country doesn’t have enough data to feed its AI ambitions. On June 9, the agency released a draft plan designed to massively expand the supply, circulation, and commercialization of high-quality training datasets across the Chinese economy. Global estimates suggest that the pool of publicly available human-generated data could shrink dramatically between now and 2032. China is essentially trying to build a strategic reserve before the well runs dry. What the plan actually covers The NDA’s blueprint spans nine core industrial sectors, including healthcare, finance, and transportation, along with emerging fields like autonomous driving. It isn’t just about hoarding text. The plan explicitly calls for multimodal datasets covering text, code, images, audio, and video. The target deadline is 2028. By then, the NDA wants a comprehensive ecosystem of multimodal datasets ready for commercial use across those sectors. Alongside the dataset initiative, the Chinese government is pouring approximately $295 billion, roughly 2 trillion yuan, into AI-focused data centers over the next five years. The plan also dev...

Read Entire Article