The impact of China’s export of AI data on the global knowledge system and Taiwan’s response strategies
How large language models learn Chinese, and how Taiwan actively responds to the impact of China’s information environment
Quick Look
- Black Bear Academy explores the impact of China’s export of AI data on the global Chinese knowledge system.
- Research shows that state control of the information environment will affect AI output through training data.
- The author suggests that Taiwan should actively open up high-quality Chinese materials to ensure that AI can be exposed to multiple perspectives, not just China's position.
AI-generated summary
Why It Matters
Large language models are trained on public Internet data. Research shows that a country's degree of control over the information environment will affect the model's evaluation of the country's political leaders and institutions.
The New York Times headline: "China exports AI data, Beijing's discussion affects the global knowledge system" has caused concern, and many Taiwanese people believe that "AI equals information warfare."
At present, we do not know which Chinese websites are used by various commercial large-scale language models and what proportion each accounts for, because companies such as OpenAI, Anthropic, and Google have not disclosed a complete list of training materials.
What is certain is that "the open Internet is one of the important sources," but the Chinese Internet itself has a political structure that is very different from the English Internet. The vast majority of the world's Chinese-speaking population lives in China, and the Chinese government has long exercised a high degree of control over news, publishing, search engines, and online platforms.
So the question naturally arises: If a large language model reads a large amount of text left by humans, and the information environment of a certain language itself is systematically intervened by state power, will the traces left by these interventions also be learned by the model?
A study published in the magazine "Nature" in May 2026 is answering this question. The research team first compared the degree of media freedom in different countries and the tendency of large-scale language models to use local languages to answer political questions. They found: "The higher the degree of media control by the state, the more positive the model's evaluation of the country's political leaders and systems when answering in local languages."
The really important finding of this study is not at all that "U.S. AI has been controlled by the Chinese government", but it proves a more basic, almost common sense mechanism, that is, "the state's control of the information environment can further affect the output of large language models through training data."
On Taiwanese social media, you can occasionally see people reminding netizens not to answer questions from suspected Chinese public relations companies or information operation accounts. However, large-scale language models do not need to send out fake accounts and ask Taiwanese people question by question, "What do you eat for breakfast?" and "How do Taiwanese people call something?" in order to learn the language and culture of Taiwanese people.
Because the forums, blogs, public community posts, etc. accumulated over decades on the Internet in Taiwan are themselves a huge amount of language materials. The model can learn which words Taiwanese use and how to form sentences from the statistical relationships of a large number of texts. In comparison, it is an extremely inefficient data collection method to send people to operate a large number of fake accounts, wait for real people to answer one by one, and then use these fragmentary answers to train a large basic model.
Of course we still need to be vigilant. If what the other party wants to collect is the "real reaction" of "a specific age, political stance or social group" when facing a "specific problem", designing questions to induce real people to answer may still form valuable annotated data, which can be used for public opinion research, character simulation, model fine-tuning or information manipulation. But this is two different things from "China lacks Taiwanese corpus, so it must pretend to be a foreigner and ask Taiwanese people in order to train a large language model."
In fact, if we are truly worried that the Chinese knowledge of global large-scale language models will be overly affected by China’s information environment, the last strategy Taiwan should adopt is to hide its own data.
For Taiwan, instead of passively worrying about China's official stance contaminating the language model of the democratic world, it is better to think proactively: "How to make it easier for the global artificial intelligence system to obtain a large amount of high-quality and credible sources of Taiwanese Chinese materials"?
For example, government public information, court judgments, congressional records, news materials, historical and cultural collections, research results, and the large amount of text naturally produced by Taiwanese society may all form a Chinese world that is different from China's information environment.
If the Chinese government has realized that the data read by artificial intelligence in the future itself will be a kind of infrastructure, it has begun to actively provide data and AI models to the world. So what Taiwan really needs to face is not how to prevent its own language from being read by artificial intelligence, but how to ensure that when the next generation of artificial intelligence struggles to understand the Chinese-speaking world, it will not only read China’s position.
Open Questions
- What is the specific proportion of Chinese data in the training materials of major AI companies?



