Breaking
RUThe Red Stone checkpoint on the border between Russia and Belarus is temporarily closed due to a drone attackRUHigh-ranking sports officials detained in Turkey on corruption chargesITUstica massacre: the Rome Prosecutor's Office ready to withdraw the request for dismissal after the opening of FranceRUSan Jose placed Ilya Samsonov on waiversTRHeavy rain and storm warning from Istanbul GovernorshipRUThe court arrested journalist Yuri Dud in absentiaARQatar: Mediation efforts continue to bring views closer between Washington and TehranDESpain responds to housing shortage: Ban on property purchases by investment fundsESThe Government presents two decrees to stop the housing crisis with fiscal and financing measuresARThe repercussions of US sanctions and the war on daily life in Iran and the detention of a plane in IstanbulRUThe Red Stone checkpoint on the border between Russia and Belarus is temporarily closed due to a drone attackRUHigh-ranking sports officials detained in Turkey on corruption chargesITUstica massacre: the Rome Prosecutor's Office ready to withdraw the request for dismissal after the opening of FranceRUSan Jose placed Ilya Samsonov on waiversTRHeavy rain and storm warning from Istanbul GovernorshipRUThe court arrested journalist Yuri Dud in absentiaARQatar: Mediation efforts continue to bring views closer between Washington and TehranDESpain responds to housing shortage: Ban on property purchases by investment fundsESThe Government presents two decrees to stop the housing crisis with fiscal and financing measuresARThe repercussions of US sanctions and the war on daily life in Iran and the detention of a plane in Istanbul
BackThe impact of China’s export of AI data on the global knowledge system and Taiwan’s response strategies
The impact of China’s export of AI data on the global knowledge system and Taiwan’s response strategies
Tech
自由时报1 hour agoTech5 min readChinaView original

The impact of China’s export of AI data on the global knowledge system and Taiwan’s response strategies

How large language models learn Chinese, and how Taiwan actively responds to the impact of China’s information environment

Quick Look

  • Black Bear Academy explores the impact of China’s export of AI data on the global Chinese knowledge system.
  • Research shows that state control of the information environment will affect AI output through training data.
  • The author suggests that Taiwan should actively open up high-quality Chinese materials to ensure that AI can be exposed to multiple perspectives, not just China's position.

AI-generated summary

Why It Matters

Large language models are trained on public Internet data. Research shows that a country's degree of control over the information environment will affect the model's evaluation of the country's political leaders and institutions.

Font size

The New York Times headline: "China exports AI data, Beijing's discussion affects the global knowledge system" has caused concern, and many Taiwanese people believe that "AI equals information warfare."

At present, we do not know which Chinese websites are used by various commercial large-scale language models and what proportion each accounts for, because companies such as OpenAI, Anthropic, and Google have not disclosed a complete list of training materials.

What is certain is that "the open Internet is one of the important sources," but the Chinese Internet itself has a political structure that is very different from the English Internet. The vast majority of the world's Chinese-speaking population lives in China, and the Chinese government has long exercised a high degree of control over news, publishing, search engines, and online platforms.

So the question naturally arises: If a large language model reads a large amount of text left by humans, and the information environment of a certain language itself is systematically intervened by state power, will the traces left by these interventions also be learned by the model?

A study published in the magazine "Nature" in May 2026 is answering this question. The research team first compared the degree of media freedom in different countries and the tendency of large-scale language models to use local languages to answer political questions. They found: "The higher the degree of media control by the state, the more positive the model's evaluation of the country's political leaders and systems when answering in local languages."

The really important finding of this study is not at all that "U.S. AI has been controlled by the Chinese government", but it proves a more basic, almost common sense mechanism, that is, "the state's control of the information environment can further affect the output of large language models through training data."

On Taiwanese social media, you can occasionally see people reminding netizens not to answer questions from suspected Chinese public relations companies or information operation accounts. However, large-scale language models do not need to send out fake accounts and ask Taiwanese people question by question, "What do you eat for breakfast?" and "How do Taiwanese people call something?" in order to learn the language and culture of Taiwanese people.

Because the forums, blogs, public community posts, etc. accumulated over decades on the Internet in Taiwan are themselves a huge amount of language materials. The model can learn which words Taiwanese use and how to form sentences from the statistical relationships of a large number of texts. In comparison, it is an extremely inefficient data collection method to send people to operate a large number of fake accounts, wait for real people to answer one by one, and then use these fragmentary answers to train a large basic model.

Of course we still need to be vigilant. If what the other party wants to collect is the "real reaction" of "a specific age, political stance or social group" when facing a "specific problem", designing questions to induce real people to answer may still form valuable annotated data, which can be used for public opinion research, character simulation, model fine-tuning or information manipulation. But this is two different things from "China lacks Taiwanese corpus, so it must pretend to be a foreigner and ask Taiwanese people in order to train a large language model."

In fact, if we are truly worried that the Chinese knowledge of global large-scale language models will be overly affected by China’s information environment, the last strategy Taiwan should adopt is to hide its own data.

For Taiwan, instead of passively worrying about China's official stance contaminating the language model of the democratic world, it is better to think proactively: "How to make it easier for the global artificial intelligence system to obtain a large amount of high-quality and credible sources of Taiwanese Chinese materials"?

For example, government public information, court judgments, congressional records, news materials, historical and cultural collections, research results, and the large amount of text naturally produced by Taiwanese society may all form a Chinese world that is different from China's information environment.

If the Chinese government has realized that the data read by artificial intelligence in the future itself will be a kind of infrastructure, it has begun to actively provide data and AI models to the world. So what Taiwan really needs to face is not how to prevent its own language from being read by artificial intelligence, but how to ensure that when the next generation of artificial intelligence struggles to understand the Chinese-speaking world, it will not only read China’s position.

Open Questions

  • What is the specific proportion of Chinese data in the training materials of major AI companies?

Related Topics

This article was originally published by 自由时报.

Related Stories

A team from the University of Shanghai for Science and Technology develops single-beam multi-dimensional optical storage technology to achieve Pb-level storage capacity
Tech·

A team from the University of Shanghai for Science and Technology develops single-beam multi-dimensional optical storage technology to achieve Pb-level storage capacity

The team of Professors Gu Min and Zhang Qiming of the University of Shanghai for Science and Technology has developed single-beam multi-dimensional optical storage technology. By increasing the information capacity of a single storage point to 10 bits, the theoretical capacity of DVD-sized media reaches 0.4 Pb, and it has overcome the system problem of complex dual-beam collaboration.

中国新闻网
4 min read
University Exhibition Area of ​​the 26th Industrial Expo: Domestic wind power bearings, embodied robots and new varieties of river crabs unveiled
Tech·

University Exhibition Area of ​​the 26th Industrial Expo: Domestic wind power bearings, embodied robots and new varieties of river crabs unveiled

The 26th China International Industry Expo is about to open. Many universities in Shanghai are displaying scientific research results, including Shanghai Institute of Electric Power's localization plan for wind power main bearings, Shanghai Electric Power University's humanoid electric body-work robot, Shanghai Ocean University's new "White Jade Crab" strain, and Shanghai Maritime University's port digital twin platform.

中国新闻网
6 min read
The 2026 World Internet Conference Wuzhen Summit will be held in Wuzhen, Zhejiang in November
Tech·

The 2026 World Internet Conference Wuzhen Summit will be held in Wuzhen, Zhejiang in November

The 2026 World Internet Conference Wuzhen Summit is scheduled to be held in Wuzhen, Zhejiang from November 2 to 5. This summit has 27 sub-forums with the theme of "Open Source, Openness, Co-construction and Sharing", focusing on topics such as artificial intelligence governance, global digital cooperation and the development of digital intelligence technology, aiming to respond to the global "smart divide" challenge.

中国新闻网
2 min read
The 7th IGS Global Digital Cultural and Creative Conference was held in Chengdu, focusing on the upgrading of AI-enabled industries
Tech·

The 7th IGS Global Digital Cultural and Creative Conference was held in Chengdu, focusing on the upgrading of AI-enabled industries

The 7th IGS Global Digital Cultural and Creative Conference was held in Chengdu with the theme of "AI Empowerment, Ecological Integration". The conference brought together hundreds of companies to discuss the in-depth integration and industrial implementation paths of artificial intelligence in the fields of digital cultural creation, content production, cultural tourism consumption and IP development.

中国新闻网
4 min read
Liuzhou, Guangxi launches "Pension Benefit Actuary": Realizing "one person, one policy" for social security services
Tech·

Liuzhou, Guangxi launches "Pension Benefit Actuary": Realizing "one person, one policy" for social security services

Liuzhou City, Guangxi recently launched the "Pension Benefit Actuary" service, which uses big data and AI technology to provide insured people with accurate pension calculations and personalized payment plans through the "Liuzhou Social Security Code Uplink" applet and offline special windows, promoting the transformation of social security services from "passive question answering" to "active actuarial service".

中国新闻网
2 min read
More on this topicartificial intelligence