
大型語言模型如何學習中文,以及台灣在面對中國資訊環境影響時的積極應對之道
黑熊學院探討中國輸出AI數據對全球中文知識體系的影響。研究顯示國家對資訊環境的控制會透過訓練資料影響AI輸出。作者建議台灣應積極開放高品質中文資料,確保AI能接觸多元觀點,而非僅限於中國立場。
AI-generated summary
大型語言模型透過公開網路數據進行訓練。研究顯示,國家對資訊環境的控制程度會影響模型對該國政治領導人與制度的評價。
紐時標題:「中國輸出AI數據,北京論述影響全球知識體系」引起擔憂,也有不少台灣人認為「AI等於資訊戰」。
目前我們並不知道各家商業大型語言模型究竟使用了哪些中文網站、各自占多少比例,因為OpenAI、Anthropic、Google等公司並未公開完整的訓練資料清單。
可以確定的是「公開網路是重要來源之一」,但中文網路本身具有一個與英語網路很不一樣的政治結構。全球絕大多數中文使用人口居住於中國,中國政府又長期對新聞、出版、搜尋引擎與網路平台進行高度控制。
於是問題自然就浮現了:如果大型語言模型大量閱讀人類留下的文字,而且有某一種語言的資訊環境本身受到國家力量系統性介入,這些介入留下的痕跡,會不會也被模型一起學進去?
2026年5月刊登於《自然》(Nature)雜誌的一份研究,正是在回答這個問題。研究團隊首先比較不同國家的媒體自由程度與大型語言模型使用當地語言回答政治問題時的傾向,發現:「媒體受到國家控制程度愈高,模型以當地語言回答時,對該國政治領導人與制度的評價也往往更加正面。」
這項研究真正重要的發現完全不是「美國AI已經被中國政府控制」,而是證明了一個更基本、近乎常識的機制,亦即「國家對資訊環境的控制,可以透過訓練資料進一步影響大型語言模型的輸出」。
在台灣社群媒體上,時不時可以看到有人提醒網友,不要回答疑似中國公關公司或資訊操作帳號提出的問題。但大型語言模型不需要派出假帳號,一題一題詢問台灣人「你們早餐吃什麼」、「台灣人怎麼稱呼某樣東西」,才能學會台灣人的語言與文化。
因為台灣網路上數十年累積的論壇、部落格、公開社群貼文等,本身就是規模極大的語言材料。模型可以從大量文字的統計關係裡學習台灣人使用哪些詞彙、如何造句。相較之下,派人經營大量假帳號、等待真人逐一回答,再拿這些零碎回答訓練一個大型基礎模型,是極度沒效率的資料蒐集方法。
當然我們還是要有所警覺。如果對方想蒐集的是「特定年齡、政治立場或社會群體」面對「特定問題」時的「真實反應」,設計問題誘導真人回答,仍然可能形成有價值的標註資料,可以用於輿情研究、人物模擬、模型微調或資訊操作。但這與「中國缺乏台灣語料,所以必須假裝外國人來問台灣人,才能訓練大型語言模型」是兩件不同的事情。
事實上,如果我們真正擔心全球大型語言模型的中文知識受到中國資訊環境過度影響,台灣最「不」應該採取的策略,就是把自己的資料藏起來。
對台灣而言,比起消極的擔心中國官方立場污染民主世界的語言模型,不如積極思考:「如何讓全球人工智慧系統更容易取得大量、高品質而且具有可信來源的台灣中文資料」?
例如政府公開資料、法院判決、國會紀錄、新聞資料、歷史文化典藏、研究成果,以及台灣社會自然產生的大量文字,都可能構成與中國資訊環境不同的中文世界。
若中國政府已經意識到,未來人工智慧所閱讀的資料本身就是一種基礎建設,因此開始積極向全球提供資料與AI模型。那麼台灣真正需要面對的,就不該是如何阻止自己的語言被人工智慧讀到,應是如何確保當下一代人工智慧寒窗苦讀試著理解中文世界時,讀到的不是只有中國立場。

The 2026 World Internet Conference Wuzhen Summit is scheduled to be held in Wuzhen, Zhejiang from November 2 to 5, with the theme of "Open Source, Openness, Co-construction and Sharing". The summit will focus on the development of artificial intelligence, security governance and international cooperation, and will hold the "Light of the Internet" expo and a number of technology competitions, attracting participation from many countries around the world.

The team of Professors Gu Min and Zhang Qiming of the University of Shanghai for Science and Technology has developed single-beam multi-dimensional optical storage technology. By increasing the information capacity of a single storage point to 10 bits, the theoretical capacity of DVD-sized media reaches 0.4 Pb, and it has overcome the system problem of complex dual-beam collaboration.

The 26th China International Industry Expo is about to open. Many universities in Shanghai are displaying scientific research results, including Shanghai Institute of Electric Power's localization plan for wind power main bearings, Shanghai Electric Power University's humanoid electric body-work robot, Shanghai Ocean University's new "White Jade Crab" strain, and Shanghai Maritime University's port digital twin platform.

China's generative AI user base reached 700 million by the end of June, representing over half the population. Data from the China Internet Network Information Centre shows a 16% increase since late 2025, with widespread use in Q&A, media processing, and smart hardware.

The 2026 World Internet Conference Wuzhen Summit is scheduled to be held in Wuzhen, Zhejiang from November 2 to 5. This summit has 27 sub-forums with the theme of "Open Source, Openness, Co-construction and Sharing", focusing on topics such as artificial intelligence governance, global digital cooperation and the development of digital intelligence technology, aiming to respond to the global "smart divide" challenge.

The 7th IGS Global Digital Cultural and Creative Conference was held in Chengdu with the theme of "AI Empowerment, Ecological Integration". The conference brought together hundreds of companies to discuss the in-depth integration and industrial implementation paths of artificial intelligence in the fields of digital cultural creation, content production, cultural tourism consumption and IP development.