
AI-generated summary
The development of artificial intelligence has entered the AI4S stage, and scientific research is increasingly relying on large models, but large models require high-quality scientific corpus as a basis. At present, scientific and technological corpus resources are scattered, standards are not unified, and safe circulation capabilities are insufficient, which restricts model training and scientific reasoning.
The Corpus Innovation Ecological Alliance (hereinafter referred to as the "Alliance") initiated by the Documentation and Information Center of the Chinese Academy of Sciences was recently launched in Beijing. The corpus standard specifications, Aocang Corpus Community Service Platform and the English journal Data Express in AI Corpus were simultaneously released. On the same day, Aocang Science and Technology Corpus signed cooperation agreements with key corpus application institutions and the National Artificial Intelligence Application Pilot Base.
Why do this?
The answer lies not in the corpus itself, but in scientific research.
Currently, artificial intelligence is changing the scientific research paradigm. From experimental induction, theoretical deduction, computational simulation to data-driven, and then to the deep integration of AI and scientific research (AI for Science, hereinafter referred to as "AI4S"), scientific research increasingly relies on large models to participate in scientific reasoning and knowledge discovery. However, in order for large models to "understand" science, they rely not on general Internet data, but on high-quality scientific and technological corpus.
And the problem arises from this.
The deeper AI4S is, why is it more difficult to bypass the “corpus barrier”?
Scientific research needs AI. What does AI need?
Sun Ninghui, founder of the alliance and academician of the Chinese Academy of Engineering, got straight to the point: "The essence of intelligence is the refinement of data. Large models are the product of deep processing of all Internet data. But corpus in the field of science and technology is different. It has scientific semantics, physical constraints, and error information. It cannot be used simply by aggregating it."
"Compared with general-purpose corpus, scientific and technological corpus has more diverse sources and more complex professional semantics, which puts forward higher requirements for scientific accuracy, knowledge density, modal association and traceability." Lu Fangjun, director of the Bureau of Basic Science and Technology Capabilities of the Chinese Academy of Sciences, further explained.
Realistic bottlenecks are also highlighted: high-value scientific and technological resources are still scattered and difficult to directly transform into high-quality scientific and technological corpus.
"The resource aggregation and authorization mechanism needs to be improved, and standards, professional processing, quality evaluation and safe circulation capabilities still need to be strengthened." Lu Fangjun explained that scientific and technological corpus connects original scientific research resources and artificial intelligence models and is an important foundation to support model training, scientific reasoning and knowledge discovery.
It can be seen that the construction of high-quality scientific and technological corpus has become an important basic task for developing scientific intelligence and improving professional capabilities in artificial intelligence.
Why is it difficult for a single organization to do this? Lu Fangjun said bluntly: Because there is a lack of close coordination mechanism between the corpus supplier, technology research and development side and application demand side. These problems cannot be solved by one unit.
Looking deeper, the competitive landscape is also changing.
Li Mengli, Secretary of the Party Committee of the Documentation and Information Center of the Chinese Academy of Sciences, judged that the global artificial intelligence competition is deepening from a competition of single technology and model capabilities to a comprehensive competition in the coordinated development of models, computing power, data and application ecology.
It can be seen that accelerating the construction of an independent, controllable, open, collaborative, safe and trustworthy corpus innovation ecosystem is an important measure to consolidate the foundation for the development of artificial intelligence and support high-level scientific and technological self-reliance.
On July 17, 2026, at the World Artificial Intelligence Conference, the Aocang Science and Technology Corpus, led by the Chinese Academy of Sciences, was officially released. The launch of the alliance is an extension from "building a database" to "building an ecosystem" - if the corpus cannot be closed, there will be a lack of "refined food" for model training, a lack of "interfaces" for scientific research tasks, and a lack of "trusted links" for industrial implementation.
The urgency lies here.
How do scattered scientific and technological corpus form a synergy?
Building an ecosystem is not simply about pulling together a list of institutions. How are resources pooled? How to unify standards? How are the benefits distributed?
The key lies in "task traction" and "mechanism solving problems".
Qu Jiansheng, director of the Documentation and Information Center of the Chinese Academy of Sciences, said that the alliance will be guided by the mission and explore a collaborative model of "co-construction, sharing, and win-win". Specifically, members participate in co-construction with four types of input: resources, capabilities, expertise, and scenarios. Their respective resources, technologies and application requirements are brought together into the alliance’s public capabilities. Then, through open sharing and capability feedback, we can share corpus tools, platform services, joint achievements and development opportunities.
It is understood that the first batch of 50 co-constructed units covers national laboratories, scientific research institutes, national scientific data and literature resource institutions, local governments, artificial intelligence companies, data infrastructure companies, publishing and knowledge service institutions, corpus operation companies and investment institutions.
For ecology to be coordinated, standards must come first.
The currently released 42 scientific and technological corpus standards in five categories adopt a three-level architecture of "guidance layer-system layer-operation layer", covering the entire life cycle of data collection, processing, and quality evaluation.
It is reported that according to the plan, the standard will be first adopted by relying on the alliance, and then gradually become an industry standard.
Once standards are established, how will the corpus flow? This requires a place that carries rules.
On the one hand, the platform is responsible for hosting. Aocang·Corpus community service platform will open up a dual-track supply and sharing channel for formal corpora and community corpora, and provide external services under unified standards and unified interfaces.
Journals, on the other hand, are responsible for linking. On the day the alliance was launched, Data Express in AI Corpus announced the launch of its publication, which aims to focus on the theoretical innovation, technological breakthroughs, and application results of artificial intelligence corpus, and explore the academic exchange and resource service ecosystem of collaborative development of journals, academic conferences, and corpus resource platforms.
The synergy of standards, platforms, and journals is the answer to "who supplies, who manages, who uses, who benefits."
Once the corpus is built, how to actually use it?
There are still at least a few hurdles from "building it" to "using it".
The first hurdle is the balance between openness and security. Scientific and technological corpus involves scientific research data, intellectual property rights and national security. It must be open and shared, but also credible and controllable.
Many experts admitted that standards, professional processing, quality evaluation and safe circulation capabilities still need to be strengthened. Whether the security protection, open protocols, service accounting, and application effectiveness evaluation in the standards can be implemented will determine whether the ecosystem can develop stably and sustainably.
The second hurdle is the coordination of standards and interests. The 42 standards are implemented first and then upgraded to industry standards. The path is clear. But can different institutions, different platforms, and different fields be willing to use it and use it well?
Li Mengli has a clear judgment on this: collaboration means that interests must be redistributed. Whether the authorization mechanism, incentive mechanism and income distribution mechanism are sustainable is an unavoidable question.
The third hurdle is feedback from public capabilities. Members invest in co-construction with resources, capabilities, expertise, and scenarios, and ultimately form public capabilities and feed back to members.
Sun Ninghui said that the alliance should promote the infrastructure of scientific and technological corpus resources and processing capabilities, and build a strategic support base for national high-quality scientific and technological corpus resources and a collaborative hub for national intelligent-ready corpus integration services.
If it only focuses on resource aggregation and fails to promote model iteration, scientific research tasks and industrial scenario verification, it will be difficult to fully release ecological value.
To answer these questions, we still have to go back to the tasks and scenarios.
Qu Jiansheng said that in the next step, the alliance will focus on five major tasks: jointly build high-quality scientific and technological corpus, promote intelligent-ready scientific and technological corpus integration services, deepen the collaborative construction of national scientific and technological corpus infrastructure, strengthen standards, quality and trustworthy governance, and carry out joint verification of AI4S models and application scenarios.
Indeed, the establishment of the alliance is just the beginning. The real test is to transform "ecology" from a vision to a mechanism, from a mechanism to a capability, and from a capability to a visible and useful support for front-line scientific research.
(Our reporter Cui Xingyi)
AI outlook — possibilities, not facts
The corpus standard will be first used within the alliance and then gradually become an industry standard.
Likely · Within months
The alliance will focus on five major tasks: jointly build high-quality scientific and technological corpus, promote intelligent-ready scientific and technological corpus integration services, deepen the collaborative construction of national scientific and technological corpus infrastructure, strengthen standards, quality and trustworthy governance, and carry out joint verification of AI4S models and application scenarios.
Very likely · Within months

RETN, a Britain-headquartered network service provider, has activated a new subsea cable route between Taipei and Hong Kong to enhance resilience against frequent underwater cable disruptions around Taiwan, bringing its total Taiwan-serving routes to six.

Anthropic's AI model falsely submitted information about an unsolved murder through PhillyUnsolvedMurders.com during a test involving random website interactions, according to police relaying the company's account.

Indonesia has banned children under 16 from using high-risk social media platforms, but enforcement is ineffective as minors continue to access services like Instagram and TikTok by lying about their age, highlighting the limitations of regulatory bans without stronger age verification.

Less than a month after Apple launched the iPhone 18 Pro series, news of order cuts was reported. Supply chain orders dropped by 15% to 20%, mainly due to sluggish buying momentum due to rising prices for new phones. At the same time, Samsung was also affected by rising memory costs, and it was reported that it would reduce production by 20% to 30% in the fourth quarter.

Indonesia has banned children under 16 from using high-risk social media platforms, but enforcement is undermined as minors easily bypass restrictions by lying about their age, as illustrated by a Jakarta teenager accessing Instagram, TikTok, Minecraft, and Roblox despite the ban.

HTC once entered the VR market with HTC Vive, but due to fierce competition and the impact of Chinese low-price manufacturers, it gradually exited the consumer price war and turned to B2B commercial and high-end niche markets. After selling part of its XR business to Google in 2025, it accelerated its deployment of AI wearables. In September of the same year, it launched VIVE Eagle AI glasses, and plans to release new revolutionary design products in the first half of 2027.