What kind of scientific corpus is needed for large models to "understand" science?
Quick Look
- The Corpus Innovation Ecological Alliance initiated by the Documentation and Information Center of the Chinese Academy of Sciences was launched in Beijing.
- It aims to solve the problem of the dependence of large-scale models on high-quality scientific corpora in scientific research and promote the development of AI for Science through standard specifications, Aocang Corpus community service platform and Data Express in AI Corpus journal.
AI-generated summary
Why It Matters
Artificial intelligence is changing the paradigm of scientific research. Scientific research increasingly relies on large models to participate in scientific reasoning and knowledge discovery. However, to "understand" science, large models require high-quality scientific corpus rather than general Internet data.
Scientific research increasingly relies on the participation of large models. But to ‘understand’ science, large models do not rely on general Internet data——
Scientific research needs AI, what does AI need?
The Corpus Innovation Ecological Alliance (hereinafter referred to as the "Alliance") initiated by the Documentation and Information Center of the Chinese Academy of Sciences was recently launched in Beijing. The corpus standard specifications, Aocang Corpus Community Service Platform and the English journal Data Express in AI Corpus were simultaneously released. On the same day, Aocang Science and Technology Corpus signed cooperation agreements with key corpus application institutions and the National Artificial Intelligence Application Pilot Base.
Why do this?
The answer lies not in the corpus itself, but in scientific research.
Currently, artificial intelligence is changing the scientific research paradigm. From experimental induction, theoretical deduction, computational simulation to data-driven, to the deep integration of AI and scientific research (AI for Science, hereinafter referred to as ‘AI4S’), scientific research increasingly relies on large models to participate in scientific reasoning and knowledge discovery. But to ‘understand’ science, large models rely not on general Internet data, but on high-quality scientific and technological corpus.
And the problem arises from this.
The deeper AI4S is, why is it more difficult to bypass the ‘corpus barrier’
Scientific research needs AI. What does AI need?
Sun Ninghui, founder of the alliance and academician of the Chinese Academy of Engineering, got straight to the point: ‘The essence of intelligence is the refinement of data into steel. The large model is the product of in-depth processing of all Internet data. But the corpus in the scientific and technological field is different. It has scientific semantics, physical constraints, and error information. It cannot be used by simply aggregating it. ’
‘Compared with general-purpose corpus, scientific and technological corpus has more diverse sources and more complex professional semantics, which puts forward higher requirements for scientific accuracy, knowledge density, modal association and traceability. ’ Lu Fangjun, director of the Bureau of Basic Science and Technology Capabilities of the Chinese Academy of Sciences, further explained.
Realistic bottlenecks are also highlighted: high-value scientific and technological resources are still scattered and difficult to directly transform into high-quality scientific and technological corpus.
‘The resource aggregation and authorization mechanism needs to be improved, and standards, professional processing, quality evaluation and safe circulation capabilities still need to be strengthened. ’ Lu Fangjun explained that scientific and technological corpus connects original scientific research resources and artificial intelligence models, and is an important foundation to support model training, scientific reasoning and knowledge discovery.
It can be seen that the construction of high-quality scientific and technological corpus has become an important basic task for developing scientific intelligence and improving professional capabilities in artificial intelligence.
Why is it difficult for a single organization to do this? Lu Fangjun said bluntly: Because there is a lack of close coordination mechanism between the corpus supplier, technology research and development side and application demand side. These problems cannot be solved by one unit.
Looking deeper, the competitive landscape is also changing.
Li Mengli, Secretary of the Party Committee of the Documentation and Information Center of the Chinese Academy of Sciences, judged that the global artificial intelligence competition is deepening from a competition of single technology and model capabilities to a comprehensive competition in the coordinated development of models, computing power, data and application ecology.
It can be seen that accelerating the construction of an independent, controllable, open, collaborative, safe and trustworthy corpus innovation ecosystem is an important measure to consolidate the foundation for the development of artificial intelligence and support high-level scientific and technological self-reliance.
On July 17, 2026, at the World Artificial Intelligence Conference, the Aocang Science and Technology Corpus, led by the Chinese Academy of Sciences, was officially released. The launch of the alliance is an extension from "building a database" to "building an ecosystem" - if the corpus cannot be closed, there will be a lack of "refined food" for model training, a lack of "interfaces" for scientific research tasks, and a lack of "trusted links" for industrial implementation.
The urgency lies here.
How do scattered scientific and technological corpus form a synergy?
Building an ecosystem is not simply about pulling together a list of institutions. How are resources pooled? How to unify standards? How are the benefits distributed?
The key lies in ‘task traction’ and ‘mechanism solving the problem’.
Qu Jiansheng, director of the Documentation and Information Center of the Chinese Academy of Sciences, said that the alliance will be guided by the mission and explore a collaborative model of 'co-construction, sharing, and win-win'. Specifically, members participate in co-construction with four types of input: resources, capabilities, expertise, and scenarios. Their respective resources, technologies and application requirements are brought together into the alliance’s public capabilities. Then, through open sharing and capability feedback, we can share corpus tools, platform services, joint achievements and development opportunities.
It is understood that the first batch of 50 co-constructed units covers national laboratories, scientific research institutes, national scientific data and literature resource institutions, local governments, artificial intelligence companies, data infrastructure companies, publishing and knowledge service institutions, corpus operation companies and investment institutions.
For ecology to be coordinated, standards must come first.
The currently released 42 scientific and technological corpus standards in five categories adopt a three-level structure of 'guidance layer - system layer - operation layer', covering the entire life cycle of data collection, processing, and quality evaluation.
It is reported that according to the plan, the standard will be first adopted by relying on the alliance, and then gradually become an industry standard.
Once standards are established, how will the corpus flow? This requires a place that carries rules.
On the one hand, the platform is responsible for hosting. Aocang·Corpus community service platform will open up a dual-track supply and sharing channel for formal corpora and community corpora, and provide external services under unified standards and unified interfaces.
Journals, on the other hand, are responsible for linking. On the day the alliance was launched, Data Express in AI Corpus announced the launch of its publication, which aims to focus on the theoretical innovation, technological breakthroughs, and application results of artificial intelligence corpus, and explore the academic exchange and resource service ecosystem of collaborative development of journals, academic conferences, and corpus resource platforms.
The synergy of standards, platforms, and journals is exactly the answer to ‘who supplies, who manages, who uses, and who benefits’.
Once the corpus is built, how to actually use it?
There are still at least a few hurdles from ‘building’ to ‘using’.
The first hurdle is the balance between openness and security. Scientific and technological corpus involves scientific research data, intellectual property rights and national security. It must be open and shared, but also credible and controllable.
Many experts admitted that standards, professional processing, quality evaluation and safe circulation capabilities still need to be strengthened. Whether the security protection, open protocols, service accounting, and application effectiveness evaluation in the standards can be implemented will determine whether the ecosystem can develop stably and sustainably.
The second hurdle is the coordination of standards and interests. The 42 standards are implemented first and then upgraded to industry standards. The path is clear. But can different institutions, different platforms, and different fields be willing to use it and use it well?
Li Mengli has a clear judgment on this: collaboration means that interests must be redistributed. Whether the authorization mechanism, incentive mechanism and income distribution mechanism are sustainable is an unavoidable question.
The third hurdle is feedback from public capabilities. Members invest in co-construction with resources, capabilities, expertise, and scenarios, and ultimately form public capabilities and feed back to members.
Sun Ninghui said that the alliance should promote the infrastructure of scientific and technological corpus resources and processing capabilities, and build a strategic support base for national high-quality scientific and technological corpus resources and a collaborative hub for national intelligent-ready corpus integration services.
If it only focuses on resource aggregation and fails to promote model iteration, scientific research tasks and industrial scenario verification, it will be difficult to fully release ecological value.
To answer these questions, we still have to go back to the tasks and scenarios.
Qu Jiansheng said that in the next step, the alliance will focus on five major tasks: jointly build high-quality scientific and technological corpus, promote intelligent-ready scientific and technological corpus integration services, deepen the collaborative construction of national scientific and technological corpus infrastructure, strengthen standards, quality and trustworthy governance, and carry out joint verification of AI4S models and application scenarios.
Indeed, the establishment of the alliance is just the beginning. The real test is to transform "ecology" from a vision to a mechanism, from a mechanism to a capability, and from a capability to a visible and useful support for front-line scientific research.
(Our reporter Cui Xingyi)
What to Watch
AI outlook — possibilities, not facts
The Corpus Innovation Ecological Alliance will promote the gradual advancement of scientific corpus standards into industry standards
Likely · Within months
The alliance will carry out joint verification of AI4S models and application scenarios
Likely · Within months
Open Questions
- When will corpus standards be elevated to industry standards?
- How do different institutions coordinate the benefit distribution mechanism?
- How does the Aocang Corpus community service platform achieve a balance between openness and security?






