The publisher added a statement prohibiting use for AI training on the copyright page, triggering controversy over the boundary between copyright and AI training.
Quick Look
- Some publishing houses have added a statement prohibiting unauthorized use for AI training on the copyright page of their books, reflecting the publishing industry's proactive demands for copyright protection.
- In judicial practice, courts tend to determine infringement from the output end in AI infringement cases because it is difficult to prove the use of the green end.
- The industry calls for the establishment of authorization mechanisms and collective licensing to balance innovation and rights protection.
AI-generated summary
Why It Matters
With the development of generative artificial intelligence, the demand for corpus for large model training has increased. Some works have been used for AI training without knowledge, authorization, and remuneration, causing copyright disputes. Some publishers express their attitude by adding a statement prohibiting use for AI training on the copyright page.
With the rapid development of generative artificial intelligence, some works are being quietly "fed" to large models without the knowledge, authorization, and remuneration of the copyright holders. How to prove infringement of large model training end? Where are the boundaries of “fair use” of a work? Is the relevant platform a neutral party or a content user?
"It is prohibited to use the contents of this book for artificial intelligence training, and violators will be prosecuted." Recently, some readers discovered that such a sentence quietly appeared on the copyright page of some books from a publishing house. This new statement, which has not appeared in previous editions, expresses the copyright owner's clear attitude towards "unauthorized collection of content in the book for AI training."
Currently, generative artificial intelligence is developing rapidly, and the conflict between the huge demand for corpus for large model training and copyright protection is gradually emerging. A problem facing creators is that their works are being quietly "fed" to large models without their knowledge, authorization, or remuneration.
As the core driving force of a new round of technological revolution and industrial transformation, the development of AI is inseparable from the supply of massive data. But when these data contain works protected by copyright, how should the rights of creators be protected?
When AI training encounters "copyright defense battle"
"Although it is difficult to ban books from being used for AI training with a statement, at least it shows an attitude and respects copyright." Cheng Cheng, who has been working as an editor in a publishing house for many years, told the Workers' Daily reporter that published books have been reviewed and reviewed three times. The content is accurate and the writing is standardized. They are high-quality corpus for large-scale model training. "Sitting back and seizing it" is a disregard for intellectual achievements.
Talking about if the books he edited and published were used for AI training, Cheng Cheng said, "I think most practitioners would be unwilling to do so without knowledge and without compensation. In the future, more and more books may be marked with 'no use for AI training without permission'."
However, opposing the unauthorized use of training is not the same as opposing AI itself. Cheng Cheng mentioned that currently, many publishing houses are using their own published books to train their own models, hoping to form a corpus in a certain professional field.
In May this year, 22 publishing and media organizations including China Encyclopedia Press jointly issued the "Proposal for the Construction of High-Quality Corpus for Artificial Intelligence", advocating the principle of "authorize first, use later" and work together to create an authoritative and genuine corpus that can be verified and commercially available. From a single publisher's statement to an industry collective initiative, it reflects that the industry's protection of copyright is moving from passive defense to proactive rule building.
In judicial practice, courts in many places have concluded a number of cases involving AI copyright infringement, and the tug-of-war between AI industry development and copyright protection has further come into public view.
In a copyright infringement case involving generative artificial intelligence services concluded by the Guangzhou Internet Court, the court held that a website’s AI painting function directly outputs pictures that are substantially similar to Ultraman’s works according to user prompts, and relied on members’ recharge computing power to make profits. As the direct provider of the content generation tool, the defendant violated the plaintiff’s right to copy and adapt the Ultraman works involved in the case. In addition, cases involving AI copyright infringement concluded by courts in Hangzhou, Shanghai and other places are adjudicated from different dimensions based on the circumstances of the case, striving to achieve a balance between stimulating the innovative development of the AI industry and protecting copyright.
Judicial determination faces multiple difficulties
A statement expresses the copyright owner's firm protection of the work, but it is not easy to prove that the work is used for AI training.
In the above-mentioned case concluded by the Guangzhou Internet Court, the plaintiff requested that the defendant delete Ultraman materials from the training data set, which was not supported by the court. The reason for the rejection was the lack of direct evidence to prove that the defendant actually used the plaintiff's works for training. How to prove that the work has been used when the input end is "invisible and intangible"? In Shanghai’s first large-scale AI model copyright infringement case, the Shanghai Intellectual Property Court held that only when the AI output content reproduces the original work, it constitutes an infringement of the right to copy.
"From the perspective of the operability of governance, the output end should give priority to breakthroughs." Liu Xiaochun, director of the Internet Law Research Center of the University of Chinese Academy of Social Sciences, said in an interview with Workers' Daily that compared with the training end, the infringement facts at the output end are more intuitive and the damage is easier to quantify. "Prioritizing the output side of governance is low-cost, less controversial, and has quick results. It can also leave buffer space for the industry to explore compliance models on the training side."
Whether using unauthorized works to train AI falls into "use" within the meaning of copyright law is the key to determining whether the training behavior constitutes infringement. In Hangzhou's first case involving a generative artificial intelligence platform infringing the right to disseminate information on information networks, the court held that at the content output end, the platform did not take necessary measures for user-generated infringing models and pictures, which constituted contributory infringement. However, it also stated that the use of works during the training phase was not for the purpose of reproducing original expressions and did not impair the normal use of the original work, and could be deemed fair use.
In addition to actively collected corpus, in interactive scenarios, once the content input by the user itself involves infringement, how is the responsibility distributed between the platform and the user? In this regard, Wang Hua, director of Beijing Hua Ran Law Firm, believes that platform responsibilities need to be distinguished based on their actual handling of content uploaded by users.
Wang Hua said that if the platform only uses the content uploaded by users for the current answer, without storing or training, the role is close to that of a neutral service provider. The "safe harbor principle" (note: network service providers are only obliged to take measures after knowing the existence of infringement or infringing content, such as deleting, blocking or disconnecting links) has certain application space, but it still needs to bear the duty of care that matches the information management capabilities.
"If the platform uses content uploaded by users to train or improve models, it is no longer a neutral channel, but a 'user' who actively uses the content, and should bear a higher duty of care regarding the legality of the source of the content." Wang Hua said.
Finding a dynamic balance between protection and innovation
How to find a dynamic balance between protecting copyright and supporting innovation is an issue that academia and industry continue to explore.
Zhang Hongbo, executive vice president and director-general of the Chinese Text Copyright Association, said, "We hope that the AI industry can respect content creation based on the basic principles of technology for good and people-oriented. AI training that commercially uses copyrighted works should obtain permission in advance and pay reasonable remuneration in accordance with the law." Regarding the contradiction between the immediate demand for massive data and traditional prior authorization rules, he believes that the legal status and copyright resource advantages of copyright collective management organizations can be fully utilized to establish a "package" authorization and large-scale copyright dispute resolution mechanism.
Zhang Hongbo also suggested that rights holder organizations such as copyright collective management organizations, writers associations, federations of literary and art circles, and translators associations can establish efficient dialogue and communication mechanisms with AI enterprise industry associations. Relevant authorities should strengthen coordination and administrative supervision to standardize the market order for the use of copyrighted works in AI data training.
"Protecting copyright does not mean that all AI training activities are considered infringements." Liu Xiaochun said that if the entire chain of compliance obligations is moved to the training end, it will increase the cost of intellectual property verification for small and medium-sized enterprises and weaken the vitality of innovation. She suggested that it should be made clear in the copyright law or its implementing regulations that it is only used for the AI training process and is not used independently for a specific work, which constitutes non-work use, or that it should be set as an exception in fair use that does not constitute infringement.
The reporter noticed that the rules on whether authorization is required for the use of copyrighted works in AI training corpus still need to be gradually established in practice, but it has long been a consensus that the corpus itself should be legal, and pirated and infringing content must not be used to "feed the model." In May this year, four departments including the National Copyright Administration jointly launched the "Jianwang 2026" special action to combat online infringement and piracy, clearly focusing on copyright rectification in the field of artificial intelligence and promoting the resolution of copyright compliance issues with large model training corpus.
"We are happy to see that AI can carry out legal and compliant development and application on the basis of respecting and protecting copyright, promote healthy, standardized and sustainable development of the industry, and empower the real economy." Zhang Hongbo said.
Our reporter Shi Lina
"Worker Daily" (September 17, 2026, Page 06)
What to Watch
AI outlook — possibilities, not facts
In the future, more and more books will be marked with the label "No use for AI training without permission"
Likely · Within months
Copyright collective management organizations will explore the establishment of a package authorization mechanism to standardize the use of AI training corpus
Possible · Within months
Open Questions
- How to prove direct evidence that the work is used for AI training?
- Does the use of copyrighted works during the training phase constitute fair use?
- How is the platform’s liability defined when users upload infringing content?
- How does collective copyright management achieve blanket authorization of AI training corpus?





