주메뉴바로가기본문바로가기
비즈한국 비즈한국

U.S. Court Rules 'Physical Book Scanning' for AI Training is 'Legal', South Korean Publishing Industry Reels

This article was automatically translated by AI. There may be errors compared to the original Korean article.  Read original in Korean →

[비즈한국] Anthropic, the developer of the AI chatbot 'Claude', has received a court ruling in the U.S. that its practice of scanning millions of second-hand books for AI training is legal. This verdict sets a precedent that an AI company's method of purchasing physical books and digitizing them can fall under 'Fair Use' in copyright law. Consequently, this is expected to have a significant impact on the domestic publishing market, which has been advocating for copyright protection and fair compensation for data value.

A U.S. court ruled that Anthropic's act of scanning millions of books for AI training constitutes 'transformative use' and is legal. Photo=Generative AI
A U.S. court ruled that Anthropic's act of scanning millions of books for AI training constitutes 'transformative use' and is legal. Photo=Generative AI

Anthropic Uses Scanned Second-Hand Books as Training Data

On January 27 (local time), a Washington Post report revealed the full extent of Anthropic's 'Project Panama'. Since 2023, Anthropic determined that learning from 'books'—high-quality data—was essential to improving the performance of its generative AI models. However, viewing it as practically impossible to obtain consent from individual publishers and authors, the company devised a workaround strategy of purchasing large quantities of second-hand books and digitizing them.

They secured millions of books through second-hand bookstores and libraries struggling with deficits. They then hired external firms to dismantle the books with hydraulic cutters, digitize them using high-speed scanners, and feed them into AI training. Once the training was complete, the books were sent to recycling facilities.

In June of last year, the U.S. District Court for the Northern District of California ruled that these actions by Anthropic were legally sound. The court evaluated Anthropic's use of the books for training as an 'inherently innovative' endeavor. It determined that, much like a reader who consumes books to become an author, the AI model was also learning to create something new, thus falling under 'transformative use' protected by copyright law.

The court noted that by purchasing the physical books directly, Anthropic provided direct revenue to the publishing market. It concluded that this minimized the 'negative impact on the market value of the original work,' which is one of the key criteria for determining fair use.

However, the court did not condone all of Anthropic's training processes. It ruled that Anthropic’s unauthorized downloading and use of millions of pieces of content from illegal pirate sites like 'LibGen' and 'Pirate Library Mirror' constituted copyright infringement. Anthropic settled the lawsuit by paying $1.5 billion (approximately 2.1 trillion KRW) to authors and publishers.

No Domestic Precedents Yet… Focus Turns to Lawsuit Between Three Major Terrestrial Broadcasters and Naver035420

Concerns are rising that this verdict makes it legally more difficult for copyright holders to demand that AI companies 'cease training altogether.' The domestic publishing industry appears to be on edge. While some publishers are taking individual measures, such as inserting 'machine learning prohibition' clauses in overseas licensing agreements, questions remain regarding their technical and legal enforceability.

The domestic publishing industry fears that if a structure where creators are not properly compensated becomes entrenched, the knowledge ecosystem will collapse. The photo shows the Seoul Bookstore inside 'My Friend Seoul Gallery' in Jung-gu, Seoul. Photo=Reporter Choi Jun-pil
The domestic publishing industry fears that if a structure where creators are not properly compensated becomes entrenched, the knowledge ecosystem will collapse. The photo shows the Seoul Bookstore inside 'My Friend Seoul Gallery' in Jung-gu, Seoul. Photo=Reporter Choi Jun-pil

The Korean Publishers Association (KPA) sent an official letter to its members last September requesting a joint response. Guarding against individual publishers entering into unfavorable contracts with AI companies, the KPA is pushing for the signing of a Memorandum of Understanding (MOU) that delegates negotiation rights to the association to protect the value of data assets. The KPA stated that it has currently received authorization from a significant number of publishers.

Currently, most Korean AI companies sign contracts with publishers to secure training data. However, there is a significant gap in perception regarding the value of this data during the negotiation process. Consequently, there are concerns that this ruling could lead to the degradation of intellectual property rights held by publishers and an undervaluation of the entire publishing industry's data assets.

In particular, there is deep concern that if a structure where creators are not fairly compensated becomes fixed, it could lead to the 'collapse of the knowledge ecosystem,' where the will to create is stifled and high-quality books are no longer produced. Critics argue that if scanning and learning are justified simply because a physical book was purchased—as in the case of Anthropic—the rights of creators will have no path for protection.

Kim Si-yeol, Executive Director of Copyright Policy at the KPA, expressed concern, stating, "If book data becomes a cheap reservoir for AI that can be used for the price of a single book, it is questionable who would want to publish their intellectual output as a book," and added, "If the knowledge ecosystem collapses, even AI, which relies on high-quality data, could stagnate."

In domestic law, clear legal principles or precedents regarding 'transformative use' for AI training have not yet been established. In this regard, the ongoing generative AI copyright lawsuit between the three major terrestrial broadcasters and Naver is expected to be an important milestone. The broadcasters argue that news content is for portal exposure, not for AI training. Conversely, Naver claims that even if news content was used as AI training data, it falls under 'research and development of new services' according to their terms of service. They also contend that reporting facts is not subject to copyright protection.

Kim Kyung-hwan, Managing Partner at Minhoo Law Firm, said, "Since there is no precedent for transformative use in domestic law, we must watch related lawsuits to see if the same standards (as in the U.S.) will be applied in Korea," adding, "If we only consider the aspect of industrial development, creators' rights could be stifled, so appropriate compensation is necessary."

This article was automatically translated by AI. There may be errors compared to the original Korean article.
김민호 기자

중화학공업·에너지 분야를 담당하고 있습니다. 지속가능한 사회와 삶에 관심이 많습니다.

goldmino@bizhankook.com
저작권자 ⓒ 비즈한국 무단전재 및 재배포 금지