주메뉴바로가기본문바로가기
비즈한국 비즈한국

Exclusive
National Institute of Korean Language ‘Corpus’ Project Copyright Lawsuit, Woongjin Loses Second Trial

This article was automatically translated by AI. There may be errors compared to the original Korean article.  Read original in Korean →

[비즈한국] A landmark copyright infringement ruling has been handed down regarding the process of building AI training data by a public institution. Copyright holders, whose works were used without authorization for the National Institute of Korean Language’s “Everybody’s Corpus” project, have won both the first and second trials in a damages lawsuit filed against the service provider, Woongjin, and its subsidiary, Woongjin Bookcen.

The court determined that having the right to publish and distribute e-books does not grant the right to copy and process those works into AI training data. This ruling is being evaluated as a reaffirmation that even in data-building projects for public purposes, works cannot be used without the copyright holder's permission, and that such actions cannot be justified by “fair use” alone.

Copyright holders whose works were used without authorization for the National Institute of Korean Language’s “Corpus” project won both the first and second trials in a damages lawsuit against Woongjin and Woongjin Bookcen. Photo courtesy of the National Institute of Korean Language.

Liability for Damages Acknowledged After Copyright Controversy

On the 24th of last month, the 8-3 Civil Division of the Seoul Central District Court ruled partially in favor of the plaintiffs in the appellate trial of a damages lawsuit filed by 78 copyright holders of books published by "Communication Books" against Woongjin and Woongjin Bookcen. The court ordered compensation of approximately 3.44 million won out of the approximately 5.5 million won total claimed by the plaintiffs. Woongjin had appealed after the first trial ruled partially in favor of the plaintiffs in February of last year, but the appellate court upheld the original judgment.

This lawsuit stemmed from the "National Institute of Korean Language Corpus incident" that emerged in September 2022. At the time, Woongjin had signed a 3.06 billion won service contract with the National Institute of Korean Language for a "Written Corpus Source Material Collection Project." Woongjin subsequently outsourced the supply and organization of the works to its subsidiary, Woongjin Bookcen, which was responsible for building and delivering the source materials based on content from "Booktopia," an e-book company it had acquired in 2010.

A corpus is a linguistic database constructed to allow computers to analyze books, newspaper articles, and dialogues. The National Institute of Korean Language has been promoting this project for Korean language research and dictionary compilation, and recently, it has also been used as core training data for generative AI and Large Language Models (LLMs).

The court determined that it was difficult to see the rights to use Booktopia content for the National Institute of Korean Language’s corpus project as having been comprehensively succeeded, and that it did not constitute fair use. Pictured is the Woongjin headquarters located in Cheonggyecheon-ro, Jung-gu, Seoul. Photo by Reporter Park Jung-hoon.

The corpus incident surfaced due to collective backlash from the publishing industry. The industry claimed copyright infringement, stating that approximately 16,000 types of e-books were used in the project without the permission of the copyright holders. Although the Korean Publishers Association, the National Institute of Korean Language, and Woongjin Bookcen later reached an agreement, this lawsuit is a separate damages claim filed by individual copyright holders at the time.

“Copyright Principles Apply to AI Data Construction”… Fair Use Claims Rejected

The appellate court did not accept Woongjin’s claim of “comprehensive succession of rights,” stating it was difficult to view Woongjin Bookcen as having comprehensively inherited the rights to copy and publicly transmit the works from Booktopia. The gist is that the contracts typically signed between publishers and authors are limited to the right to publish the original work or provide it as an e-book, and it is difficult to see this as including the right to copy, process, and provide it as a separate product.

Regarding the claim of fair use, the court judged that “the circumstances argued by the defendants alone are insufficient to view the use of the works in question as fair use under copyright law.”

A corpus is a linguistic database used as core training data for generative AI and Large Language Models (LLM). Photo from the National Institute of Korean Language’s Everybody’s Corpus website.

While the court acknowledged that the National Institute of Korean Language’s corpus project was implemented for a public interest purpose, it pointed out that the service contract carried out by Woongjin was for profit. Furthermore, it noted that the contract explicitly required the contractor to “verify rights with copyright holders, enter into usage permission contracts, and spend at least 80% of the service fee on copyright processing,” yet the works were processed without obtaining the plaintiffs' permission.

The amount of damages was calculated by multiplying 70% of the e-book’s list price by five copies and the plaintiff’s stake in each work.

This ruling confirms that general copyright principles apply equally to AI and language data-building projects promoted by public institutions. The court also noted that as the demand for database construction for AI development grows, a market for the licensing of works has already formed. It has made clear that it is difficult to limit the legitimate interests of copyright holders solely for the public interest purpose of AI development.

Following the corpus incident, the Korean Publishers Association, the National Institute of Korean Language, and Woongjin Bookcen signed an agreement in 2023, the main contents of which included shortening the usage period of works, paying additional usage fees, and forming an operating committee. The agreement involves shortening the period of use and revising the system so that subsequent use depends on the publisher’s choice and separate contracts.

As Woongjin has decided not to appeal to the Supreme Court, this lawsuit appears to be concluded with the appellate court’s decision. A Woongjin official stated, “We respect the court’s decision and have no plans to appeal.”

This article was automatically translated by AI. There may be errors compared to the original Korean article.
강은경 기자

기술과 산업을 취재하고 씁니다.

gong@bizhankook.com
저작권자 ⓒ 비즈한국 무단전재 및 재배포 금지