Tokyo publishers and rights holders have sounded new warnings about how large language model developers obtain training material, after distributor and industry sources told reporters this week that an international AI company placed unusually large orders for secondhand books in Japan. The buying pattern, which appears aimed at assembling raw text for machine reading and digitization, has reopened debate over whether current industry practices and government guidance provide adequate protection for creators and publishers.
What happened this week
On October 9, reporting based on interviews with book distributors, digitisation services and industry contacts described concentrated purchases of used titles bound for addresses associated with commercial digitisation. The orders covered a broad range of fiction and non fiction, and involved quantities that distributors said were far larger than ordinary secondhand trade. Industry participants raised concerns that bulk purchases are being used to assemble datasets without seeking licences from rightsholders.
Publishers said the pattern echoes earlier cases in other markets, where companies acquired large volumes of printed material to create text corpora for model training, or sent physical books to third party digitisation houses. Rights groups and industry associations noted that even when the underlying law allows certain text and data mining activity, those legal permissions do not resolve tensions around creator remuneration, attribution or commercial reuse.
Tokyo's principle code and the disclosure debate
The new friction lands against a backdrop of recent Japanese government guidance setting expectations for generative AI developers. In August the national Intellectual Property Strategy body finalised a non binding “principle code” that urges AI system developers and service providers to disclose the models they use and the training data or collection methods they relied on, under a comply or explain regime. The code is aimed at improving transparency and reducing intellectual property conflicts while avoiding heavy handed legal obligations that could chill innovation.
That principle code does not mandate disclosure in every circumstance, and it makes clear that sensitive trade secrets or safety critical details can be withheld. But publishers say voluntary guidance is not a sufficient remedy for sudden, large scale acquisition of copyrighted works, especially when disclosure offers no immediate mechanism for licensing, auditing or restitution.
Why the book purchases matter
Bulk acquisition of physical books has technical and legal significance. Physically buying books and sending them to a digitiser is one way to create high quality, machine readable corpora that capture the original text formatting and character encodings without relying on web scraping. For languages with abundant unique scripts and typography like Japanese, clean digitisation can materially improve model performance on morphology, punctuation, and idiomatic usages.
From a legal and commercial perspective, buying secondhand copies does not by itself resolve copyright concerns. Many copyright systems, including Japan's, contain exceptions or allowances for text and data mining intended for research, but those exceptions often do not cover subsequent commercial distribution of outputs or the large scale reuse of expressive content. Publishers worry that models trained on their texts may reproduce or generate content that undermines sales and licensing markets for authors, manga artists and academic publishers.
Publishers demand clearer rules and stronger enforcement
Industry groups representing publishers have called on both government and private sector actors to close gaps in practice and policy. Their requests fall into three broad areas: faster, clearer public disclosure about which works are included in training corpora when a rights holder asks; a practicable route to negotiate licences for commercial model training; and better auditability for datasets used by AI systems that are offered to Japanese users.
Those demands reflect a wider international debate. In Europe and North America rightsholders have already pressed for copyright clarity through litigation, regulatory proposals and voluntary licensing initiatives. Japan's approach so far favors a light touch, using guidelines to nudge behaviour while continuing to rely on existing copyright law and targeted administrative guidance.
What this means for AI developers and Japanese policy
For companies building language models, the episode underscores that dataset provenance matters beyond pure legality. Reputational risk, contract friction with service providers, and potential downstream takedown or legal claims can all stem from opaque data acquisition practices. Firms that operate at scale and across borders face a patchwork of expectations around disclosure and rights management, and voluntary codes do not yet create a single standard.
For policymakers in Tokyo, the current moment poses a choice between stronger mandatory requirements that could impose compliance costs on developers, or a more collaborative path that would pair the existing principle code with industry supported licensing mechanisms and clearer takedown or notice processes. The government has said it prefers to balance innovation with protection for creators, but rights holders argue that voluntary guidance must be backed by procedures that produce concrete outcomes.
Where the story goes from here
Publishers and trade associations in Japan say they will press for rapid engagement with both AI companies and the ministries that authored the principle code. Legal challenges from rights holders remain an option, and the episode could trigger calls for more prescriptive regulation if voluntary transparency does not produce measurable change. Meanwhile, AI developers may respond by updating procurement chains, adding provenance labels, or seeking commercial licences where necessary to reduce conflict.
Whatever follows, the dispute highlights a persistent policy friction of the generative AI era: how to preserve creative markets and author rights while allowing researchers and companies to assemble the large, diverse datasets that drive current advances. Japan's principle code and this week's reporting will likely be referenced in future negotiations about licensing, transparency and the practical limits of a voluntary regime.




