OTWopentechwire
Tech Intelligence, Openly Wired
Policy

Anthropic's Book-Buying Spree in Japan Sparks Copyright Fury

The AI company's bulk purchases of secondhand texts from distributors have ignited fresh debate over training data rights and publisher consent across Asia's third-largest book market.

KW
Kenji Watanabe
Hardware & Products Reporter · Tokyo
Oct 9, 2026
4 min read
Anthropic's Book-Buying Spree in Japan Sparks Copyright Fury
Credit: Reuters

A Quiet Acquisition Campaign

Anthropic has been purchasing physical books in bulk from Japanese distributors, a practice that appears aimed at building training corpora for its Claude language models. The orders, concentrated in secondhand inventory channels, have prompted sharp reactions from publishers who argue the practice circumvents licensing norms even if it skirts the edge of legality.

The purchases came to light after distributor-level sources noted unusual order patterns: large quantities of individual titles, spanning fiction, non-fiction, and technical works, shipped to addresses linked to digitisation service providers. While Anthropic has not commented publicly on the acquisition strategy, the pattern mirrors practices observed in other markets where AI laboratories have sought low-cost access to copyrighted text at scale.

Japanese publishers view the move as a deliberate end-run around collective licensing frameworks. Unlike academic or archival digitisation, which typically operates under negotiated agreements or statutory exceptions, the AI training use case sits in a grey zone where neither permission nor compensation flows back to rightsholders. The anger is compounded by timing: Japan's content industry has been lobbying for clearer protections as generative AI adoption accelerates across the region.

The Legal Grey Zone

Japan's copyright framework permits certain uses of copyrighted material without explicit authorisation, including temporary reproduction for computational analysis. Anthropic's strategy appears designed to test the boundaries of this provision. By acquiring physical copies through standard retail or wholesale channels, the company establishes a chain of lawful possession; digitisation for internal training purposes may then fall within analytical use exceptions, depending on how courts interpret the scope of "non-expressive" data processing.

Publishers, however, reject this framing. They argue that large-scale ingestion and reproduction for commercial model training exceeds the spirit of fair-use carve-outs, which were drafted with academic research and indexing in mind. The fact that Claude generates revenue through enterprise subscriptions and API sales further weakens the analytical-use defence, in their view.

At Opentechwire, we've tracked similar friction points across Seoul, Singapore, and Taipei, where legacy copyright regimes collide with the compute-intensive appetites of frontier labs. Japan's case is notable for the directness of the conflict: rather than scraping web archives or licensing aggregated datasets, Anthropic moved upstream to acquire physical artefacts, a tactic that both demonstrates intent and complicates the legal counter-argument.

Regional Ripple Effects

The backlash in Japan is not isolated. South Korea's publishers' association filed a formal complaint last year against multiple AI companies over dataset provenance, and Taiwan's Ministry of Culture has convened working groups to update text-and-data-mining exceptions. What distinguishes Japan is the scale of its book market and the political clout of its publishing lobby, which has close ties to the ruling coalition and a history of shaping digital-content policy.

Japan's government has signalled a willingness to intervene. Officials in the Agency for Cultural Affairs have indicated they are exploring legislative amendments that would require explicit licensing for training corpora derived from copyrighted works, regardless of the acquisition method. Such a shift would place Japan closer to the European Union's emerging framework, where opt-in consent is becoming the default for commercial AI use.

For Anthropic, the stakes extend beyond Japan. The company has positioned itself as a more responsible alternative to OpenAI and other frontier labs, emphasising constitutional AI and interpretability research. A protracted legal or reputational battle over training data provenance threatens that narrative, particularly in a region where trust in US tech platforms remains fragile after years of data-localisation disputes and content-moderation controversies.

Why Publishers Are Drawing a Line

The anger from Japanese publishers reflects deeper anxieties about value capture in the generative AI era. Book publishing operates on thin margins; advances, editing, and distribution costs are recouped slowly, and digital piracy has already eroded revenue streams. The prospect that AI labs can ingest decades of editorial investment without compensation or attribution is seen as an existential threat.

Several major publishers have begun exploring collective-action strategies, including a proposed licensing consortium modelled on music-industry precedents. The idea is to pool rights and negotiate bulk deals with AI companies, ensuring that training use generates revenue for authors and publishers. Anthropic's unilateral book-buying campaign undercuts this approach by demonstrating that labs can bypass collective frameworks entirely.

There is also a symbolic dimension. Books, unlike web content, are discrete cultural artefacts with clear authorship and established markets. The decision to acquire physical copies rather than negotiate digital rights is perceived as a calculated avoidance of the industry's preferred channels. Publishers argue that if AI companies genuinely respect intellectual property, they should engage with rightsholders directly, not exploit loopholes in secondhand distribution.

What Comes Next

Japan's content industry is now pressing for policy action on multiple fronts. In addition to copyright amendments, publishers are advocating for transparency requirements that would compel AI labs to disclose the composition of training datasets, including the titles and authors of ingested works. Such disclosures would enable rightsholders to assess whether their content has been used and pursue licensing or litigation accordingly.

Anthropic has not indicated whether it will continue bulk purchases in Japan or expand the practice to other markets. The company's silence has been strategic: acknowledging the campaign would invite regulatory scrutiny, while denying it risks credibility if documentation surfaces. For now, the ball is in the hands of Japanese lawmakers, who face pressure to act before the next wave of model releases embeds unlicensed content deeper into commercial products.

The outcome in Japan will likely influence policy trajectories across Asia. If publishers succeed in securing legislative protections or favourable court rulings, other jurisdictions may follow. If Anthropic's approach is validated, expect other labs to adopt similar tactics, accelerating the race to hoover up training data before legal windows close. Either way, the days of frictionless corpus-building are ending, and the fight over who controls and profits from textual knowledge is only beginning.

Read next
Policy

Washington's Voluntary AI Pact Leaves Enforcement Vacuum as Beijing and Brussels Forge Binding Rules

Daniel R. Whitfield · 6 min
Policy

Beijing's AI Pause Calculus: Why Frontier Model Slowdowns Miss the Governance Point

Priya Nair · 7 min
Policy

Three Country-Code Domains Exploited to Issue Fraudulent Google Certificates

Arjun S. Mehta · 5 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.