top of page

Can Your AI Vendor Tell You What It Trained On?

The answer emerged through litigation in one case, which is not the disclosure route any board should be relying on.



Sign up for my Substack daily AI newsletter here.


See my AI Training course portfolio for corporate Business Leaders here.




CONTEXT AND BACKGROUND

Court records in Bartz v Anthropic revealed that the company purchased millions of physical books, removed their bindings, scanned them and discarded the originals to build a searchable digital library for training its models, in what internal documents called Project Panama. The industry term for the process is destructive scanning, which involves cutting the spine from a book so the pages can be fed through high-speed scanners before the remaining copy is discarded. A federal judge later ruled that using legally purchased books to train models constituted fair use, while separate claims concerning downloads from online libraries were resolved through settlement.


Anthropic’s position is that Claude is trained on a mix of publicly available web data, commercially acquired datasets and internally generated data, that it purchases through regular commercial markets, and that none of its acquisition programmes acquires or destroy rare or antiquarian books.


INSIGHT AND ANALYSIS

Note how that became public. Not through a product disclosure, a customer briefing or a procurement questionnaire, but through discovery in a lawsuit brought by authors. Every enterprise customer of that model learned what it had been trained on at the same moment the litigants did, and only because somebody sued. The same reporting shows how little is known about the layer beneath. Court records established what happened to books after acquisition, not how millions of physical volumes are sourced in the first place. A Dutch antiquarian bookseller in Haarlem received an emailed spreadsheet of 3,001 titles, mostly published in 2020 and 2021 by academic publishers including Elsevier, Wiley, Routledge and Oxford University Press, with shipping directed to China. He dismissed it as phishing. The sender’s company did not respond to requests for comment.


IMPLICATIONS

The governance question is narrow and answerable. A South African bank, insurer or listed company will demand chain of custody for physical inputs, conflict minerals declarations from suppliers, and data residency assurances from cloud providers. The same organisation licenses a model whose training corpus it cannot describe, under terms that do not oblige the vendor to describe it either. Where intermediaries have been named, they have contested the characterisation. A book metadata company told Fortune that a bulk sourcing service reported elsewhere was never launched and that the relevant webpages reflected an exploratory concept rather than an active business. A Canadian firm named in European reporting told a Swiss broadcaster that its purchases formed part of its ordinary recycling and trading model. Contested accounts are unremarkable in any supply chain. Having no contractual right to ask is not.


CLOSING TAKEAWAY

Three questions belong in the next renewal. What material was this model trained on, described by category and source rather than by adjective? Under what licence or legal basis was that material obtained, and has any of it been the subject of litigation or settlement? Can the vendor evidence both without a court compelling it? None of these requires disclosure of model weights or commercially sensitive architecture, and all three would be routine in any other supply relationship. A board that cannot answer them cannot describe its own exposure to copyright claims, reputational risk or regulatory change. Litigation should not be the mechanism by which customers learn what they bought.


Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net




 
 
 

Comments


Leveraging AI in Human Resources ​for Organisational Success
CTU Training Solutions webinar

bottom of page