@awnlee jawkingI'm quasi-following ongoing discussions about this on Twitter (X.com) and some of the rumors gathered by that are:
- the worst offender is allegedly Anthropic
- they apparently believe they are REQUIRED to shred the scanned books due to weird interpretation of copyright law: so they "transfer" the data and and no new copy is made, and so they don't break copyright. Allegedly there's been a court ruling to that effect: that Anthropic are permitted to train on copyrighted material if they destroy the physical copy they acquired.
- Since they can't be bothered to check if the book is or isn't under copyright they bulk apply this to everything. Added benefits (to them) is that this is faster and cheaper and they don't have to store or handle the proceessed material afterwards.
- possible meta benefits of this are: denying the material to competitors, and information control in general.
- it's not just novels, it's anything and everything they can get their hands on.
- yes, the supposed AI watershed date is roughly 2022; it's deemed there's no major volume of printed material with AI contamination prior to that; anything older thus fetch a premium for AI training purposes.
(AI contaminated training data supposedly may be poisonous to AI training purposes potentialy leading to phenomenon known as "model collapse" where complexity of the "long tails" (fringe cases, rich language, obscure knowledge) is lost while repetition is entrenched, defeating the purpose of as large as possible training data corpus and dreams of AGI ANSI etc (of whatever interpretation).
Caveats apply: practice know as "distillation" explicitly train a new AI model on the output of a different one, often resulting in similar overall performance with smaller weights volume. While mathematical loss is obvious, the practical is unknown and unnoticeable -- at least one step down. Iterative experiments that created "model collapse" alarms have been of relatively small scale for obvious compute resource reasons. More specialized distiled models are probably fine.)
- Elon Musk claims Grok do care to check for and preserve rare/valuable copies they come across creating a library. That implies they do not subscribe to the weird copyright infringementing emptor Anthropic seems to employ by destroying everything they scan.
- while all AI companies are doing this to some extent destroying physical copies that are not rare *yet* for marginal savings on processing costs and storage/handling expenses, they are apparently not the only one mass book burners. A business practice that hunt for potentially undervalued print copies to resell at high markup may be worse because they too destroy stock they can't move relatively short-term.
- this isn't (hopefully, no guarantee with types like Anthropic) about shredding one-of-a-kind sixteen century manuscripts. The concern is more about obscure than unique, in space between privately published novels, nineteenth or early twentieth century school textbooks and things like, say, 1991 "Wordperfect for Dummies" on the low end. While one may argue that loss of the last copy of the later leads to no "loss of knowledge" long having no practical value on the face, it still is a historic artificial informing a future researcher about topics like period and trade slang or engineering and visual design practices, or other questions we may not necessarily imagine right now -- and note the design specifics are irretrievably lost even if the text is destructively scanned and preserved as part of AI training dataset (further note: those are said to not be human readable files as they are).
If this AI book burning panic leads to renewed efforts to preserve such "worthless" obscure print artifacts that may be an unexpected silver lining. And, frankly, even the more responsible/charitable AI companies themselves should be interested: there's a valid speculation the current approach may be a dead end simply because we may not possess the unique data volume necessary to train that supposedly coveted Artificial Superintelligence, and if companies end up in turf wars destructively scanning and thus privatizing volumes of obscure data it's super unhelpful to the overall effort.