Home ยป Forum ยป Author Hangout

Forum: Author Hangout

Teaching AI not to write slop

Switch Blayde ๐Ÿšซ
Updated:

I saw on the news today that AI companies are acquiring all the novels they can get their hands on that pre-date AI (I think it said written before 2022 but not sure of that date). They are putting the novels through a machine that removes the spine and then feeds the pages into the system to be read AND LEARNED from.

The goal is to use human written novels to teach AI how to do it so that AI won't write AI slop.

Replies:   awnlee jawking  irvmull
awnlee jawking ๐Ÿšซ

@Switch Blayde

I read an article on this, from the point of view of a secondhand book seller. He's doing a roaring trade, but has grave misgivings because the tech companies buying the books are taking them out of circulation, sometimes destroying the only remaining print copy of a book.

AI companies are far short of ethical.

AJ

Replies:   LupusDei
LupusDei ๐Ÿšซ
Updated:

@awnlee jawking

I'm quasi-following ongoing discussions about this on Twitter (X.com) and some of the rumors gathered by that are:

- the worst offender is allegedly Anthropic

- they apparently believe they are REQUIRED to shred the scanned books due to weird interpretation of copyright law: so they "transfer" the data and and no new copy is made, and so they don't break copyright. Allegedly there's been a court ruling to that effect: that Anthropic are permitted to train on copyrighted material if they destroy the physical copy they acquired.

- Since they can't be bothered to check if the book is or isn't under copyright they bulk apply this to everything. Added benefits (to them) is that this is faster and cheaper and they don't have to store or handle the proceessed material afterwards.

- possible meta benefits of this are: denying the material to competitors, and information control in general.

- it's not just novels, it's anything and everything they can get their hands on.

- yes, the supposed AI watershed date is roughly 2022; it's deemed there's no major volume of printed material with AI contamination prior to that; anything older thus fetch a premium for AI training purposes.
(AI contaminated training data supposedly may be poisonous to AI training purposes potentialy leading to phenomenon known as "model collapse" where complexity of the "long tails" (fringe cases, rich language, obscure knowledge) is lost while repetition is entrenched, defeating the purpose of as large as possible training data corpus and dreams of AGI ANSI etc (of whatever interpretation).
Caveats apply: practice know as "distillation" explicitly train a new AI model on the output of a different one, often resulting in similar overall performance with smaller weights volume. While mathematical loss is obvious, the practical is unknown and unnoticeable -- at least one step down. Iterative experiments that created "model collapse" alarms have been of relatively small scale for obvious compute resource reasons. More specialized distiled models are probably fine.)

- Elon Musk claims Grok do care to check for and preserve rare/valuable copies they come across creating a library. That implies they do not subscribe to the weird copyright infringementing emptor Anthropic seems to employ by destroying everything they scan.

- while all AI companies are doing this to some extent destroying physical copies that are not rare *yet* for marginal savings on processing costs and storage/handling expenses, they are apparently not the only one mass book burners. A business practice that hunt for potentially undervalued print copies to resell at high markup may be worse because they too destroy stock they can't move relatively short-term.

- this isn't (hopefully, no guarantee with types like Anthropic) about shredding one-of-a-kind sixteen century manuscripts. The concern is more about obscure than unique, in space between privately published novels, nineteenth or early twentieth century school textbooks and things like, say, 1991 "Wordperfect for Dummies" on the low end. While one may argue that loss of the last copy of the later leads to no "loss of knowledge" long having no practical value on the face, it still is a historic artificial informing a future researcher about topics like period and trade slang or engineering and visual design practices, or other questions we may not necessarily imagine right now -- and note the design specifics are irretrievably lost even if the text is destructively scanned and preserved as part of AI training dataset (further note: those are said to not be human readable files as they are).

If this AI book burning panic leads to renewed efforts to preserve such "worthless" obscure print artifacts that may be an unexpected silver lining. And, frankly, even the more responsible/charitable AI companies themselves should be interested: there's a valid speculation the current approach may be a dead end simply because we may not possess the unique data volume necessary to train that supposedly coveted Artificial Superintelligence, and if companies end up in turf wars destructively scanning and thus privatizing volumes of obscure data it's super unhelpful to the overall effort.

Replies:   awnlee jawking
awnlee jawking ๐Ÿšซ
Updated:

@LupusDei

Where's Captain Scarlet when we need him?

AIs are like the Mysterons - to create they must first destroy!

AJ

irvmull ๐Ÿšซ

@Switch Blayde

100% certainty that Anthropic (or another company they set up) will be selling the contents piecemeal once they have destroyed the remaining physical copies.

Back to Top

 

WARNING! ADULT CONTENT...

Storiesonline is for adult entertainment only. By accessing this site you declare that you are of legal age and that you agree with our Terms of Service and Privacy Policy.


Log In