Home ยป Forum ยป Author Hangout

Forum: Author Hangout

Teaching AI not to write slop

Switch Blayde ๐Ÿšซ
Updated:

I saw on the news today that AI companies are acquiring all the novels they can get their hands on that pre-date AI (I think it said written before 2022 but not sure of that date). They are putting the novels through a machine that removes the spine and then feeds the pages into the system to be read AND LEARNED from.

The goal is to use human written novels to teach AI how to do it so that AI won't write AI slop.

awnlee jawking ๐Ÿšซ

@Switch Blayde

I read an article on this, from the point of view of a secondhand book seller. He's doing a roaring trade, but has grave misgivings because the tech companies buying the books are taking them out of circulation, sometimes destroying the only remaining print copy of a book.

AI companies are far short of ethical.

AJ

Replies:   LupusDei
LupusDei ๐Ÿšซ
Updated:

@awnlee jawking

I'm quasi-following ongoing discussions about this on Twitter (X.com) and some of the rumors gathered by that are:

- the worst offender is allegedly Anthropic

- they apparently believe they are REQUIRED to shred the scanned books due to weird interpretation of copyright law: so they "transfer" the data and and no new copy is made, and so they don't break copyright. Allegedly there's been a court ruling to that effect: that Anthropic are permitted to train on copyrighted material if they destroy the physical copy they acquired.

- Since they can't be bothered to check if the book is or isn't under copyright they bulk apply this to everything. Added benefits (to them) is that this is faster and cheaper and they don't have to store or handle the proceessed material afterwards.

- possible meta benefits of this are: denying the material to competitors, and information control in general.

- it's not just novels, it's anything and everything they can get their hands on.

- yes, the supposed AI watershed date is roughly 2022; it's deemed there's no major volume of printed material with AI contamination prior to that; anything older thus fetch a premium for AI training purposes.
(AI contaminated training data supposedly may be poisonous to AI training purposes potentialy leading to phenomenon known as "model collapse" where complexity of the "long tails" (fringe cases, rich language, obscure knowledge) is lost while repetition is entrenched, defeating the purpose of as large as possible training data corpus and dreams of AGI ANSI etc (of whatever interpretation).
Caveats apply: practice know as "distillation" explicitly train a new AI model on the output of a different one, often resulting in similar overall performance with smaller weights volume. While mathematical loss is obvious, the practical is unknown and unnoticeable -- at least one step down. Iterative experiments that created "model collapse" alarms have been of relatively small scale for obvious compute resource reasons. More specialized distiled models are probably fine.)

- Elon Musk claims Grok do care to check for and preserve rare/valuable copies they come across creating a library. That implies they do not subscribe to the weird copyright infringementing emptor Anthropic seems to employ by destroying everything they scan.

- while all AI companies are doing this to some extent destroying physical copies that are not rare *yet* for marginal savings on processing costs and storage/handling expenses, they are apparently not the only one mass book burners. A business practice that hunt for potentially undervalued print copies to resell at high markup may be worse because they too destroy stock they can't move relatively short-term.

- this isn't (hopefully, no guarantee with types like Anthropic) about shredding one-of-a-kind sixteen century manuscripts. The concern is more about obscure than unique, in space between privately published novels, nineteenth or early twentieth century school textbooks and things like, say, 1991 "Wordperfect for Dummies" on the low end. While one may argue that loss of the last copy of the later leads to no "loss of knowledge" long having no practical value on the face, it still is a historic artificial informing a future researcher about topics like period and trade slang or engineering and visual design practices, or other questions we may not necessarily imagine right now -- and note the design specifics are irretrievably lost even if the text is destructively scanned and preserved as part of AI training dataset (further note: those are said to not be human readable files as they are).

If this AI book burning panic leads to renewed efforts to preserve such "worthless" obscure print artifacts that may be an unexpected silver lining. And, frankly, even the more responsible/charitable AI companies themselves should be interested: there's a valid speculation the current approach may be a dead end simply because we may not possess the unique data volume necessary to train that supposedly coveted Artificial Superintelligence, and if companies end up in turf wars destructively scanning and thus privatizing volumes of obscure data it's super unhelpful to the overall effort.

Replies:   awnlee jawking
awnlee jawking ๐Ÿšซ
Updated:

@LupusDei

Where's Captain Scarlet when we need him?

AIs are like the Mysterons - to create they must first destroy!

AJ

irvmull ๐Ÿšซ

@Switch Blayde

100% certainty that Anthropic (or another company they set up) will be selling the contents piecemeal once they have destroyed the remaining physical copies.

Justin Case ๐Ÿšซ

@Switch Blayde

First thing to do is teach AI what a man and woman is and isn't.
The AI algorithms are the worlds worst to use improper pronouns. (almost as bad as blue haired liberals)

Next thing would be to teach it to avoid RPG fantasy subjects that (should) only appeal to teenagers in mom's basement.
Reality has a place, and that place is true literature. Everything else is slop in my book.

As long as they don't use the "human written" slop that has infested SOL in the past year or so the idea might actually work. That stuff is a waste of bandwidth for sure.
And if it churns out more of what I just waded through in the "New Stories" list... we are all doomed.

Michael Loucks ๐Ÿšซ

@Justin Case

First thing to do is teach AI what a man and woman is and isn't.

One reason I use grok for everything. A simple statement in the agent file:

"Respect period understanding of gender/sex along with language, gender roles, and other social norms for the period covered by the dateline provided"

Solves the problem completely for my proofreading and critique skills. Not sure about crafting prose because I don't allow it to do that (because it totally sucks at it, even with training).

awnlee jawking ๐Ÿšซ
Updated:

@Justin Case

I came across this story synopsis:

When courier pilot Neris Vael and her sentient ship are thrown through an impossible dimensional accident, they crash far beyond every world they know. Daniel Mercer, a widowed farmer, makes a choice before he understands what has landed on his property: he treats the strangers at the edge of his woods as neighbours. What follows is not a story of invasion or conquest. It is a story of repairs. Necessity becomes trust. Trust becomes friendship. Friendship becomes a household.

The synopsis is obviously AI-generated - only an AI would put a 'Not X but Y' construct in a synopsis.

And how do you throw something 'through an accident'?

Despite that, the story piques my interest. I know of at least two stories on SOL, human-written, that have addressed this sort of scenario (Colin Barrett and Vincent Berg), and they were both pretty decent. Graybyrd also has an excellent series involving crash-landed aliens but it has a very different tenor.

A long time ago I even attempted my own variant, but it's totally unpublishable. The aliens gave the protagonist a mind-reading device as a reward for his finding everything they needed to repair their ship, and it proved a double-edged sword.

AJ

irvmull ๐Ÿšซ

@Switch Blayde

AI crap is what you would expect from a couple of generations of people who were told that "everybody's special" and where everybody get a trophy, no matter how badly they perform. People with no self-respect.

A bunch of basement dwelling, role-playing gamers, in all likelihood.

Most of their 'work' is about as creative as feeding a Sinatra recording thru an auto-tuner to make it sound like Britney Spears.

Although, come to think of it, that would at least be funny. Most AI can't even manage that.

Replies:   mrherewriting  julka
mrherewriting ๐Ÿšซ
Updated:

@irvmull

I doubt that they are role-playing gamers.

Everyone I knew who did played D&D came before AI and they weren't concerned with trophies or participating in anything that wasn't creative.

And the only "trophy" they wanted was to see their friends having fun with what they created.

If you're talking video games, there was and is no participation trophies for these guys. You were either good, or you were trash-talked to death until you quit or got good. (Most of us were athletes too.)

No one said "Good job" to people who couldn't pull their weight.

Participation Trophies & AI slop are for people who want to be a part of everything by putting in the least amount of effort possible.

julka ๐Ÿšซ

@irvmull

AI crap is what you would expect from a couple of generations of people who were told that "everybody's special" and where everybody get a trophy, no matter how badly they perform. People with no self-respect.

I love how this is always presented as a dunk on the people who got the trophy and there's absolutely zero thought given to questions like "who gave out participation trophies" because that seems like the party who's really to blame here.

I will anticipate your counter argument of "oh the kids would have thrown a tantrum if they didn't get a trophy" yeah man they're kids, they have tantrums, it's the job of adults to not give in on that and you fucked it up, nice job blaming the kids for your failures.

Replies:   Sarkasmus
Sarkasmus ๐Ÿšซ

@julka

I love how this is always presented as a dunk on the people who got the trophy and there's absolutely zero thought given to questions like "who gave out participation trophies" because that seems like the party who's really to blame here.

Yes! Thank you!

I have the same thought every time I see people complain about how all the young ones are addicted to social media on their phones these days. Well... who fucking WROTE those addiction-focused algorithms facebook and Co introduced and then got perfected by TikTok!? Certainly not the 12-year-olds glued to their phones.

Replies:   Grey Wolf
Grey Wolf ๐Ÿšซ

@Sarkasmus

At least as much to the point: who gave them the phone?

Yes, there are reasonable reasons for kids of some ages to have phones. That doesn't mean 'install whatever you want' or 'use it whenever you want'.

And that's on the parents. Facebook, TikTok, etc aren't handing out phones to kids, nor are them preventing parents from taking said phones away.

zathronas ๐Ÿšซ
Updated:

@Switch Blayde

this is my first story written with A.I and here is what I noticed:

1. The longer the story, the more mistakes it makes. My first few chapters were around 40% rewrite, now its more 70%.
2. Too much repetition.
3. Despite giving it extensive details of characters, it keeps depicting one character for another.

I'll finish this story with A.I but I doubt I'll use it again.

A.I Story are easy to detect, for now. Look for phrase repetition.

jimq2 ๐Ÿšซ

@Switch Blayde

I just read that one (or more) of the AI companies is (are)
programming specifically to eliminate or at least limit em-dashes, as that is one of the most frequent indicators of AI output.

Replies:   Michael Loucks
Michael Loucks ๐Ÿšซ

@jimq2

as that is one of the most frequent indicators of AI output.

And one that is completely unreliable and inaccurate.

Replies:   julka
julka ๐Ÿšซ
Updated:

@Michael Loucks

Did you see the arxiv paper on StoryScope[1], by any chance? I haven't thought about their methodology deeply enough to form a really strong opinion on their conclusions, but at least in terms of identifying hallmarks of AI fiction they have some interesting points. They basically condense a story into a set of narrative "features", where a feature is a closed-form question with discrete answer choices, on the idea that stylistic watermarks are easier to ablate with light editing than narrative watermarks.

They ended up with some interesting results (and again, I haven't reasoned about their methodology enough to have a strong opinion about validity here, so you would be well-served to read the paper yourself and have a think); AI fiction tends to over explain themes, while human fiction tends to leave space to infer; AI fiction has much tighter causal chains and fewer sub-plots, while human fiction tends to explore non-linear narratives and have more messy loose ends, that sort of thing.

It's some interesting stuff and may provide a useful vocabulary to talk about AI writing detections in a way that's more robust than the basic style hallmarks that everybody thinks of.

[1]: https://arxiv.org/abs/2604.03136

Replies:   Michael Loucks
Michael Loucks ๐Ÿšซ

@julka

It's some interesting stuff and may provide a useful vocabulary to talk about AI writing detections in a way that's more robust than the basic style hallmarks that everybody thinks of.

I'm reading it now. It's solid (so far), though I'm skeptical that the features are durable against future training/improvements.

Replies:   julka
julka ๐Ÿšซ

@Michael Loucks

For sure the durability against model improvement is an extremely open question. I'd love to read a reproduction of the paper in six months with a bunch more models to see how it holds up over time.

Mushroom ๐Ÿšซ

@Switch Blayde

The goal is to use human written novels to teach AI how to do it so that AI won't write AI slop.

Not possible, because in the end it will always be simply throwing things together in accordance to it's algorithm.

It will never be "real", it will never "think". Just as it was over five decades ago it is just a simulation that follows programming.

For example, no matter how good "AI Voice" has gotten in the last two decades, it is still incredibly obvious to me and I can detect if a narration is real or synthetic in a minute or less.

AI can not really "learn", it can at best simulate "learning" along the parameters built into the programming.

Grey Wolf ๐Ÿšซ

@Mushroom

At the risk of repeating myself (okay, fine - to repeat myself): Define 'think'. Define 'real'. And define 'AI'.

The first two of those terms aren't easy to define, particularly if one includes the notion (fairly popular in both physics and philosophy at the moment) that our universe is deterministic in one way or another. The third is often conflated with current LLMs.

If our universe is deterministic, human beings are, by definition, simply throwing things together in accordance with the algorithm of the universe, and are a simulation of decision-makers, following programming.

At that point, the argument simply becomes that we have better programming. A difference of quality, perhaps, but not kind.

That doesn't mean 'AIs' (e.g. LLMs built by the current design of LLMs) will ever achieve 'AGI' or 'think.' It seems very unlikely that they will. But that is hardly the only class of AI.

There is no necessary reason to believe that a computational device can't 'think.' At some level of computation, they can or we can't, most likely. There is no model of computation more powerful than a Turing Machine (except for some corner-case models which are not believed to be physically possible to construct). That includes human cognition.

Mind you, there are theories that human cognition exceeds the limitations of Turing Machines in some ways - but those theories are currently hypothetical and untestable.

On your last point, there exist LLM 'AIs' that can learn within a reasonable understanding of 'learning' (storing and retrieving information and acting based on that information). To really dig into it, we would also need to define 'learning' in an unambiguous manner, and that's about as easy to do as defining 'thinking' in a way that distinguishes human thought from Turning Machines. A Turing Machine can learn no better than an LLM can be designed to learn, and (again) there is no testable model of human cognition that exceeds what a Turing Machine is capable of.

The TLDR here is: be careful of broad statements based on undefined terms. They really don't mean very much.

awnlee jawking ๐Ÿšซ

@Mushroom

For example, no matter how good "AI Voice" has gotten in the last two decades, it is still incredibly obvious to me and I can detect if a narration is real or synthetic in a minute or less.

I found Marc Nobbs' test very illuminating. There was one human passage and four versions rewritten by AI. IMO, one of the AI versions was very obvious, but without being told beforehand that one was human and the other four were AI, it would have been nigh on impossible.

In the right circumstances, human writing and AI writing can be virtually indistinguishable.

AJ

Replies:   Pixy I
Pixy I ๐Ÿšซ

@awnlee jawking

In the right circumstances, human writing and AI writing can be virtually indistinguishable.

Just a year ago, I could easily tell an AI image of a human apart from an actual image of a human. And the same with AI songs.

Now? I can't tell the difference for both.

And I'm not alone. The BBC news site did one of their little 'quizzes' a few months ago, where the task was to tag pictures of people AI or real. The majority of participants (myself included) were wrong, often tagging the real pictures as AI ๐Ÿซฃ

And it's very apparent on Youtube, with many posts on purely AI-generated music where the comments section below is full of thirst comments from (unsurprisingly), men saying how much they love the singer's voice and how beautiful the singer is on the (AI) cover art. Completely oblivious to the fact the singer doesn't exist.

As I have mentioned elsewhere, the last bastion of being able to be right, was in video, but the difference in quality of animated music videos on Youtube over the last six months has been a little shocking. There are a few Youtube channels I regularly frequent for my taste in music, where just a few months ago, the videos were bloody obviously AI. The videos being released in the last couple of weeks, I can't tell. Well, I can, just, because A: I know from the channel history all the videos are AI and B: there are no long pan shots (or whatever the technical term is) the cuts are all the same length, and C: the people in them are all different for each (whatever it's called, scene?). However, even the latter, C is starting to become obsolete because now the people in each segment are starting to become the same. So continuity between scenes is improving and no doubt, the length of shots will also subsequently increase.

Jo-AnneWiley ๐Ÿšซ

@Switch Blayde

Boy, I don't get it. I worked hard to learn to write. I went to school to learn to write. I forged a career from writing. Through both fiction and commercial writing, I have supported a comfortable lifestyle. Regardless of a project or an assignment, I write everyday. I love the craft of writing. Why would I deny myself that pleasure? I really don't get it...
Jo

Replies:   Switch Blayde
Switch Blayde ๐Ÿšซ
Updated:

@Jo-AnneWiley

I love the craft of writing. Why would I deny myself that pleasure? I really don't get it...

That has nothing to do with AI. Many years ago when I received feedback from a submission editor that I didn't understand, I began my journey to learn the craft of writing fiction. All those things like head-hopping, show don't tell, passive voice, POV, strong verbs, etc., etc. made me appreciate the craft. I was so excited that I kept bringing up in this forum things I learned. My posts were met with resistance and scorn so I realized most writers here don't have your passion for the craft.

Taking that to the next level, there are people who not only don't want to bother learning the craft, but they don't want to bother writing. They simply use AI to do the writing for them.

But the news report was about the AI companies actually trying to teach their products how to craft fiction better.

Replies:   Michael Loucks  Pixy I
Michael Loucks ๐Ÿšซ

@Switch Blayde

My posts were met with resistance and scorn so I realized most writers here don't have your passion for the craft.

Resistance to the demands of 'experts' (not you, but the submission editors) has zero to do with passion for the craft.

Imagine e e cummings trying to be published today. His submissions would never make it past the gatekeepers.

Replies:   awnlee jawking
awnlee jawking ๐Ÿšซ

@Michael Loucks

passion for the craft

Passion for the craft makes you a writing expert.

Passion for writing makes you an author.

AJ

Pixy I ๐Ÿšซ

@Switch Blayde

My posts were met with resistance and scorn so I realized most writers here don't have your passion for the craft

I have another take on that. Humans fear what they do not understand. I was always of the impression that your intention was good, but your presentation...well...

It reminds me somewhat of how a child trying to convey an idea in their head, gets frustrated and angry, because they lack the ability to convey and disseminate their thoughts/feelings. And when they get angry and frustrated, their ability to converse drops even further, and they end up in a 'doom loop'.

That, in my humble opinion, is what makes amateur writers struggle to comprehend the 'proper' technical ability of 'the craft'. They simply lack the comprehension of the basic tenets, which are the rules for the more complex structures. Or to put it simply. "You need to learn to walk before you can run."

Because they don't understand, they slip back into a more defensive posture, which, in human terms, is fear and anger. This is complicated further by being an instinctual reaction rather than a reasoned reaction/response. So yeah, it can be seen as 'resistance or scorn' but most likely, it's just ignorance manifesting itself in ways the individual can deal with/comprehend.

CreepyUnclePete ๐Ÿšซ

@Switch Blayde

There are many utilities that claim to detect AI writing, but most are absolute garbage.

Shakespear and the Bible are more likely to be detected as AI-produced than results from Chat GPT or Copilot.

Most of the 'AI detectors' can't tell if something was written by Grok or Charles Dickens, but many authorities believe the results.

Back to Top

 

WARNING! ADULT CONTENT...

Storiesonline is for adult entertainment only. By accessing this site you declare that you are of legal age and that you agree with our Terms of Service and Privacy Policy.


Log In