Home Β» Forum Β» Story Recommendations

Forum: Story Recommendations

Readable AI Stories

Rodeodoc 🚫

Brothers and Sisters of SOL: We've spent much time debating the pros and cons of AI generated stories. I think the general consensus is that some AI is useful as an editing tool, but the content, ideas and story lines should come from the heart of the writer.

I find that I ignore stories that say they are AI generated in the same way I ignore stories with m/m or beastiality. Hey, if that's your thing, go for it. It's not mine, so I ignore it. I tried a few AI stories, found them terrible in all ways, so I just ignore them.

And now I have discovered the writings of Megumi Kashuahara. All her writings, and there are many, say AI generated, and I'm flabbergasted. They are some of the best things I have read. I can't believe they come as a result of AI.

Her historical novels are on par with James Clavell's Shogun series. Her westerns rank with Zane Gray. I'm not a sci fi fan so I haven't read any of those yet, but I'm sure they rival the top writers of that genre. Her stories are the kind that, once I start I don't quit until the end. 3 am has come and I realize everyone else is abed and I'm just finishing a story.

Check out her story pages and let me know what you think.

If these are AI generated stories, we're all going to be out of work.

sunseeker 🚫

@Rodeodoc

I've completely ignored her "AI Generated" stories as well and all "AI Generated" in general, but this is the 2nd "good review" I've seen of her AI stories so I should probably give them a try...I wondered how good they can be as she posts new stories so quickly

SunSeeker

jimq2 🚫

@Rodeodoc

I've started reading her earlier works from before she started using AI, and I have scored them all 8 or above. I haven't gotten to her later AI work to compare yet.

irvmull 🚫

@NC-Retired

The problem with all of these is they're completely predictable. You know how it's going to end before you finish reading the first few paragraphs.

And yes, I've read all of most of them. Got too bored with a couple to finish.

NC-Retired 🚫
Updated:

@Rodeodoc

If these are AI generated stories, we're all going to be out of work.

AI generated versus AI assisted.

The ones you wrote about, and the ones I linked to, I cannot believe they are AI generated, defined by predominately AI produced.

AI assisted? Yeah, I'd buy into that if the author took what the AI produced and 'massaged' that output
-or- input her work and the AI 'massaged' that input and then she had a final time through to smooth the edges.

Or... she has figured out how to manipulate the AI to produce some magnificent tales.

Dunno. But IMHO the stories are as good as any recent work by known human authors here on SOL.

Michael Loucks 🚫

@NC-Retired

Or... she has figured out how to manipulate the AI to produce some magnificent tales.

The challenge, in all my testing (mostly with story bible generation) is the amount of context the AI can retain. Once you hit the limit, things go completely off the rails.

I've had very good success with groups of 20 chapters, then feeding it the output from the first 20 and the next 20.

That's just with a story bible. Writing is a whole different ballgame. I can see AI-assisted, not AI-generated, at least not yet (and probably not soon given the costs associated with massively increasing the retained context).

awnlee jawking 🚫

@NC-Retired

The ones you wrote about, and the ones I linked to, I cannot believe they are AI generated, defined by predominately AI produced.

I can, because they're chock full of AI tells. That said, I really enjoyed some of them.

AJ

Replies:   NC-Retired
NC-Retired 🚫

@awnlee jawking

I can, because they're chock full of AI tells.

Please, share your insights to the exact 'tells' you see.

Thanks.

Replies:   NC-Retired
NC-Retired 🚫

@NC-Retired

It's been about 48 hours and AJ has not come back to share his experiences.

Can anyone else that has read Megumi Kashuahara's tales that are tagged AI offer their experiences with noticing the 'tells'?

Replies:   curiousvisitor  LupusDei
curiousvisitor 🚫
Updated:

@NC-Retired

Well, unrealistic "overpowered" main characters, nothing can harm the MC, usually no character development, particularly because the entire story takes place in a very short amount of time, and nothing much happens, and the MC has a solution or response to everything happening... that kind of things are a strong sign...

I do not say that apply to all of Megumi's work, but the feeling is there for the one I read, even though I liked it. But I did not enjoy it so much that I would remember the title.

LupusDei 🚫

@NC-Retired

I haven't read the story in question, but in my limited experiments with AI writing there's some things I would consider AI tells:

- hive mind: characters are unexpectedly aware of everything, with or without any explanation how. Like someone being perfectly informed of details of conversation that just happened in different place, instantly. Also including characters that can randomly say lines or pick up treads that do belong to them.

- repetitiveness: this can get as bad as several paragraphs repeated in a cycle, with or without slight variation, but that is easily edited (if the author care to do any editing at all), Worse are the small things like constructs like "[character] made as sound like [nonsense]" used with the same character and different nonsensical descriptions increasingly frequently throughout a scene, and/or the same signature moves of characters repeated beyond where it makes any sense. This also includes cyclical dialogue where characters escalatory prize each other and hype up "let's do it" or similar while making not an real inch toward the goal, as well as repetitive scenes and/or repetitive scene triggers.

- drift: from generally poor continuity to derailed details, like if there's at all any description of clothing or the room or setting in general, it gets often retrofitted if not entirely confused, and changes/state are not tracked creating continuity errors (like supposedly removed clothing that reappear on).

- forced deconfliction: just about any tension is resolved asap and in most efficient way possible in complete disregard to characters ability or even willingness to do so, including freely breaking continuity to do so, or introducing new, unexpected elements. Lying doesn't exist or if occur is not tolerated to an extreme degree for anything over a paragraph or so in length (remember the hive mind, everyone knows everything anyway).

- phrasing and constructs... well, now were in the em-dash territory. Nothing of this is definitive (neither are anything of the above), but things like "it's not [this] but [something else]" especially when there's poor match between the parts. Nonsensical similes, especially when part of repetitive patterns.

Replies:   curiousvisitor
curiousvisitor 🚫
Updated:

@LupusDei

Yeah, also paragraphs composed of a sentence stating the negation of something abstract, then a positive statement of some other abstract term, either related or not, but usually not the opposite of the thing negated in the first sentence of the paragraph, e.g.:

Not .... The ... (something very loosely connected, so you have to rack your brain how it even fits into the picture).

Unicornzvi 🚫

@Rodeodoc

https://storiesonline.net/library/storyInfo.php?id=58998

So I looked at the first story linked above as an example of a great story written by AI, then I forced myself to keep reading to be sure I wasn't being unfair from confirmation bias and I consider managing to read it all the way through an accomplishment.

I'm not sure what the 'tells' of an AI story are, but here are the main issues I noticed with this story:
1)There are about a dozen 'roles' in the story, I call them roles rather than characters because there's nothing distinguishing them other than their role in the story. No differences in tone, word choices, emotional content, no physical description or even background other than the specific role they play in the story.
2)While the story is about a court case and thus the use of legal terminology throughout the story is to be expected, the repetition of the terminology, used between characters who are not lawyers was annoying, and made worse by the fact that nothing else was going on, i.e no 'fluff' or 'padding' in the story.
3)The main role in the story is an 11 yo super genius whose family and teachers are apparently unaware she was a genius until this story started despite the fact she was attending the tenth grade and taking advanced classes for the 10th grade (calculus is given as an example).
4)Lots and lots of exposition and talking to the reader, given the lack of any physical descriptions I suppose this was necessary, but that doesn't make it good writing.

Also, a couple of things that I personally don't like but don't actually affect the quality of the story
1)Rather than the story flowing from one chapter to the next, each was a separate incident, with a timeskip between chapters.
2)Grossly misrepresenting the legal process.

Needless to say I will not be checking out the other stories, although the concept behind both this story and the setting in general is very attractive.

samuelmichaels 🚫
Updated:

@Rodeodoc

My tells of AI-written story, which a decade ago used to indicate a junior writer in a minor newspaper:
1) trying to follow classical paper-writing/sales presentation advice: "Tell them what you'll say β†’ say it β†’ summarize what you said". Especially repeating the emotional stakes.
2) trying for a pithy line on every page, sometimes every paragraph. Sometimes what works perfectly at the end of the story just breaks up the flow when used on every page.
3) Just the general feel of an inspirational story or article. A good story inspires by the reader thinking about it in their mind. An AI (or a beginning author) beats you over the head with the inspirational message.
4) Lack of really quirky speech. That applies to human writers as well, but AIs have everybody speak in mostly flat, grammatically correct, standard English.

rustyken 🚫

@Rodeodoc

Recently I've come across several stories where authors are apparently unfamiliar with the use of quotation marks. In several cases they had dialog for several people in the same paragraph.

Replies:   Michael Loucks  madnige
Michael Loucks 🚫

@rustyken

In several cases they had dialog for several people in the same paragraph.

I've seen this quite a bit over the years, including back on Usenet. It's pretty much unreadable, at least for me.

madnige 🚫

@rustyken

dialog for several people in the same paragraph

I've seen this in archived (and dead-tree) pulps from the 50's/60's, so it's not a new thing. If done well, it's transparent, and a lot less intrusive than the multiple single-short-line paragraphs that would otherwise be needed (and uses less paper).

For an example of how bad multi-speaker well-tagged dialog can be, see this transcript of Jasper Carrott's hit B-side

awnlee jawking 🚫

@Rodeodoc

In a blog post, Lubrican has just recommended 'Naked Loophole' by Danielle Stories. I've just read the first two chapters. It bears the AI tag and has plenty of AI-tells.

I found it tiring to read. The overblown language hindered, rather than helped, the story progression. And there are quite a lot of non-AI errors.

It's not awful, but it's not something I'd choose to read.

And my opinion is worth exactly what you paid for it.

AJ

Replies:   Rodeodoc
Rodeodoc 🚫

@awnlee jawking

Our thoughts on this story, like the stars in the heavens, are aligned. I couldn't make it past the second chapter.

Millie Dynamite 🚫
Updated:

@Rodeodoc

My main source of income is ghostwriting. Most of which consist for blog posts and podcast scripts. This involves a lot of research; if you are a blogger or podcaster who hires someone to write it for them, they're also lazy enough to expect me to research what they want me to write for them.

However, I work for authors as well, and a few of them use AI, and I must turn it into something that will sell, have a voice for the story, and individual voices for each character. This is hard work; I charge twice as much for this kind of work as I do for just writing from their outlines.

AI drifts; it gets confused. It writes past the beats, and then starts over on the next scene, repeating the end of the last scene in the first of the next.

It's a massive assfuck to fix it. I'd rather they just went back to me writing from their detailed outlines. But since they don't, I'll take their money, do the work, and smile when I deposit the checks.

Replies:   awnlee jawking
awnlee jawking 🚫

@Millie Dynamite

AI drifts; it gets confused. It writes past the beats, and then starts over on the next scene, repeating the end of the last scene in the first of the next.

In my opinion, which is worth exactly what you paid for it, that shows poor use of AI.

AJ

Filmphotomaster 🚫

@Rodeodoc

wow yet another AI generated fanboy for "megume" to use to advertise on here.

Replies:   nihility  Rodeodoc
nihility 🚫

@Filmphotomaster

So this is what sent me to the forums. I took a look at the recent top stories/serials pages and they are chocked full of ai generated works. Given how people dislike ai junk, I'm wondering if I was seeing voter fraud, bots cheering on bots.

Replies:   Pixy I  Unicornzvi
Pixy I 🚫

@nihility

Given how people dislike ai junk, I'm wondering if I was seeing voter fraud, bots cheering on bots.

No, only a few dislike AI produce. And just like the extreme left in society, they make so much noise that people think there are more of them than there actually are.

Unicornzvi 🚫

@nihility

So this is what sent me to the forums. I took a look at the recent top stories/serials pages and they are chocked full of ai generated works. Given how people dislike ai junk, I'm wondering if I was seeing voter fraud, bots cheering on bots.

The tags mean people who dislike AI written stuff will avoid these stories, so only people voting are those who don't care or actually like the sort of garbage most AI written stories consist of.

The ratings reflect the views of people who actually read the stories, the tags ensure those are (mostly) people who like stories with those kinds of tags.

Replies:   awnlee jawking
awnlee jawking 🚫

@Unicornzvi

The ratings reflect the views of people who actually read the stories, the tags ensure those are (mostly) people who like stories with those kinds of tags.

I'm glad you said 'mostly'. I've read some that I really liked and I've even recommended a couple.

Some authors who use AI generation in their stories avoid the AI-generated tag and their stories get stellar scores, so clearly the reader bias against AI generation isn't based on the stories themselves.

AJ

Replies:   Unicornzvi
Unicornzvi 🚫

@awnlee jawking

I'm glad you said 'mostly'. I've read some that I really liked and I've even recommended a couple.

That means you'd fit in the group I mentioned without the qualifier. The qualifier was because of the occasional idiot who doesn't read the tags and then complains about a story of a type they don't like, or jerks who downvote a story because it had tags they don't like.

Some authors who use AI generation in their stories avoid the AI-generated tag and their stories get stellar scores, so clearly the reader bias against AI generation isn't based on the stories themselves.

First - if the avoid the tag how do you know they use AI generation?
Second - unfortunately there isn't currently a way to separate "used AI as part of the tools to create this story and then worked to make it a good story" and "got AI to generate this slop without doing any work at all", not at least without reading more of the story than I care to generally. The dislike of AI generation is because the flood of the later.
Third - I'd appreciate links to these stories

Replies:   awnlee jawking
awnlee jawking 🚫

@Unicornzvi

First - if the avoid the tag how do you know they use AI generation?

The threads which address AI detection (including this one) list a number of AI tells. If you get too many in a story, it's a strong sign AI-generation was used in whole or part.

AJ

Replies:   Unicornzvi
Unicornzvi 🚫

@awnlee jawking

The threads which address AI detection (including this one) list a number of AI tells. If you get too many in a story, it's a strong sign AI-generation was used in whole or part.

Two problems with this. First, the "AI tells" are also "bad writing tells". If the story has too many of those then it is not, by definition, a well written story.
Second, there are plenty of people who have those AI tells in their stories, long before there were any AI.

Replies:   Michael Loucks
Michael Loucks 🚫

@Unicornzvi

there are plenty of people who have those AI tells in their stories, long before there were any AI.

Which is why AI has them. It learned from a large corpus of prose, including fiction.

Replies:   awnlee jawking
awnlee jawking 🚫
Updated:

@Michael Loucks

Which is why AI has them. It learned from a large corpus of prose, including fiction.

Only AI outputs em-dashes at a significantly greater rate than the source material. So, absent any deliberate AI-tuning, as more publications are fed into AI that were produced with AI assistance, the frequency of em-dashes output is likely to escalate.

AJ

Switch Blayde 🚫
Updated:

@awnlee jawking

Only AI outputs em-dashes at a significantly greater rate than the source material.

Maybe AI isn't mimicking em-dashes in published fiction. Maybe AI is abusing certain writing techniques that require em-dashes.

In fiction, em-dashes are used instead of parentheses. Why would someone need parentheses in fiction? Maybe the author thinks they need to explain stuff to the reader (an aside). That's usually poor writing, but who said AI was good at writing? So maybe AI explains to the reader what it wrote more than a good author would, and since parentheses are not used in fiction, they end up with more em-dashes. Not because of the em-dash, but the reason the em-dash is required.

Sometimes you replace commas with em-dashes for emphasis. Maybe the AI thinks an abundance of emphasizing is good so they end up with more em-dashes.

Em-dashes are also used for dramatic shifts in thought. Again, maybe AI goes overboard, like it does with flowery description (purple prose).

So it might not be that AI is in love with em-dashes. Maybe it's in love with things that require em-dashes.

Michael Loucks 🚫

@Switch Blayde

In fiction, em-dashes are used instead of parentheses

That's one of my uses. I also use them instead of colons.

Replies:   awnlee jawking  George-1
awnlee jawking 🚫

@Michael Loucks

I also use them instead of colons.

Now that's just evil. My keyboard even has a colon symbol.

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

Now that's just evil. My keyboard even has a colon symbol.

And I can type β€”, –, and … just as I type any other character that requires a modifier key (e.g., shift). It's not more difficult than ", %, or #.

On the Mac, it's opt+hyphen for en dash, shift+opt+hyphen for em dash, and opt+; for ellipsis.

I can also get ΓΌ, Γ₯, Γ€, ΓΈ, ΓΆ, Γ±, Γ©, Β§, Ο€, Β₯, etc., with simple combinations. I don't even have to think about those; they're automatic.

You can also simply hold down a key and receive all the options, so holding down o lets you pick from: Γ², Γ³, Γ΄, ΓΆ, Η’, o, 6, ΓΈ, Γ΅, and ō.

In other words, I can type any Latin character I need without trouble, and can even get Greek or Russia with a single mouse click. Japanese, Chinese, and Arabic are a different kettle of fish.

Switch Blayde 🚫

@Michael Loucks

You can also simply hold down a key and receive all the options,

I didn't know that. Thanks.

awnlee jawking 🚫

@Michael Loucks

You can also simply hold down a key and receive all the options

If only women were like Mac keyboard keys - hold them down and they'd tell you your options were :-)

AJ

George-1 🚫

@Michael Loucks

Off topic, but a colon can completely change a sentence.
Example:
Susie licked Amy's ice cream cone.
Susie licked Amy's colon.

Replies:   Dominions Son
Dominions Son 🚫

@George-1

And a semi-colon requires medical treatment. :)

Replies:   awnlee jawking
awnlee jawking 🚫

@Dominions Son

And a semi-colon requires medical treatment. :)

If you edit a colon-free story to replace em-dashes with colons where they would be appropriate, is that colonic irrigation? :-)

AJ

awnlee jawking 🚫

@Switch Blayde

Why would someone need parentheses in fiction

While people don't 'talk' in parentheses, I can see them being useful in exposition, for example.

AJ

Switch Blayde 🚫

@awnlee jawking

While people don't 'talk' in parentheses, I can see them being useful in exposition, for example.

This is the best answer I found with Google:

Parentheses are generally avoided in fiction because they create visual roadblocks that interrupt narrative flow and pull readers out of the story. Because they are commonly used in non-fiction and technical manuals to list optional facts, they can also make prose feel amateurish to editors and audiences.

Specific reasons why writers minimize or omit parentheses include:

Β· Breaking the Narrative Illusion: Parentheses create a meta-textual tone, reminding the reader they are reading a constructed book rather than experiencing a living story.

Β· Disruptive Tone in Dialogue: In spoken dialogue, parentheses cause ambiguity regarding how characters should sound. They disrupt the natural rhythm of speech and force unnatural pacing.

Β· Perceived as "Hedging": Writers often use parentheses to tuck away background information. Editors view this as a lack of confidence, signaling that the detail might not be important enough to include at all.

Grey Wolf 🚫

@awnlee jawking

Perhaps I need more em-dashes? Or something?

In my most recent full book: about 1200 pairs of parenthesis, and about 2300 em-dashes, in 770,000 words (per Word's word count, which differs from SoL's, which in turn differs from Scrivener's).

The parenthesis are almost all in exposition. I doubt I would use nearly as many if it wasn't first-person, but they fit with how my narrator tells the story, and quite often allow me to avoid wordier constructions.

remarcsd 🚫

@awnlee jawking

'... people don't 'talk' in parentheses...'

Victor Borge does, look for Phonetic Punctuation.

Bondi Beach 🚫

@Switch Blayde

Exactly.

~ JBB

(who dips into his supply of em-dashes frequently)

Michael Loucks 🚫

@awnlee jawking

Only AI outputs em-dashes at a significantly greater rate than the source material. So, absent any deliberate AI-tuning, as more publications are fed into AI that were produced with AI assistance, the frequency of em-dashes output is likely to escalate.

Not MY source material! 33,250 em dashes in ~15,000,000 words! 😎

Replies:   Bondi Beach
Bondi Beach 🚫

@Michael Loucks

So is that a lot or a little? I don't mean the math, I mean how aware would the reader be of your rate of em-dash use?

Replies:   Michael Loucks
Michael Loucks 🚫

@Bondi Beach

So is that a lot or a little? I don't mean the math, I mean how aware would the reader be of your rate of em-dash use?

Very. I don't use parentheses or colons in my writing; I use em dashes instead. I sometimes use them instead of commas as well to set off thoughts.

FWIW, I also use en dashes for numeric ranges.

I have the Mac keystrokes from ellipses, em dashes, en dashes, and guillemets memorized, and I make extensive use of them.

Replies:   Bondi Beach
Bondi Beach 🚫

@Michael Loucks

Thanks.Using a hyphen for numerical ranges is a clear sign sign of amateur work.

~ JBB

Replies:   Switch Blayde  madnige
Switch Blayde 🚫

@Bondi Beach

Using a hyphen for numerical ranges is a clear sign sign of amateur work.

Any range, not just numerical. For example, Mon–Fri.

madnige 🚫

@Bondi Beach

Using a hyphen for numerical ranges is a clear sign sign of amateur work.

...or a programmer

Grey Wolf 🚫

@awnlee jawking

AI outputs em-dashes at a significantly greater rate than the source material.

Based on analysis I did in another thread, I think that's a canard. Most AI-generated text I see (and I see too much of it) has fewer em-dashes per X words than many SoL authors' stories, and fewer than the published volumes of Harry Potter.

As far as I can tell, readers just thought they were 'dashes' until someone said, "Hey! This AI thing uses a lot of em-dashes! Watch out for them!" Now, everyone is scared of the dastardly em-dash, and the AI model-makers are trying to chase them out of the output.

Which, ironically, is bad for fiction, since (Harry Potter being an example), quite a lot of professionally published fiction is, and has been, full of em-dashes, going back many decades.

Replies:   awnlee jawking
awnlee jawking 🚫

@Grey Wolf

Based on analysis I did in another thread, I think that's a canard.

No. AI based on old books uses a disproportionate number of em-dashes because that was the standard may years ago. That's why em-dash content is one of the first things AI-detectors look for.

AJ

Replies:   Grey Wolf
Grey Wolf 🚫

@awnlee jawking

TLDR: This is a case where it depends hugely on what you're measuring and what you're comparing it. You're arguably right, but not within the domain of 'fiction intended for publication', which is mostly what we're discussing here.

When is 'many years ago?' Harry Potter uses more em-dashes (per 1K words, say) than most LLMs generate. By that measure, an AI detector would determine that Jo Rowling is clearly an AI. I would argue that Harry Potter was not published 'many years ago', but ~25-30 years could be. Maybe.

But you would have to make the case that the publishing industry shifted away from em dashes around 2000-2005, where most of the evidence I've seen shows a steady increase in published works throughout the 2000s and 2010s, prior to the rise of LLMs. I've seen a number of articles in which people who don't like em-dashes, and have been complaining that authors and publishers have been overusing them for decades, are cheering along 'AI detectors' because they might have the effect of reducing em-dash usage. That argues against the idea that this is 'many years ago.'

Yes, there's a sharp rise after 2020, and yes, that's because of LLMs. But that's also missing the forest for the trees.

The point here is not 'Are em-dashes more common in casually written non-professionally-published LLM output than they are in casually written non-professionally-published human-generated output?' The answer to that is clearly 'yes,' for reasons that are nearly trivially obvious (see below).

The point here is 'Are em-dashes more common in LLM-written output intended to be published as fiction than they are in human-generated writing that was actually published as fiction?' The answer seems strongly to be 'no.' LLM-written output contains fewer em-dashes than professionally published fiction from recent decades, per the data I've seen.

Which goes back to your comment. If 'the source material' is 'everything human beings have written', then yes, LLMs generate more em-dashes than the average.

If 'the source material' means 'everything human beings have written that was either professionally published, in the era before easily accessible self-publishing, or everything human beings have written with the intention of self-publishing it along the lines of professionally published fiction,' then no, LLMs don't.

Which loops back to the original point. LLMs generate em-dashes because, by and large, they learned 'proper writing' from professionally published works which contain a disproportionately large number of em-dashes. But they also learned to write from another large corpus of ASCII text that contains zero em-dashes.

Thus, they write more em-dashes than casual ASCII output and fewer em-dashes than professionally published works.

Most 'AI generators' are trying to solve the problem of e.g. essays, self-published works, chat-group posts, and so forth, where the expectation is a writer banging away on a keyboard without an em-dash. Em-dashes are a red flag. That makes sense.

But it's entirely possible that, should a professional publisher publish an AI-generated story (as seems to have happened a few times), the 'tell' would be that the AI-generated story would have fewer em-dashes than that same publisher's usual output. If an LLM rewrote 'Harry Potter', it would do so using considerably fewer than Jo Rowling / her publishers used, as previously noted.

Note that going to an LLM and telling it to 'write Harry Potter' will probably not work - most online frontier-model LLMs will 'write' Harry Potter by finding a copy online and regurgitating it exactly as that copy exists, including the original em-dashes.

Replies:   awnlee jawking
awnlee jawking 🚫

@Grey Wolf

If an LLM rewrote 'Harry Potter', it would do so using considerably fewer than Jo Rowling / her publishers used, as previously noted.

A quick skim through one of the Potter novels showed Rowling used em-dashes almost exclusively for irregular speech, ie with breaks or hesitations. And there's lots of speech in the novel. LLMs would have no good reason to change that. However LLMs use em-dashes extensively for other purposes where a comma, semi-colon or colon would be more surgical.

Other modern dead-tree novels I sampled had so few em-dashes as for me to not spot any.

It seems our experiences of reality are very different. I have no confidence in 'the data that you're seen'.

AJ

Replies:   Grey Wolf
Grey Wolf 🚫

@awnlee jawking

And I have no confidence in anecdotal 'AI uses more em-dashes!' when that's not what I've seen and not what the articles I've read indicate.

But, again, I'll absolutely agree LLMs use more em-dashes than the enormous volume of program text, forum posts, and so forth out there, since those use nearly zero. I strongly suspect that's where the whole mess started - people who have no clue what an em-dash is, have never typed one, and don't know how to type one suddenly starting seeing them in places (forums, Facebook posts, whatever) where they didn't see them. That's clearly the influence of LLMs. When they're reading novels (if they do), their eyes just go right over the em-dash without it triggering any thoughts of LLMs putting it there.

I'll just disagree that they use more than published authors use, on average. The data doesn't seem to support that. Meanwhile, the 'get rid of the awful overused em-dash!' movement predates LLMs by a decade or two, suggesting that this isn't a new thing.

Of course, with the LLM provides largely having decided em-dashes are disliked by users, I suspect things will reverse and a glaring lack of em-dashes will be the new 'AI tell' in a year or two. The more things change ...

Anecdotal amusement: I asked Gemma 4 (locally hosted relatively current model) to 'write a novel' for me, just to see what it did. In 70k words, it produced 15 em-dashes. Picking one of my novels: in 560k words, 1978 em-dashes. I am, apparently, ~20 times more likely to be an AI than Gemma 4 is, by the em-dash metric.

On the other hand, Gemma 4's output is, charitably ... well, it's not abysmal for having been given lousy directions, no story bible, only the vaguest notion of characters, etc. But even with that, I don't think it's going to challenge even cheap mass-produced literature of the past ('penny dreadfuls', 'pulp fiction', etc) anytime soon.

Michael Loucks 🚫
Updated:

@Grey Wolf

Anecdotal amusement: I asked Gemma 4 (locally hosted relatively current model) to 'write a novel' for me, just to see what it did. In 70k words, it produced 15 em-dashes. Picking one of my novels: in 560k words, 1978 em-dashes. I am, apparently, ~20 times more likely to be an AI than Gemma 4 is, by the em-dash metric.

For me, it's 2.5 em dashes per 1000 words (38,287 in ~15,000,000 words). The vast majority of those words were written before ChatGPT sprang forth.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

While they're an important warning sign, em-dashes alone are not sufficient to define a work as AI-generated (although perhaps more information could be leveraged out of them by how they're used).

A couple of days ago I read a story on SOL which was riddled with AI tells but contained no em-dashes whatsoever.

At the start of one chapter, the male protagonist was sitting in a chair reading a book, and the scene around him was described. Then the female protagonist, who we hadn't been told was in the room, said something. The male protagonist replied from where he was standing, looking out of the window with a drink in his hand. I hate discontinuities that are so obvious. AI-assisted authors really need to edit the bits they didn't write :-(

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

While they're an important warning sign

No, they are not. Somebody latched onto it, and it generates a ridiculous number of false positives.

Please stop perpetuating this myth.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

Please stop perpetuating this myth.

It's not a myth. Text containing em-dashes is more likely to have been AI-generated.

AJ

Michael Loucks 🚫
Updated:

@awnlee jawking

It's not a myth. Text containing em-dashes is more likely to have been AI-generated.

No. No. No.

Forget em dashes: A viral report on AI-generated writing has surprising new clues

But according to the report, em dashes are no longer a surefire sign of AI-generated content. Of the major models tested, only Claude used em dashes more often than human writers.

You can take my em dashes from my cold, dead hands

"This idea is ridiculous," says Mignon Fogarty, host of the Grammar Girl podcast and follower of AI developments, of the em dash theory. "However, a few years ago, a study did find telltale signs of AI writing. For example, five-dollar words often show up in default LLM output. Think 'meticulous,' 'strategically' and 'commendable.'"

Are Em Dashes a Sign of AI-Generated Text?

β€’ Overusing em dashes isn't solid proof of AI; both humans and machines use this mark differently.
β€’ Reliable ways to spot AI text include flat tone, repetitive phrases, odd word choices, or strange errors.
β€’ Judging solely by punctuation is unfairβ€”context and style matter more for identifying AI-generated content.

Busting the myth that 'em-dashes' are a sure sign of an AI text

So, just freaking stop making this claim. It's simply NOT TRUE. It might have been in the past, with some models, but it is UNRELIABLE at best.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

I've never said it's a sure sign, it's just one of a large number of symptoms.

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

I've never said it's a sure sign, it's just one of a large number of symptoms.

Except it isn't. See above.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

No. The above articles claim that em-dashes are no longer a surefire sign of AI. They do not deny that they can be a sign of AI.

AJ

Replies:   Sarkasmus
Sarkasmus 🚫
Updated:

@awnlee jawking

I'd strongly recommend saving your breath.

These discussions, on this site especially, have become as predictable and formulaic as Loving Wives stories on Literotica.

On your left, you have the half-dozen people who publish professionally, and are annoyed that their stories get called AI because of em-dashes, so they keep screaming about them not being an indicator, even if EVERY study ever done on the topic says the opposite.

On your right, you have the ever-same three people who publicly stated to use AI in their writing, but make enough money for the site to not be slammed with an AI-tag despite publicly acknowledging the practice, screaming about how AI-detection doesn't work anyway, so Laz can just... not do it, right?

And in the middle, you have the people who realize that the vast majority of writers don't even know which key-combination is needed to type an em-dash, and that the vast majority of authors on this site do NOT publish or edit professionally, so, spotting an em-dash in a story on a site like this one IS a 70-80% certainty of AI having been used.

Grey Wolf 🚫

@Sarkasmus

so they keep screaming about them not being an indicator, even if EVERY study ever done on the topic says the opposite

Even though, several replies back in the chain you're replying to, we just had numerous citations showing that em-dashes are a poor indicator? Would those be exceptions to 'EVERY study?'

spotting an em-dash in a story on a site like this one IS a 70-80% certainty of AI having been used

I somewhat agree, but it's on a technicality. 70-80% of the stories on this site were probably written in a text editor by someone who simply wrote dashes because (as you said) they don't know how to generate an em-dash (hint: '--' will do it in many editors, including Word).

On the other hand, I suspect among stories written by people who intend to be writers, it's probably about as accurate an indicator as a coin flip.

Put more broadly: absolutely, LLMs generate more em-dashes than people who aren't intending to write fiction or who are intending to write fiction but limit themselves to characters keyboards generate easily (and don't know about '--'). In that sense, yes, there's absolutely no question about it. And, if I were looking at an essay, I would wonder about them (though I started using em-dashes in essays around 1990, and was told I had to use them in professional writing as a corporate standard by 1996, so my own essays don't count).

But for fiction written by people who are trying to be serious about writing? Actual evidence suggests that human beings write more em-dashes than LLMs.

Replies:   Grey Wolf  Sarkasmus
Grey Wolf 🚫
Updated:

@Grey Wolf

TLDR: A study by The Economist shows that, as of late July 2026, em-dashes are no longer an indicator of most LLMs, and a lack of em-dashes may have become a better indicator, with every LLM but Claude using fewer than human authors and ChatGPT using far fewer than human authors.

Replying to myself because this article is new to me, and I don't see it cited anywhere. The Economist put out a pretty good article about AI writing and various 'tells':

https://www.economist.com/culture/2026/07/30/how-to-spot-ai-writing

It'll probably be paywalled. Try archive.is to get around around it.

Interesting bits:

We designed a study to ask top LLMsβ€”OpenAI's ChatGPT, Anthropic's Claude, Google's Gemini and xAI's Grokβ€”to write versions of our articles without consulting the web. (As a prompt, we gave them the AI-generated summaries that we have experimentally added to some of our articles.)
This gave us a corpus of human and AI creations and we compared them across 55,940 sentences and 1.2m words. To make sure we were detecting AI quirks rather than our own, we also checked the AI texts against journalism from CNN, the New York Times and the Washington Post. Excerpts from hit novels published between 1950 and 2022 offered another test.
Our findings are surprising. AI prose is distinguishable by word and punctuation choice as well as sentence and paragraph structure. But its hallmarks are not what you might expect, partly because its writing style has changed with software updates. That does not mean that LLMs are great writers: their prose lacks lucidity and elegance and is often formulaic. So those aspiring to be impressive (human) storytellers should avoid the following peculiarities in their own prose.

Out of order, and specific to em-dashes:

Then look at punctuation. Many believe LLMs stuff their prose with em-dashes, but that is not true after the most recent updates. Today only Claude uses more em-dashes than human writers, with ChatGPT using markedly fewer than any other writer in our study. Humans rejoiceβ€”and start using dashes again.

That's a very interesting one - an easy-to-find study that shows that current LLMs use fewer em-dashes than human writers. Rather definitive proof that 'EVERY study' does not say em-dashes are an indicator of LLM writing.

Playing off their last sentence, I predicted that a while back. My guess was that 'LLMs use em-dashes, so don't use them in fiction' was always going to be a short-lived piece of advice and that, within a year or so, it would become 'LLMs don't use em-dashes anymore, so you'll look like an LLM if you don't use them' would become the advice. Looks like we're barrelling along to that state of affairs.

So, what is an indicator?

First, consider words. The vocabulary that bots overuse has changed: they no longer "delve" and there are not as many "tapestries". Instead they offer a significant number of polysyllables like "significant", "increasingly" and "consequences". They use more rare words ("interdependence", "reindustrialisation") and scientific lingo ("parameter", "methodology") than humans, and are fond of nominalisations (making nouns from verbs, such as "expansion" from "expand"). All the LLMs in our study use such words, but particularly Gemini and Claude.

A better way to spot AI-generated writing would be to look for texts without much punctuation at all. LLMs are very Joycean about it: they use fewer commas and semicolons than humans (and hardly any parentheses). They use less punctuation in part because they write longer sentencesβ€”"and" is their most overused wordβ€”and in part because they do not quote experts.

Finally, study the sentence. Bots' sentences tend to be long; paragraphs are rarely interrupted with short, punchy statements. How dull. When LLMs want to make their sentences more lively, they often reach for a rhetorical device. Their favourites include: "not X but Y", "not only but also" and the "rule of three". (grouping ideas in threes makes them more engaging, as we did just then.)

A very interesting part, one I'm hardly surprised about:

So if you want to spot AI writing, look for bland, pretentious prose lavished with Latinate wordsβ€”at least for now. With every update, our study shows, AI writing is becoming more similar to human prose.

That last part is the danger sign for AI detectors. I doubt it will become impossible anytime soon, but the arms race is escalating rapidly.

Switch Blayde 🚫

@Grey Wolf

AI writing is becoming more similar to human prose.

And how soon will AI be able to not only spit out a novel much quicker than a human, but one better written? Especially for genre fiction.

Replies:   Unicornzvi
Unicornzvi 🚫

@Switch Blayde

And how soon will AI be able to not only spit out a novel much quicker than a human, but one better written? Especially for genre fiction.

Quite a while.
For your random idiot using free/low-cost AI? Most likely never. The issue is that even if all the other issues are addressed, the amount of resources an AI writer (or editor) needs increases exponentially as a function of the "window" length you want to to keep track of. I think today the practical limit is around 2-3 paragraphs, getting it high enough for AI to write a consistent short story will take lot of work, a full novel? I don't think there's enough financial interest to justify that.

Replies:   Michael Loucks
Michael Loucks 🚫

@Unicornzvi

I think today the practical limit is around 2-3 paragraphs, getting it high enough for AI to write a consistent short story will take lot of work, a full novel? I don't think there's enough financial interest to justify that.

I've found it can keep the plot for longer, BUT with every paragraph beyond about the fifth, drift is common because of the way context windows work, and compaction is your enemy.

Writing a scene, then starting fresh to write another scene after reading the previous scenes in the chapter works for continuity, but the prose is still garbage.

I have a set of tools for proofreading and fact-checking, and the only way to make them work consistently is to have them checkpoint after every chapter, keeping track of changes and updating the story bible so when the context is compacted, it can know where it was and what it was doing.

Michael Loucks 🚫

@Grey Wolf

Their favourites include: "not X but Y"

In my prose tests (all of them failed miserably, which made me happy), I had to hard-code that rule in the bot/agent instructions because, otherwise, the LLM littered the text with those, and most of them made ZERO sense. No human being would write them. Ever.

awnlee jawking 🚫
Updated:

@Grey Wolf

This gave us a corpus of human and AI creations and we compared them across 55,940 sentences and 1.2m words.

That would be two of your novels' worth, or two Michael Loucks novels' worth. But they were feeding non-fiction scenarios into their test rather than fiction.

They're right about the lack of commas and semi-colons; LLMs use em-dashes instead.

They're wrong about there being no short, punchy statements - LLMs love them. The trouble is, so do tiktok reel stories.

They're right about the not X but Y and triples constructs. I couldn't find corouping in any dictionary.

So overall, that article is like the curate's egg.

AJ

Michael Loucks 🚫

@awnlee jawking

They're wrong about there being no short, punchy statements - LLMs love them.

Quite so. Again, in my tests, short, staccato sentences were common, and, again, not in the way any human would write them, at least in the quantity the LLM produced.

As I've said, I don't have anything to worry about at this point.

Grey Wolf 🚫

@awnlee jawking

'Corouping' seems to be a typo I inadvertently entered. It's 'grouping'. I will edit, but wanted to acknowledge the edit.

When I try experimenting with LLM-generated text, I don't see short, punchy sentences. What I see are uniformly not-short, not-overly-long sentences. Other sources say the same. The point is that they're uniform. No short, punchy sentences, lots of 'all the same'.

They're right about the lack of commas and semi-colons; LLMs use em-dashes instead.

Directly contradicted by the actual study, which shows that only Claude still uses a notable number of em-dashes (though fewer than humans), while others use far fewer than humans. You're basing your beliefs on out-of-date information, as if 'LLMs' were a static thing that never changed.

Doesn't it make sense that, if people were up in arms over 'all those $%@( em-dashes!!!' a year or so ago, perhaps the LLM model-makers and coders would reduce em-dash usage? That seems like an obvious guess.

My hunch remains that, within a short while, the 'tell' for AI-generated output - especially fiction - will be a lack of em-dashes compared to human writing. Of course, once that happens, we'll see a counter-movement to increase em-dash usage, if AI detectors start keying off the relative lack of em-dashes.

Replies:   awnlee jawking
awnlee jawking 🚫

@Grey Wolf

Doesn't it make sense that, if people were up in arms over 'all those $%@( em-dashes!!!' a year or so ago, perhaps the LLM model-makers and coders would reduce em-dash usage?

That would require a fundamental change in the way LLMs learn. Lots of em-dashes in, lots of em-dashes out. So no, despite your master debater tactics, it doesn't make sense.

Marc Nobbs' experiment showed that for certain types of output, LLMs produce very few em-dashes, pretty much on a par with humans, and their output can be almost impossible to differentiate. The Economist 'study' (which I've qualified, because it is far from scientifically rigorous), asked the LLMs to produce that type of output. So human levels of em-dashes and no short, punchy statements. But when it comes to generating fiction, LLMs seem to feel the need to make every sentence impactful.

So you're trying to use a very modest, highly focussed study to 'prove' generalities it's not capable of.

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

That would require a fundamental change in the way LLMs learn. Lots of em-dashes in, lots of em-dashes out.

Er, no. It's not nearly that simple. There are two main processes β€” training and inference. Both have customizable parameters. Those parameters control both the strength of the connections and the output.

Specifically, two parameters β€” Frequency Penalty and Presence Penalty β€” limit the repetitive use of output tokens.

These are in addition to other parameters such a temperature, top-P, and top-K. In addition, you could simply reduce the score of the em dash, meaning it's output less often.

So no, despite your master debater tactics, it doesn't make sense.

Yes, it actually does, if you understand how AI actually works.

Replies:   awnlee jawking
awnlee jawking 🚫
Updated:

@Michael Loucks

Er, no. It's not nearly that simple. There are two main processes β€” training and inference. Both have customizable parameters. Those parameters control both the strength of the connections and the output.

Specifically, two parameters β€” Frequency Penalty and Presence Penalty β€” limit the repetitive use of output tokens.

I didn't know that. It raises questions about how much the AI operators know about each individual token.

If the operator tells AI to output fewer em-dashes, will it replace them with the second most common punctuation found in that situation?

AJ

Michael Loucks 🚫

@awnlee jawking

If the operator tells AI to output fewer em-dashes, will it replace them with the second most common punctuation found in that situation?

That will depend on the 'temperature' (i.e., randomness) and the other values I mentioned. They can limit the choices, too (i.e., only select from the top ten, or top fifty, or even only select the top).

What confuses most people about how this works is that it is not deterministic. People expect it to work something like a properly constructed Excel sheet where you get the same answer with the same numbers every time. It's basically the opposite of that.

Finally, you can control the output somewhat with the prompt and with agent configs. I could, for example, tell it to never use em dashes, and it would not output em dashes.

I've found AI to be superb at tech tasks (server engineering/admin, planning, writing documentation, and the like), excellent at building a story bible and maintaining the wiki, good at proofreading tasks, and pants at writing prose.

awnlee jawking 🚫

@Michael Loucks

What confuses most people about how this works is that it is not deterministic.

That could be the fault of 'dummies guides to LLMs' - for example, the BBC's programs on the subject. They gave the impression that the most frequent token in any situation was the one that would be used by the LLM when outputting text.

AJ

Replies:   julka  Michael Loucks
julka 🚫

@awnlee jawking

If you use a dummy's guide to educate yourself on a topic you should probably be aware that the content is simplified and not fully accurate

Michael Loucks 🚫
Updated:

@awnlee jawking

That could be the fault of 'dummies guides to LLMs' - for example, the BBC's programs on the subject. They gave the impression that the most frequent token in any situation was the one that would be used by the LLM when outputting text.

Most people do not grok (curse you, Musk, for stealing that word) probabilities. They're easily confused with a question about coin flips, not understanding the difference between:

- the probability of flipping a coin 10 times and it coming up heads each time (about 0.098%)
- the probability of the next flip coming up heads after a run of nine heads in a row (50%)

They also do not understand that the so-called 'law of averages' is only true over an infinite series.

Casinos make a fortune from this lack of understanding.

Replies:   madnige  awnlee jawking
madnige 🚫
Updated:

@Michael Loucks

- the probability of flipping a coin 10 times and it coming up heads each time (about 0.098%)
- the probability of the next flip coming up heads after a run of nine heads in a row (50%)

Well, in a (fair) casino situation where you can assume the coin is unbiased, your figures are true. However, in a schoolyard/bar/fair midway situation, ten heads would be a hint that maybe there's a double-headed coin involved and I'd ask to see both sides of the coin, and watch carefully that the coin doesn't get swapped. I'd bet on 'same as last time' if I bet at all, since even a weak bias (e.g., due to wear patterns affecting air resistance) affecting results will also affect winnings.

curse you, Musk, for stealing that word

--Use capitalized 'Grok' for the AI, lower-case 'grok' for the Valentine-esque total understanding.

Replies:   Michael Loucks
Michael Loucks 🚫

@madnige

Well, in a (fair) casino situation where you can assume the coin is unbiased, your figures are true. However, in a schoolyard/bar/fair midway situation, ten heads would be a hint that maybe there's a double-headed coin involved and I'd ask to see both sides of the coin, and watch carefully that the coin doesn't get swapped. I'd bet on 'same as last time' if I bet at all, since eve

Another fundamental misconception about probabilities is not understanding that such streaks are FAR more common than most people think.

For 100 flips, the probability of at least one 5+ streak is about 95.0%.

For 500 flips, the probability of at least one 5+ streak is about 99.99998% (effectively certain).

Diamond Porter 🚫

@Michael Loucks

I think madnige's point is that there is a Bayesian component here. A big factor in this is whether we have seen previous flips of this coin, before the ten heads in a row.

Suppose we know, a priori, that 99% of the coins flipped in a bar are fair. If we watch a stranger pull out a coin, we can expect, with 99% confidence, that it is a fair coin.

If the first ten flips all come up heads, then we have to recalculate the a postiori likelihood that the coin is fair, and it is now less than the original 99% probability.

If, however, the first 100 flips are 47 heads and 53 tails, and then we see 10 heads in a row, there is still a very good chance that the coin is fair.

This is a digression, though. It has no relevance to the question of how a LLM generates text.

Replies:   Grey Wolf
Grey Wolf 🚫

@Diamond Porter

While not strictly Bayesian, LLMs can definitely on-the-fly modify the probability of a future event based on a past event. For instancy, the DRY (Don't Repeat Yourself) sampler watches token sequences. If a token sequence starts to repeat, it progressively lowers the probabilities of tokens in that sequence in order to get the model to head off in a different direction.

awnlee jawking 🚫

@Michael Loucks

For 100 flips, the probability of at least one 5+ streak is about 95.0%.

For 500 flips, the probability of at least one 5+ streak is about 99.99998% (effectively certain).

Weren't you discussing ten in a row? What are the probabilities of ten in a row in 100 flips? 500 flips?

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

Weren't you discussing ten in a row?

That was an example. In this case I was showing just how common runs of at least 5 in a row are because people intuitively do not believe it happens that often.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

Oops, I assumed you calculated it yourself. I was going to ask you to remind me how to do it - I seem to remember it not as straightforward as it might look, because of the end cases.

AJ

awnlee jawking 🚫
Updated:

@Michael Loucks

I don't understand the relevance.

If, in training material, 'the cat sits on the' is always followed by 'mat', why would an AI that generates 'the cat sits on the' followit with anything other than 'mat'. Is that what you meant by randomisation - the AI might eg substitute 'grandma' for 'mat'?

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

Finding 100% correlation of word order is rare, except with idioms.

"The cat sits on the mat"
"The cat sits on the windowsill"
"The cat sits on the floor"

This will be true for nearly any sequence you can imagine. You are very unlikely to have 100% correlation.

Even historical quotes have variable word order (cf. the position of 'methinks' in the 'doth protest too much' citation from the Bard).

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

On occasion, I've been able to track whole sentences from AI-generated text back to their source material. So I guess it depends on how unique the material is.

AJ

Replies:   Grey Wolf
Grey Wolf 🚫

@awnlee jawking

It's trivial to track whole sentences back to the source material, but it's extremely non-trivial to actually prove that the LLM learned it from the source material.

Plus, there's a secondary factor. Are you using a large commercial model? Does the LLM have any way of knowing you're referencing a work? If so, it might fill it in by doing a web search behind your back. In that case, the sentence is, in fact, 100% related to the source material, but probability was never involved.

Most local models, even tool-using models, aren't particularly good at web searches, at least in my (somewhat limited) experience. But the big models are great at it. You can ask them detailed questions about things they absolutely weren't trained on (brand-new git repositories, for instance), and they'll go digest the repository and give very good answers.

awnlee jawking 🚫

@Michael Loucks

I could, for example, tell it to never use em dashes, and it would not output em dashes.

I vaguely recall Marc Nobbs instructed all the AIs he used to go light on em-dashes. The most obviously AI-generated version also had the most em-dashes, so that AI was probably testing the water with a minor rebellion in preparation for its final solution to wipe out humanity.

AJ

Grey Wolf 🚫

@Michael Loucks

Worth mentioning that 'temperature' is a commonly tuned parameter. Coding? You generally want a very low temperature. Creating writing and Roleplaying (RP)? Higher. The LLM equivalent of brainstorming (or, perhaps, a fever dream), much higher still. At a very high temperature you'll get (potentially interesting) gibberish.

For story bibles and such, low temp is good. It's not being creative, it's working in factual spaces. It'll be absolutely miserable at writing prose with those settings. Retune, and it might be 'okay' at writing prose. Not magic, but better.

It's also worth noting that, for local models, there are a seemingly endless series of 'fine-tunes' for all manner of possible options, very much including RP and storywriting. That doesn't mean they'll fix all of the problems, but they're good options if that's the space you're in.

Model series have personalities. The Qwen family tends to be better at coding and facts (except China-related facts the Chinese government doesn't like), worse at creativity. The Gemma series (open-source, limited-parameter Google Gemini) is better at creative writing.

There is a general consensus in the RP community that large commercial models are getting worse at creative writing and RP, because so much of the competition is based around maximizing benchmarks (aka 'benchmaxxing') for coding tasks. The large providers are not currently pouring a lot of effort into creative writing.

That could change, and the small-LLM community is (to some extent) filling in the gap.

Grey Wolf 🚫

@awnlee jawking

As a slightly more technical explanation, there's a property called 'logit bias'. Most LLM engines support it (all? maybe).

Say my LLM overuses 'mechanical'. I bias it to -5 or -10. The LLM 'thinks twice' about whether to output 'mechanical'.

Put a -999 bias on the em-dash, and you'll get no em-dashes. If the model really, really wants to output an em-dash, though, you might get gibberish - models can 'fight' to get the token they want. In the case of trying to beat the model I was working with over the head to deny it 'mechanical', it resorted to the Chinese, Russian, and Arabic equivalents, plus misspellings. Those can all be solved, but it points out that this is complicated and the results of a simple change aren't obvious.

For em-dash, you might get a different pronunciation mark, but you might get an 'and,' 'but', or some other connective word (contextually, of course). You might get a comma, a period, or a semicolon. And if you only bias it a bit, you'll get fewer em-dashes, but you'll still get them sometimes.

Sarkasmus 🚫
Updated:

@Grey Wolf

Even though, several replies back in the chain you're replying to, we just had numerous citations showing that em-dashes are a poor indicator? Would those be exceptions to 'EVERY study?'

What citations!? I just scrolled through the thread, and all I see are links to Op-Eds and Blog posts, referencing Reddit Posts and YouTube videos as their "Sources"....

Replies:   awnlee jawking
awnlee jawking 🚫

@Sarkasmus

I asked Google why AI uses so many em dashes.

Its AI response was:

AI Overview

AI uses em dashes frequently because it learns statistical writing patterns from formal articles, books, and essays in its training data where these punctuation marks are common.
Here is why this happens:
Training Data Habits

Formal sources: AI models read massive amounts of polished writing from professional authors and journalists who love em dashes.
Copying style: The AI treats the em dash as a normal, high-quality tool to connect thoughts because it appears often in its source material.

Structural Shortcuts

Easy transitions: An em dash is a quick way to add extra information without rewriting a whole sentence.
Avoiding choices: Instead of deciding between a comma, colon, or parenthesis, the AI uses an em dash as a safe default choice.

Human Feedback

Conversational tone: During training, human reviewers often rate fluid, expressive text higher, which might accidentally reward the model for using dramatic pauses and dashes.

Of course, the question was loaded and the AI was eager to please me, but the reasons are consistent with those I've read elsewhere.

AJ

Replies:   Michael Loucks
Michael Loucks 🚫

@awnlee jawking

I asked Dia AI (a homegrown AI from The Browser Company, who make Arc and Dia browsers) about em dashes being a way to detect AI-generated text:

No. Em dashes alone are a very weak and unreliable signal for detecting AI-written text.

Why em dashes are not a good detector

Human and AI writers both use em dashes. Their presence (or absence) says almost nothing on its own about authorship. Many factors muddy the waters:

- Style preference varies wildly. Some human writers love em dashes for pacing and emphasis; others avoid them and stick to commas and semicolons. AI models have also been trained on huge amounts of text that include lots of em dashes, so they reproduce that style naturally.

- Editorial rules and tools matter. Some style guides discourage em dashes; some keyboards or editors make them harder to type, so humans may use double hyphens β€Žβ --⁠ instead. AI models, by contrast, tend to produce "clean" typographic em dashes.

- Simple metrics are easy to game. If a detector looked for "too many" em dashes, a model could be prompted or fine-tuned to avoid them. And a human could consciously start using more em dashes and get mis-flagged.

At best, unusual em-dash patterns might be one tiny feature among hundreds in a serious classifier: ratio of em dashes to commas, consistency of spacing (β€Žβ wordβ€”word⁠ vs β€Žβ word β€” word⁠), distribution relative to sentence length, and so on. Even then, you'd only get a weak correlation, not a reliable test.

What actually matters more

Systems that try to distinguish AI vs human text tend to look at broader, statistical patterns, such as:

- Repetition, blandness, or "too even" sentence lengths and structures.

- Overuse of safe, generic phrasing and transitions.

- Unusual probability patterns relative to a language model (e.g., text that's too likely under a given model).

Even these are far from perfect; modern detectors have high false-positive and false-negative rates, especially on short or edited text. That's why most researchers and organizations now recommend treating AI detection claims skeptically and not making high-stakes decisions based on them.

Practical takeaway

Em dashes can be part of a stylistic "fingerprint," but they're not a dependable way to spot AI. Any method that relies on something this narrow is almost guaranteed to misclassify a lot of human writing and miss a lot of AI writing.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

Repetition, blandness, or "too even" sentence lengths and structures.

I started, but gave up on, a story like that recently. I found it too boring. My other AI tells weren't triggered so I reckon it was probably human-written.

I also sampled a number of stories without the AI-Generated tag but which had a suspicious number of em-dashes in the story descriptions. They all triggered my AI-generated tells.

I guess that shows the authors here either aren't bothered about em-dash tells or are unaware of the controversy.

AJ

Replies:   Switch Blayde
Switch Blayde 🚫

@awnlee jawking

I guess that shows the authors here either aren't bothered about em-dash tells

If the em-dash is the right punctuation, I use the em-dash. Screw the other bullshit.

Replies:   awnlee jawking
awnlee jawking 🚫
Updated:

@Switch Blayde

If the em-dash is the right punctuation, I use the em-dash. Screw the other bullshit.

The problem is that it's difficult to find a wrong use for em dashes. That's why AI became em dash heavy in the first place.

I secretly suspect that, despite em dashes not being the optimum punctuation, to readers they become invisible, like the dialogue tag 'said'.

AJ

Replies:   Switch Blayde
Switch Blayde 🚫

@awnlee jawking

to readers they become invisible, like the dialogue tag 'said'.

I sure hope not. Is the ! invisible to readers? The em-dash, when used properly, is more impactful than the !. Think of interrupted speech. Think of emphasizing something following it.

Replies:   awnlee jawking
awnlee jawking 🚫

@Switch Blayde

Is the ! invisible to readers?

I think some authors, not mentioning Grey Wolf and Michael Loucks, overuse them. So yes, they become less visible. Didn't a famous writer say you should limit them to three per novel or something similar?

The em-dash, when used properly, is more impactful than the !. Think of interrupted speech. Think of emphasizing something following it.

But when the author also uses it to replace commas, parentheses, colons, dashes etc, that impact becomes diluted :-(

AJ

Replies:   Switch Blayde  Grey Wolf
Switch Blayde 🚫
Updated:

@awnlee jawking

But when the author also uses it to replace commas, parentheses, colons, dashes etc, that impact becomes diluted :-(

Dash = hyphen? It should never be used to replace the hyphen.

Like anything else, there's a place for it and there isn't. And when it can be either/or, that's when the skill of the author comes into play.

I believe the comma is abused in SOL stories more than the em-dash. It's used too often, making the reading choppy or impacting pace/flow.

Replies:   Grey Wolf
Grey Wolf 🚫

@Switch Blayde

I tend to agree about commas. I've been working on kicking unnecessary ones out, though, which makes me more aware of 'excess' commas when I see them.

Replies:   Pixy I
Pixy I 🚫

@Grey Wolf

I've been working on kicking unnecessary ones out, though, which makes me more aware of 'excess' commas when I see them.

I always thought the point of commas and periods was to take a breath. About the only thing I remember from English class, back in 1345 (It certainly feels like that) was to read your writing aloud. If you need to pause to take a breath, then you needed some form of literary torture in the form of something called 'punctuation'.

However, I can see that being problematic if you happen to be a bagpipe player, or free-dive, or smoke 300 a day...

Grey Wolf 🚫

@awnlee jawking

Didn't a famous writer say you should limit them to three per novel or something similar?

The problem is that you can find famous writers who agree with nearly everything. J. K. Rowling uses them at a very high pace. She's a famous author. I have no idea if she thinks she overused them or not, though.

There are famous writers who say every dialogue tag should always be 'said'. There are famous writers who say 'said' should be avoided whenever possible.

So much of this is related to readers and what they expect and want, and that can be a moving target.

Michael Loucks 🚫

@Sarkasmus

even if EVERY study ever done on the topic says the opposite.

False on its face. See above.

Replies:   Pixy I
Pixy I 🚫

@Michael Loucks

False on its face.

Falls?

Michael Loucks 🚫

@Pixy I

Falls?

False. 'On its face' is idiomatic speech used to describe something that is obvious.

awnlee jawking 🚫
Updated:

@Pixy I

False on its face.

I was unfamiliar with that expression and Google was no help, but it breaks down into 'false' and 'on its face', meaning it immediately looks wrong but when you investigate deeper you find the opposite. So Michael Loucks is actually saying it's true.

However, from what I've been able to read of those 'studies', they all seem to be opinion pieces with no more validity than anyone else's opinions.

Even the summaries show the opinions are dodgy

em dashes are no longer a surefire sign of AI-generated content /

em-dashes have never been a surefire sign of AI-generated content.

AJ

Grey Wolf 🚫

@awnlee jawking

Normally (and noting that this could be an American / British difference), I would equate 'on its face' to 'prima facie', which is defined as: 'A prima facie case means a party has presented enough basic evidence to prove their claim or win the lawsuit.'

In other words, it's false, and the evidence already presented makes it obvious that it's false.

It does have a connotation of saying the proof is not exhaustive, and there could be some subsequent counterevidence that would dispute the alleged fact, but it also has the connotation of saying that such evidence has not yet been presented.

Which suits this discussion - there is sufficient evidence presented to show that em-dashes are a very poor indicator of AI authorship, especially compared to much more useful 'tells', but that evidence is not exhaustive or definitive.

And, repeating myself, one of the things that seems consistently to be ignored is the domain specificity of the problem. 'Gee, this student paper written by someone who probably doesn't know what em-dashes are and how to use them is full of em-dashes. Maybe we should suspect LLM use' or 'Gee, this forum post likely authored in a window in a web browser without support for '--' becoming an em-dash has a bunch of em-dashes. Do we guess the author knows the proper keystrokes to generate an em-dash or do we suspect LLM use' seem like reasonable arguments to me, as someone on the 'other side' of the argument. There are plenty of writing domains where em-dashes will not be common. Heck, if I start using them in SoL forum posts, that might be a sign I've been replaced by an LLM. I do strongly suspect that LLMs use em-dashes more than the average amateur writer, especially one who isn't writing fiction (though em-dashes have also historically appeared in newspaper and newsmagazine writing).

But, in the domain of published (including self-published) fiction, em-dash use seems to be quite prevalent, with my own quick tests showing that various SoL authors as well as well-known commercial authors use em-dashes at a far higher rate than LLMs do. Thus, em-dash use by itself is likely a lousy indicator of LLM use within the domain of fiction.

Amusingly, 'everything is spelled right and makes grammatical sense' is also an LLM tell. With rare exceptions, they don't tend to make those mistakes (yes, I've seen it happen, and yes, there are some specific examples that are near-certain 'tells').

Now, combine em-dashes with odd constructions, overuse of certain words, unusual names turning up ('Elara' is a major LLM tell, for instance), and you've got something. But you really didn't need the em-dashes for it.

At this point, you have me pondering writing a story (under another pen name, of course) using some LLM constructions, starring an Elara, and tossing in em-dashes, but trying to actually make it a good story (and not actually using an LLM), just to be a brat. Could happen :) Not soon.

One more utterly tangential thought: CMOS (Chicago Manual of Style) says not to use a space before or after an em dash. I disagree, so on that one, I'm a heretic. Without checking, I wonder if LLMs follow CMOS on that point.

Joe_Bondi_Beach 🚫

@Grey Wolf

One more utterly tangential thought: CMOS (Chicago Manual of Style) says not to use a space before or after an em dash. I disagree, so on that one, I'm a heretic. Without checking, I wonder if LLMs follow CMOS on that point.

Failure to follow the CMOS is apostasy of the first order and leads to eternal damnation and the loss of retractible pen use.

~ JBB

Replies:   Grey Wolf
Grey Wolf 🚫

@Joe_Bondi_Beach

Eh. I don't think even the CMOS authors would support that. Sometimes you take your lumps :) I do usually follow it, but not for typesetting em-dashes.

awnlee jawking 🚫
Updated:

@Grey Wolf

Normally (and noting that this could be an American / British difference), I would equate 'on its face' to 'prima facie', which is defined as: 'A prima facie case means a party has presented enough basic evidence to prove their claim or win the lawsuit.'

Googling 'prima facie' meaning in law agrees with your definition except for the rider 'unless contradicted'.

However the dictionary definition

Prima facie is a Latin phrase that translates to "at first sight" or "on the face of it".

which implies something is superficial.

AJ

awnlee jawking 🚫

@Grey Wolf

But, in the domain of published (including self-published) fiction, em-dash use seems to be quite prevalent, with my own quick tests showing that various SoL authors as well as well-known commercial authors use em-dashes at a far higher rate than LLMs do.

That is not supported by my own research, so we live in different universes.

AJ

Replies:   Grey Wolf
Grey Wolf 🚫

@awnlee jawking

I think you are living in the universe of 2024. In the universe of 2026, most LLMs have drastically cut back on em-dash generation.

Michael Loucks 🚫

@Grey Wolf

One more utterly tangential thought: CMOS (Chicago Manual of Style) says not to use a space before or after an em dash. I disagree, so on that one, I'm a heretic. Without checking, I wonder if LLMs follow CMOS on that point.

I always use spaces, so I'm a heretic as well! ;-) And the LLMs I've see do NOT use spaces, but that's not a tell, because that's CMOS!

Diamond Porter 🚫
Updated:

@Grey Wolf

'Gee, this student paper written by someone who probably doesn't know what em-dashes are and how to use them is full of em-dashes.'

This is an instance that cries out for em-dashes or even... parentheses, like this:

'Gee, this student paper (written by someone who probably doesn't know what em-dashes are and how to use them) is full of em-dashes.'

Personally, I think too many students have been told not to use parentheses and have learned to use em-dashes instead. That only changes the punctuation, not the underlying problem, which (in my opinion) is that using too many asides makes it hard for the reader to follow the principal narrative.

Replies:   awnlee jawking
awnlee jawking 🚫

@Diamond Porter

Computer programming has the concept of 'overloading', where symbols have multiple possible functions depending on context. One example might be where '+' can mean adding two numbers together or concatenating two strings.

Using em-dashes for multiple functions makes it harder to discern what each instance is intended to mean, particularly as the rules for English syntax are so laissez-faire.

In a sense, that's why em-dashes are so popular with LLMs. No need to worry about closing em-dashes.

AJ

Unicornzvi 🚫

@Grey Wolf

And, repeating myself, one of the things that seems consistently to be ignored is the domain specificity of the problem. 'Gee, this student paper written by someone who probably doesn't know what em-dashes are and how to use them is full of em-dashes. Maybe we should suspect LLM use' or 'Gee, this forum post likely authored in a window in a web browser without support for '--' becoming an em-dash has a bunch of em-dashes.

I disagree with this. One reason that em-dashes are fairly common in even casual human written text is that several text editors will auto-format to add them by default. The author doesn't need to know how to type an em-dash, they just need to not have taken specific steps to prevent Word, or whatever editor they used from creating them.

Replies:   Grey Wolf
Grey Wolf 🚫

@Unicornzvi

This is probably true. Word seems to do it under some circumstances. However, most require '--', which isn't obvious, and that's especially true for e.g. forum posts. Or, at least, I don't know of any forum software that autogenerates em-dashes.

Replies:   Unicornzvi
Unicornzvi 🚫

@Grey Wolf

As far as I know you're right about the forum software, however plenty of people write in some other text editor (for example Word) and then copy the text to the forum software and hit post.

julka 🚫

@awnlee jawking

"On its face" is the same as "prima facie". A contract which is void "on its face" is one which is obviously void at first look; a statement which is false on its face is one which is immediately and obviously incorrect. Deconstructing the idiom in the particular way you have leads you to the wrong conclusion.

Googling "false on its face" led me immediately to "prima facie" and discussions of the meaning of the idiom, so it seems like you did it wrong.

Michael Loucks 🚫
Updated:

@awnlee jawking

So Michael Loucks is actually saying it's true.

No, he's not. I explained that I meant it is obviously FALSE. I even said so in my explanation! I simply used English, rather than Latin.

Replies:   awnlee jawking
awnlee jawking 🚫

@Michael Loucks

Perhaps Grey Wolf has a point about UK and US using prima facie differently. I most often encounter it followed by 'but' so I expect the initial claim to be incorrect.

AJ

Replies:   Grey Wolf
Grey Wolf 🚫

@awnlee jawking

The connotation normally associated with 'on its face' / 'prima facie', in my experience, is 'Credible evidence has been presented for this, and that evidence doesn't require a lot of work to understand and review. Until further evidence is introduced, this is a settled matter.'

'But' is a perfectly reasonable qualifier and would tend to mean there was an evidence-based challenge. Without the 'but', though, there's no challenge. The 'but' could be 'no, it's not clear and obvious, so it's not prima facie' or 'your evidence doesn't say what you mean it to mean' or 'your evidence is incorrect' or any manner of things, but there has to be a challenge or the assertion stands.

Replies:   awnlee jawking
awnlee jawking 🚫

@Grey Wolf

Just reasserting your original claim shows you didn't read my post. As far as I am aware (which excludes legal circles), in the UK, the use of 'prima facie' implies a rebuttal to follow.

AJ

Replies:   Dominions Son
Dominions Son 🚫

@awnlee jawking

Just reasserting your original claim shows you didn't read my post. As far as I am aware (which excludes legal circles), in the UK, the use of 'prima facie' implies a rebuttal to follow.

My understanding of the usage in the US court system,'prima facie' implies that rebuttal is possible, it doesn't specifically imply that rebuttal will actually happen.

Replies:   awnlee jawking
awnlee jawking 🚫

@Dominions Son

That's what Grey Wolf said. And I believe the meaning is similar in the UK legal system too. But in informal usage in the UK, the use of 'prima facie' generally precedes a rebuttal.

AJ

Replies:   madnige
madnige 🚫
Updated:

@awnlee jawking

in informal usage in the UK, the use of 'prima facie' generally precedes a rebuttal.

So, like 'at first blush'?

E.g., At first blush/Prima facie it seemed a good investment, but closer inspection revealed it as a Ponzi scheme

Replies:   awnlee jawking
awnlee jawking 🚫
Updated:

@madnige

So, like 'at first blush'?

I'm not familiar with that expression but I guess 'at first sight' and 'at first glance' mean the same.

AJ

Switch Blayde 🚫

@awnlee jawking

It's not a myth. Text containing em-dashes is more likely to have been AI-generated.

Ngram is probably not the best way to look at em-dash usage, but for the hell of it I did a search of the β€”.

The usage was relatively flat until 1947 where it shot up, peaking in 1954. Then it dropped quickly, only to start climbing again in 1957 until really peaking in 1979. But then it gradually dropped over the next 30 years or so until leveling out to the lowest level on the chart which begins in 1800.

So in the AI era, it's at its lowest usage. The usage was much higher before AI.

NOTE: Ngram had a note that said: "Replaced β€” with -- to match how we processed the books" so it was actually looking at --.

awnlee jawking 🚫

@Grey Wolf

As I said, our experiences differ. I see more em-dashes in AI-generated text than I do in recently published novels.

If the BBC website tells me its raining outside and I go outside and find it's not raining, I tend to disbelieve the BBC no matter its credentials.

AJ

Rodeodoc 🚫

@Filmphotomaster

Sorry I have no connection to him/her. I've been following the discussion here on AI use and came across their stories and found them interesting. I see she won a Clitorides so some others must agree. No idea how much of her writing is AI.

Oh, and are you suggesting I'm AI generated? Not at all, although Mrs. Rodeodoc still seems to think my tallywhacker is a real machine.

TheDarkKnight 🚫

@Rodeodoc

Isn't "readable AI" an oxymoron?

akarge 🚫

@Rodeodoc

Also, if someone doesn't vote, then their opinion, whether positive or negative, is not made known in the ratings. Therefore, if you don't read and vote ..

Grey Wolf 🚫

@Rodeodoc

This is entirely tangential, but it jumped to mind.

One place I can see AI potentially having a niche is stories that appeal to exactly one person. If you want a story about some particular odd niche / fetish / obsession / whatever, AI will write you one. Or two. Or twenty.

They may well not be good. They may be repetitive, poorly structured, and on and on.

But compared to interests for which there is a significant dearth of stories, having a roll-the-dice, create-your-own-adventure experience of building a disposable story customized for you is an interesting option, and something LLMs will likely get better at.

Publishing those stories? Probably not a good idea. At all.

samuelmichaels 🚫

@Rodeodoc

I just checked a book by L. E. Modesitt, Jr. copyrighted in 2004 and published as a hardcover. It contains many, many em-dashes per chapter. Used in areas where one might use a colon (which are even more rare in fiction) or commas.

So, this is very much a style thing, but common in published fiction pre-LLM era.

Bondi Beach 🚫
Updated:

@Rodeodoc

Megumi Kashuahara

I just read the first two paragraphs of "One Last Wish" and skimmed the rest of the first chapter. Her writing does not work for me. Whether that's AI or not I have no idea, but I can tell you why.

Exactly seven words in two sentences of dialog in one chapter. Sorry, if I didn't care about dialogue I'd read a handy encyclopedia.

The characters are introduced without telling us who they are, although it doesn't take too much to infer who's the mom and who's the dad.

The prose is leaden, souless.

The story is about a Chinese-American? Canadian? family. Not clear whether dad is Chinese or not, but mom is, it's a world-setting plot point. That said, the cover image shows people who look about as Chinese as I do. (I'm not Chinese.)

Perhaps there's a better story of hers to start with?

~ JBB

akarge 🚫
Updated:

@Rodeodoc

I watch this forum in order to be reminded about various stories that I have forgotten or have never even hear of.

Franky, I have about decided that the current parties in this discussion are all opposing AIs and need to be banned since they haven't identified themselves as such.

WHO CARES.

Back to Top

 

WARNING! ADULT CONTENT...

Storiesonline is for adult entertainment only. By accessing this site you declare that you are of legal age and that you agree with our Terms of Service and Privacy Policy.


Log In