# Why AI stories feel flat: what the research found

> Sixteen studies on how AI fiction differs from human fiction: calmer feelings, tidier endings, and stories that all end up alike.

- Author: [Nathanael](https://www.linkedin.com/in/naelnathanael71/), Co-founder & developer, MiraeStudio.id
- Published: 2026-10-06
- Updated: 2026-10-06
- Canonical: https://miraestudio.id/articles/why-ai-stories-feel-flat
- Language: English ([Bahasa Indonesia](https://miraestudio.id/id/articles/why-ai-stories-feel-flat.md))
- Publisher: [MiraeStudio.id](https://miraestudio.id)

In March 2026 Hachette pulled a horror novel called [Shy Girl](https://www.nytimes.com/2026/03/19/books/ai-fiction-shy-girl.html) after a detector flagged it as about 78% AI-generated. It was the first commercially published novel cancelled over AI. Around the same time, roughly one in five of 14,000 self-published Amazon novels came back as largely AI-written. So what does give AI fiction away?

The common answer is that AI writing is "too consistent" and "has no soul". We read the research behind that claim in full, appendices included. Some of it holds up and some of it is backwards. The part that holds up is useful to anyone who writes with AI, us included.

## Ask for a story twice, get the same story

Give a model the same prompt twice and the two stories come out close to each other. A [2026 study](https://arxiv.org/abs/2606.17350) tested this on Claude Opus 4.6, GPT-5.2, Gemini 3.1 Pro and others. A model's second story was the closest match to its first **80 to 90%** of the time. For two human writers answering the same Reddit prompt, it was **about 31%**. Raising the temperature or asking the model to "make it different" moved the numbers a little, and GPT-5.2 not at all.

One study asked gpt-4o-mini for 50 stories from each of 236 countries. Out came [one plot, over and over](https://arxiv.org/abs/2507.22445): a woman leaves the stressful city, returns to her hometown, organises a festival and saves the community. 23 of the 50 American stories had titles starting "The Last Train". 20 of the 50 Norwegian ones were called "The Whispering Pines".

The largest study, [StoryScope](https://arxiv.org/abs/2604.03136), compared 61,608 stories of about 5,000 words from human authors and five models. AI stories sit in one shared region of story space. Human stories spread out much wider.

## The feeling is there, just calmer

"No emotion" is the wrong diagnosis. [Narrative Flattening](https://arxiv.org/abs/2605.27878) measured the emotions sentence by sentence and found the feelings still there, just quieter. Against New Yorker fiction, the finished OLMo model cut conflict from **20% to 7.7%** of sentences, cut surprise and curiosity from **21% to 13%**, and raised neutral sentences from **29% to 45%**. Sadness and anxiety stayed at or slightly above human levels.

The raw base model, before any chat training, overshoots the other way: **79%** of its sentences carry conflict or surprise. So the flattening comes from the training that turns a base model into a polite assistant, and that training trims the extremes.

## Feelings explained, endings tidied

StoryScope could tell AI stories from human ones at **93.2%** accuracy without looking at the words at all, only at how the story is built. The differences:

- **The narrator states the theme** in 77% of AI stories, against 52% of human ones.
- **No subplots** in 79% of AI stories, against 57% of human ones. Human stories also use more flashbacks and broken timelines.
- **Morally mixed main characters** in 59% of human stories, against 38% of AI ones.
- **AI shows emotion through the body** (a tight throat, a racing heart) in 81% of stories, against 38%. Humans more often just name the feeling.
- Each model has habits. Claude's plots barely escalate, GPT likes gossip as a plot device, and Gemini ends tidily after a long wind-down.

Expert readers saw the same thing earlier. In [Art or Artifice](https://arxiv.org/abs/2309.14556), professional writers ran 14 creativity tests on stories. New Yorker stories passed **84.7%** of them, GPT-4 and Claude about **28 to 30%**. One judge wrote that the AI "rarely knew how to end a story": each one kept growing in scope until it was about legacy and the community. Others flagged "overtelling instead of showing" and dialogue with no subtext.

## Where "too consistent" breaks down

At the level of words, the claim is backwards. A [2026 stylometry study](https://arxiv.org/abs/2608.27855) found AI-written text has more varied vocabulary than human text, about two standard deviations more. The sameness is in plot and feeling.

Detection also fails in both directions. When AI only edits a human draft, the classifier that catches AI writers spots the edit just **8.5 to 32%** of the time. Going the other way, detectors flagged essays by non-native English writers as AI [61% of the time](https://arxiv.org/abs/2304.02819). The best commercial detector, Pangram, gets close to 0% false positives on clean human text, but flagged [64 to 80%](https://arxiv.org/abs/2608.11256) of human abstracts that had only been lightly AI-edited.

## Many readers don't notice, or don't mind

Lay readers picked the author of a poem correctly only [46.6% of the time](https://www.psypost.org/people-cannot-tell-ai-generated-from-human-written-poetry-and-they-like-ai-poetry-more/), below a coin toss, and rated the AI poems higher. Across 1,471 stories, lay readers preferred the AI version [57.6% of the time](https://arxiv.org/abs/2506.03310). Creative-writing experts preferred it in 2.2% of cases.

Two more results complicate the "no soul" story. When a model was [fine-tuned on one author's complete works](https://arxiv.org/abs/2510.13939), MFA-trained readers had 8× the odds of picking the AI excerpt as truer to that author's style. And simply [labelling a story as AI](https://academic.oup.com/joc/article-abstract/74/5/347/7756907) lowered how absorbed readers were, whoever really wrote it.

## If you write with AI

There is a cost when people lean on AI too. In a [Science Advances study](https://www.sciencedaily.com/releases/2024/07/240712222127.htm), one AI idea made writers' stories better, and also 10.7% more alike. For a brand, that means copy that sounds like everyone else's. What to check:

1. **Cut the moral.** If the last line explains what the piece meant, delete it. The reader got it.
2. **Leave something open.** Not every thread needs to resolve. Leave in the trade-off, the question without an answer, the person who was partly wrong.
3. **Keep the friction.** Keep the complaint, the bad month, the client who said no. Calm, positive copy with no conflict is the default, and it reads like one.
4. **Name things.** Name the street, the date, the number, the person. Vague allusion is a model habit; humans name their references about twice as often.
5. **Break the rhythm.** One short section. One long one. A paragraph that stops early.

Sampling settings help a little. [Verbalized sampling](https://arxiv.org/abs/2510.01171), where the model is asked to list several possible answers with odds, made stories 1.6 to 2.1× more diverse. None of these tricks closed the gap with human writers.

## Want copy that sounds like you?

We are [MiraeStudio.id](https://miraestudio.id), a small design and development studio in Indonesia. The words on a website are part of what we design, and this research is the checklist we hold them to. If your website reads like every other site in your field, [send us a brief on WhatsApp](https://wa.me/6289503386642). Tell us what you do and who it's for. We'll tell you what we'd change.

## References

- Russell et al., [StoryScope: Investigating idiosyncrasies in AI fiction](https://arxiv.org/abs/2604.03136), COLM 2026.
- Li, Zhu, Wu, Bao, Evans, [Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction](https://arxiv.org/abs/2605.27878), arXiv 2026.
- Thennal D K, Hatzel, [Do Large Language Models Always Tell The Same Stories?](https://arxiv.org/abs/2606.17350), arXiv 2026.
- Xu et al., [Echoes in AI: Quantifying lack of plot diversity in LLM outputs](https://www.pnas.org/doi/10.1073/pnas.2504966122), PNAS 2025.
- Rettberg, Wigers, [AI-generated stories favour stability over change](https://arxiv.org/abs/2507.22445), Open Research Europe 2025.
- Chakrabarty et al., [Art or Artifice? Large Language Models and the False Promise of Creativity](https://arxiv.org/abs/2309.14556), CHI 2024.
- Marco et al., [Pron vs Prompt](https://arxiv.org/abs/2407.01119), EMNLP 2024.
- Shan, Lee, Hao, [AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not](https://arxiv.org/abs/2608.27855), arXiv 2026.
- Liang et al., [GPT detectors are biased against non-native English writers](https://arxiv.org/abs/2304.02819), Patterns 2023.
- Karr et al., [Pangram and GPTZero on academic abstracts](https://arxiv.org/abs/2608.11256), arXiv 2026.
- Porter, Machery, AI-generated poetry is indistinguishable from human-written poetry, Scientific Reports 2024 ([summary](https://www.psypost.org/people-cannot-tell-ai-generated-from-human-written-poetry-and-they-like-ai-poetry-more/)).
- Marco et al., [The Reader is the Metric](https://arxiv.org/abs/2506.03310), arXiv 2025.
- Chakrabarty, Ginsburg, Dhillon, [Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers](https://arxiv.org/abs/2510.13939), arXiv 2025.
- Chu, Liu, [AI disclosure and narrative persuasion](https://academic.oup.com/joc/article-abstract/74/5/347/7756907), Journal of Communication 2024.
- Doshi, Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances 2024 ([summary](https://www.sciencedaily.com/releases/2024/07/240712222127.htm)).
- Zhang et al., [Verbalized Sampling](https://arxiv.org/abs/2510.01171), arXiv 2025.
- The New York Times, [Shy Girl and AI fiction](https://www.nytimes.com/2026/03/19/books/ai-fiction-shy-girl.html), 19 March 2026.

## Questions

### Can you tell if a story was written by AI?

Often, across many stories: AI stories resemble each other and share habits like stated themes and tidy endings. For a single text, no method is reliable enough to accuse someone. Detectors miss AI-edited text and flag some human writers, especially non-native English writers.

### Is AI writing really emotionless?

No. The research finds emotion is still there but calmer: less conflict and surprise, more neutral sentences, and feelings explained rather than shown. The flattening comes mainly from the chat training that makes models polite assistants.

### Are AI stories more repetitive in their wording?

Not at the word level. One 2026 study found AI text uses more varied vocabulary than human text. The repetition is in plots, structure and emotional range.

### How do I make AI-assisted writing sound more human?

Change the choices, not the words: cut the closing moral, leave something unresolved, keep real conflict, name specific people, places and numbers, and vary the length of sections. Never add made-up details to sound human.
