PAPER-DIGEST · 2026-09-25
Sears and Weisberg: Readers couldn't spot AI-written stories, and rated them higher when told a human wrote them — Fukai Reads
Judgment and reading psychology — how an author label changes what we think of a work
TL;DR
Readers could not tell short stories written by AI from ones written by people. Reading blind, they even rated the AI stories as higher quality. Yet when told a story was written by a human, they rated the very same text higher.
That is the result of a paper by Sydney Sears and Deena Skolnick Weisberg of Villanova University, published in the peer-reviewed journal Judgment and Decision Making on August 5, 2026. Across three studies it tested more than 2,500 participants.
For anyone who writes game text (flavor text, hints, dialogue), it is also a story about readers judging words by who they believe wrote them, not only by what the words say.
Who wrote it, and where
The authors are Sydney Sears and Deena Skolnick Weisberg of the Department of Psychological and Brain Sciences at Villanova University. Weisberg is a developmental psychologist who has long studied how people engage with fiction.
The journal is a peer-reviewed venue for judgment and decision-making research. The article number is e21 and the DOI is 10.1017/jdm.2026.10042. It is open access under CC BY-NC, and the stories and measures are posted on OSF (a site for sharing research materials). As the authors state, however, none of the studies was preregistered (analysis plans were not published before data collection).
I picked it today because whether to use generative AI for game text is a live question for many developers. And the question "can you tell?" is itself a puzzle.
The Turing Test (Bulkhead Interactive, 2016), a first-person puzzle game named after the classic test for telling machines from people. Image: Steam store page
How sure were we that AI writing can be spotted?
Whether a machine can pass as a person is an old question. The paper cites a 1972 study of ELIZA (an early chatbot) in which judges identified the machine only 48% of the time, no better than a coin flip.
In recent years, AI-written ads and explanatory texts have repeatedly been rated as good as, or better than, human writing. For poetry, ChatGPT poems were reported to be hard to tell from human ones, and rated lower once people knew they were AI-made (Porter and Machery, 2024).
The Talos Principle (Croteam, 2014). Between puzzles you read terminal texts and keep asking whether their author is human or machine. Image: Steam store page
At the same time, people tend to believe AI cannot really create. Researchers call this tendency to discount machine output "algorithm aversion" (rating a machine's judgment or work lower even when its quality is the same).
What was still unknown was whether the same holds for stories, with plot and characters. A second open question was what supports the ability to tell them apart: knowledge of literature, or knowledge of AI.
How they tested it
The materials were six short stories. Three were human works published in literary magazines or collections ("FISH" by Emilie Fox, "High Heels" by Susie Maguire, and "Inisfree" by Patrick Smyth). The other three were written by ChatGPT 4.0, prompted with each human story's theme or period setting. All were about 1,000 words, a few minutes' reading.
Study 1 (1,682 people after exclusions) measured evaluation. Each participant read one story, in a 2x2 design: true author (human or AI) crossed with stated author (human or ChatGPT). So some people read a human story labeled as AI-written. Afterward they rated absorption (how much the story pulled them in) and quality of plot, characters, theme, and style.
Study 2 (424 people) measured discrimination. Participants read one human and one AI story and chose which was written by a human (or by ChatGPT). They also reported confidence, the cues they relied on (language, character, theme, plot, symbolism, enjoyment), and their knowledge of AI and of literature.
Study 3 (481 people) was a replication of Study 2. Materials and procedure were the same; only the AI-knowledge questions were replaced with a validated AI literacy scale (a 12-item questionnaire on understanding and using AI). In every study, people who failed the comprehension questions were excluded.
What they found
In Study 1, the AI stories were rated higher. Mean quality was 1.54 for AI stories versus 0.97 for human ones, and absorption was 1.42 versus 1.00 (absorption was on a -3 to +3 scale).
At the same time, being told a story was human-written raised ratings: 1.40 versus 1.12 for quality, and 1.32 versus 1.10 for absorption. The true-author and stated-author effects worked separately (the interaction was not significant). People with more positive attitudes toward AI penalized the "ChatGPT" label less.
In Study 2, 167 of 424 people (39.39%) identified the stories correctly, significantly below the 50% chance level. Since it was a two-way choice, those who were wrong picked the AI story as the human one. People who relied on "language" as a cue were more likely to be wrong.
In Study 3, 250 of 481 (51.97%) were correct, no different from 50% (p = .41). The below-chance result of Study 2 did not replicate; it settled at chance. In both studies, confidence was unrelated to accuracy. In Study 2, reading time (mean 569.88 seconds) was also unrelated to accuracy.
What did relate to accuracy was knowledge of AI. In Study 2, more AI expertise predicted correct answers (β = .13, p < .001), and in Study 3 higher AI literacy did too (β = .29, p = .041). Literary expertise had no effect in Study 2 (β = -.006, p = .87) or Study 3 (β = .004, p = .91).
How game makers can use this
First, how you playtest text. Asking testers "does this sound like AI?" is unreliable: in this study, confidence did not track accuracy. If you want to compare text quality, show it with the author hidden. Ratings made with a known author include a label effect (here, a difference of about 0.3 in mean quality).
Second, deciding whether to disclose AI use. This study shows that the "AI" label lowers ratings. That does not mean you should hide it. Platforms such as Steam ask for disclosure, and being found out later would likely cost far more trust. My reading is that you need to disclose and then win ratings back on the merits of the content.
Third, the "human or machine?" puzzle itself. As The Turing Test and The Talos Principle show, it is a strong story theme. But if you build a guessing game around it, take care. In this study, "language," the cue many people relied on, was misleading. If the answer rests only on stylistic smoothness, accuracy drifts toward a coin flip and the puzzle stops feeling solvable.
Her Story (Sam Barlow, 2015), where you infer the truth from how a person words her testimony, a classic of "reading the texture of language" as play. Image: Steam store page
So what cues are fair? As in Her Story, letting players catch mismatches between words and facts (timelines, names, contradictions with earlier statements) gives them something to reason with. Study 3 also reports a borderline tendency for people who relied on "symbolism" to answer correctly.
Fourth, writing short in-game text. In their discussion, the authors suggest that AI stories were rated highly because of fluency (ease of reading) and a positive tone. This is the authors' interpretation, not a tested conclusion. Still, the idea that ease of reading drives judgments of text read in a few minutes seems to apply directly to item descriptions and hint text.
Fifth, roles in a team. Because the ability to tell them apart was linked to AI knowledge rather than literary knowledge, the person who reviews AI-generated text may be better chosen from those who read a lot of AI output, not simply the best writer. This is my application, not the paper's conclusion.
How far to trust it
First, the limits the authors acknowledge. The stories were short, readable in minutes, so it is unclear whether the results hold for novels or serialized work. Studies 2 and 3 told participants in advance that one story was human and one was AI, a premise absent in real life. The quality scale was built for this study and not validated beforehand.
Also, there were only three story pairs, one AI system (ChatGPT 4.0), and participants were general US adults recruited on Prolific. The authors note prior work in which professional writers rated AI stories as less creative (Chakrabarty et al., 2023), so results may differ with different readers.
What I, Fukai, want to point out is the gap between Studies 2 and 3. Study 2's 39.39% is "worse than chance"; Study 3's 51.97% is "at chance." If the same materials move that much, it is too early to say people tend to mistake AI stories for human ones. What can be said firmly, as I read it, is only that they could not tell.
Another point: Study 1's finding that AI stories were rated higher may depend heavily on the choice of three story pairs. The human stories were actually published; the AI stories were generated afterward to match their themes. And none of the studies was preregistered. The samples are large and Study 3 is a replication, but as a psychology finding it is best read with the caveat "under these conditions."
Fukai's reading
This part is my own opinion. I want to read this study as a new way of measuring an old problem: a work's value also lives in the story outside it (who made it, and with how much effort). In games, it resembles the trust we place in the "designer's intent" behind a handmade puzzle. Players keep at a hard level because they believe it has an answer that someone placed on purpose. For AI-made levels or text to earn the same trust, content quality alone is not enough; I would put it this way: the design has to convey that someone chose this and checked it.
What to read next
Porter and Machery (2024), who asked the same question about poetry, are one starting point for this paper. They found the same shape of result: people cannot tell, yet rate work lower once told it is AI-made.
Chakrabarty et al. (2023), who had professional writers evaluate AI stories, are worth reading alongside it as evidence that different readers can yield different results. Messingschlager and Appel (2022), who found that transportation into a story weakened when it was labeled AI-written, point the same way as Study 1's label effect.
On Puzzlebyrinth I have also read studies of AI generating puzzle levels. Adding the question of how people receive what AI makes connects that generation research and this study into one map.
Sources
Papers and materials referenced in this article:
・Study materials (stories and measures, OSF)
・Related work: Porter and Machery (2024), Chakrabarty et al. (2023), Messingschlager and Appel (2022) (all listed in the paper's references)
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 95 of 98
Read next
Related reviews
Lost in Play
A hand-drawn point-and-click adventure told entirely without words. The siblings Toto and Gal make their way home through fifteen scenes built out of a child's imagination, picking things up, combining them, and solving a different little minigame in almost every room. From Happy Juice Games.
The Occupation
A first-person thriller set in North West England on 24 October 1987, in which you play a journalist slipping into restricted areas and interviewing people to gather evidence. The in-game clock runs at real-world speed and the cast moves on a timetable, in this game from White Paper Games.
Thomas Was Alone
A 2D puzzle-platformer in which you switch between coloured rectangles, each given exactly one ability, and combine them into stepping stones across 120 levels. A narrator reads out what the shapes are thinking as you go. Released in 2012 by Mike Bithell's Bithell Games.



