<\/span><\/h2>\nThree things, worth watching rather than assuming.<\/p>\n
The poetry result does not sit in an undisputed literature, and the authors say so themselves. They note that their findings contrast with earlier work in which participants were<\/em> able to distinguish professional poets from human-out-of-the-loop generated poems, and with other work finding participants merely at chance rather than below it. Earlier studies also generally found generated poems evaluated more negatively, the opposite of what these experiments produced. The most likely reading is that the models moved, not that the earlier researchers were wrong, but that is an interpretation and not something the data establishes.<\/p>\nBoth studies use general-population samples doing an unfamiliar task under artificial conditions. Neither tested whether a domain expert evaluating work in their own domain<\/em>, with time and stakes, performs differently. Porter and Machery tested poetry experience among poetry readers, which is closer, but self-reported poetry background is not the same as professional expertise. The null on expertise is real and replicated, but it is a null on the expertise these studies measured.<\/p>\nThe Jones and Bergen result is a preprint at the time of writing, and the two populations disagreed meaningfully on several exploratory measures. The LLaMa persona condition, for instance, was 45% among undergraduates and 65% among Prolific workers, a 20-point spread on the same condition. Population effects here are not small.<\/p>\n
And both studies describe a specific generation of systems. If detection tooling, provenance standards, or content credentials become reliable and widespread, step 1 of the protocol changes from “stop trying” to “check the signature.” Nothing in the current data anticipates that, but nothing rules it out either.<\/p>\n
None of that changes what to do this quarter. Stop using authorship as a proxy for quality, because you cannot recover authorship. Strip labels before you judge. Ask what is new rather than whether it reads well.<\/p>\n
<\/span>FAQ<\/span><\/h2>\nAre people really worse than random at spotting AI writing?<\/strong>
\nIn this study, yes, and significantly so: 46.6% accuracy across 16,340 judgements where guessing would give 50%, chi-squared(1) = 75.13, p < 0.0001. The authors interpret below-chance performance plus above-chance agreement between participants as evidence of a shared but inverted heuristic, not of random answering.<\/p>\nDoes knowing a lot about AI help?<\/strong>
\nOn the available evidence, no. Jones and Bergen found no significant effect of level of knowledge about language models or frequency of chatbot interaction in either of their two studies, and note that accuracy was homogeneous even among people who research these systems. Porter and Machery found no poetry-experience variable with a significant positive effect on accuracy.<\/p>\nWhat tells do people wrongly rely on?<\/strong>
\nFirst-person pronouns, contractions, and family or personal topics. A computational analysis across six experiments identified these as the heuristics readers use to infer human authorship, and the authors then showed experimentally that because the heuristics are predictable they can be targeted, producing text rated as more human than human. If you catch yourself thinking a passage feels human because it says “I” and uses contractions, that is the documented failure mode.<\/p>\nSo confidence is always inversely related to accuracy?<\/strong>
\nNo, and the honest answer is messier. Porter and Machery found a significant negative effect of confidence (b = -0.021673, p < 0.0001). Jones and Bergen found self-estimated accuracy positively correlated with real accuracy among undergraduates (p = 0.03) but not among Prolific participants (p = 0.45). Confidence is not a reliable signal in either direction, which is the usable conclusion.<\/p>\nDoes this mean AI writes better than humans?<\/strong>
\nNo. It means readers rated it higher on the dimensions that reward smoothness, 13 of 14 of them, with rhythm the largest at d = 0.847. On originality, the one dimension requiring the reader to detect something new, genuinely AI-authored poems showed no significant advantage (d = 0.040, p = 0.098). That is a finding about what readers reward, not about literary merit.<\/p>\nShould I label AI-assisted work?<\/strong>
\nThat is an ethics question this data cannot settle, but it can tell you the cost. Labelling work as AI-generated reduced quality ratings by 0.814 points on a seven-point scale, d = -0.508. Readers penalise the label independently of the work. Anyone deciding on disclosure should price that in rather than be surprised by it.<\/p>\nWhat is the single most useful change?<\/strong>
\nEvaluate blind. The label effect is roughly three-quarters the size of the actual quality effect, and it is the one variable you fully control. Removing bylines before judging work you intend to act on costs nothing and removes a bias of measured size.<\/p>\n<\/span>Sources<\/span><\/h2>\nPorter and Machery. AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Scientific Reports.<\/p>\n
Jones and Bergen. Large Language Models Pass the Turing Test. arXiv preprint, Computation and Language.<\/p>\n
Jakesch, Hancock and Naaman. Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences, 2023.<\/p>\n
Kobis and Mossink, human-in-the-loop and human-out-of-the-loop paradigms for evaluating machine-generated poetry, as characterised in Porter and Machery.<\/p>\n
Note on scope: the combined participant figure of 6,518 across the three detection studies is a CEOtudent sum of published sample sizes, used only to indicate scale. The three studies differ in design, population and outcome measure and are not pooled statistically anywhere in this article.<\/p>\n
\nThis content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"Two large preregistered studies, run by different labs in completely different domains, converge on the same uncomfortable finding: people cannot reliably identify machine-generated work, expertise does not help, and confidence runs the wrong way. Worse, the evaluation itself is roughly three-quarters as sensitive to the label on a piece of work as to the work. This piece assembles both datasets, derives what the presentation layer is actually worth against the substance layer, and sets out the only probes that measurably improved accuracy.<\/p>\n","protected":false},"author":1,"featured_media":325915,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4599,18],"tags":[],"class_list":["post-325909","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-gelisim","category-strateji"],"_links":{"self":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts\/325909","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/comments?post=325909"}],"version-history":[{"count":0,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/posts\/325909\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/media\/325915"}],"wp:attachment":[{"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/media?parent=325909"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/categories?post=325909"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ceotudent.com\/en\/wp-json\/wp\/v2\/tags?post=325909"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}