Statistics

AI Poetry Statistics: Authorship, Readership, and Writing Studies

AI poetry statistics covering authorship judgments, memorization, reader reach, publishing activity, and human-versus-AI writing studies.

AI poetry research now reaches beyond whether a poem sounds convincing. Studies measure how often readers identify machine-generated work, how models reproduce poems, how large the poetry audience is, and where human and AI writing differ.

Table of contents

AI poetry identification statistics

The main test of AI poetry recognition in this dataset comes from a non-expert reader study. Readers identified AI-generated poems with 46.6% accuracy (Scientific Reports / DOAJ). That result places authorship recognition below a reliable half-and-half distinction in the reported study, making the reader’s experience especially important for poetry guides that discuss machine-assisted writing.

The discrimination study collected 16,340 total poem judgments (Scientific Reports / Nature PDF). Its central authorship effect was statistically significant at chi2(2, N=16,340)=247.04 (Scientific Reports / Nature PDF). The reported finding was not simply that some poems were difficult to classify. Participants were more likely to judge AI-generated poems as human-authored than actual human-authored poems (Scientific Reports / DOAJ).

The authors reported that AI-generated poems received more favorable ratings for rhythm and beauty (Scientific Reports / DOAJ). In the same research, poetry was described as one of the last remaining domains where generative AI had not reached indistinguishability before the study (Scientific Reports / Nature PDF). That statement is the researchers’ framing of the field at the time of the study, not a timeless claim about every model or poem.

A useful comparison from the same study is the reported odds ratio. The odds that a human-written poem would be judged human-authored were roughly 75% of the odds for an AI-generated poem being judged human-authored, with OR=0.758 (Scientific Reports / Nature PDF). The statistic describes the direction of the judgment effect; it does not mean that every reader or every poem followed the same pattern.

The study recorded 16,340 poem judgments, while reported identification accuracy was 46.6% (Scientific Reports / Nature PDF; Scientific Reports / DOAJ).

Who took part in the authorship studies

The first study recruited 1,634 US-based participants (Scientific Reports / Nature PDF). The median age in Study 1 was 37 (Scientific Reports / Nature PDF). Its reported participant profile was 49.6% male, 48.5% female, and 1.9% non-binary or preferred not to say (Scientific Reports / Nature PDF).

The second study recruited 696 US-based participants (Scientific Reports / Nature PDF). The median age in Study 2 was 40 (Scientific Reports / Nature PDF). The reported profile was 50.4% male, 46.6% female, and 3.0% non-binary or preferred not to say (Scientific Reports / Nature PDF).

MeasureStudy 1Study 2
US-based participants1,634696
Median age3740
Male participants49.6%50.4%
Female participants48.5%46.6%
Non-binary or preferred not to say1.9%3.0%

These figures identify the populations reported for the two studies. They should not be generalized to all poetry readers, all English-language readers, or all people who use generative AI. The dataset does not supply a broader population estimate. `n

What language models memorized

A separate memorization study tested poems from 60 American poets (Cornell Chronicle). ChatGPT successfully retrieved 72 of 240 tested poems (Cornell Chronicle). PaLM retrieved 10 of 240 tested poems (Cornell Chronicle). Pythia retrieved 0 entire poems, and GPT-2 also retrieved 0 entire poems in the same study (Cornell Chronicle).

ModelEntire poems retrieved
ChatGPT72 of 240
PaLM10 of 240
Pythia0
GPT-20

The results compare retrieval counts within the reported test set. They do not measure the total number of poems in any model’s training data, nor do they establish how often a model generates text that resembles a poem without retrieving an entire poem.

The Cornell study reported that inclusion in the Norton Anthology of Poetry was the most reliable predictor of ChatGPT memorization (Cornell Chronicle). That finding connects memorization performance with the prominence of the source poems in the study. It does not provide a general memorization rate for every poet, anthology, or publishing venue.

For readers evaluating AI poetry, the distinction between authorship judgment and memorization is important. The first body of research measures whether people can distinguish human and AI authorship. The Cornell study measures whether tested systems could retrieve entire poems. These are related questions about AI and poetry, but they are not the same measurement.

The scale of poetry readership

Poetry also has a measurable audience independent of AI-generated writing. The Academy of American Poets reported that 18,015,092 individuals worldwide read poems on Poets.org in 2021 (Academy of American Poets). Poets.org traffic increased 10% versus the same period before the pandemic (Academy of American Poets).

The Academy listed the top 10 metro areas for Poets.org poetry readers in ranked order, from New York, NY through Seattle-Tacoma, WA (Academy of American Poets). The supplied research preserves that ranking description but does not provide the individual metro-area counts.

A separate Poetry Matters source said that more than 1,000,000 people in the United States read poems at Poets.org each year (Poetry Matters). This annual US figure and the 2021 worldwide figure describe different geographies and time frames. They should be kept separate rather than combined.

The Academy’s 2016 annual report recorded more than 140,000 Poem-a-Day readers (Academy of American Poets annual report). In the same year, Poets.org reached more than twenty million individuals through Academy programs (Academy of American Poets annual report). The annual report also said its magazine distribution reached more than 9,000 individuals, with readers in every state (Academy of American Poets annual report).

These figures show several different ways of counting poetry engagement: website readers, subscribers, program reach, and magazine distribution. They are not interchangeable audience totals.

Poetry publishing and participation

The Poetry Matters source reported 365,000 students participated in Poetry Out Loud in 2014 (Poetry Matters). It also reported that more than 160,000 poetry fans followed poetsorg.tumblr.com, while more than 93,000 readers subscribed to Poem-a-Day (Poetry Matters).

The same source recorded 140,000 individuals who had attended the Dodge Poetry Festival since 1986 (Poetry Matters). Although that is a historical audience measurement, it helps show the scale of participation tracked by poetry organizations.

The publishing ecosystem in the Poetry Matters figures included 2,900 new poetry books and poetry-related texts featured in Poets House’s 2013 annual Showcase (Poetry Matters). It also included 931 poetry journals actively publishing poems and 278 small poetry presses actively publishing books of poems (Poetry Matters).

Institutional and educational activity was similarly broad. Guidestar tracked 855 nonprofit organizations that presented or supported poetry (Poetry Matters). There were 224 conferences and residencies offering poetry programs, and 221 graduate writing programs in poetry attended by thousands of students (Poetry Matters). The dataset also reported 100 poetry venues and sites that regularly hosted poetry slams (Poetry Matters).

Other public-facing measures included more than 200 poems displayed in New York City subway cars since 1992 through Poetry in Motion and more than 35 US cities with a local Poet Laureate position (Poetry Matters). These are counts of programs, organizations, venues, and civic positions, not estimates of AI adoption.

National programs and institutional support

The Academy said National Poetry Month was inaugurated in 1996 and described it as the largest literary celebration in the world (Poets.org). More than 60 poetry partners and sponsors supported the observance (Poets.org).

The Academy of American Poets was founded in 1934 (Academy of American Poets). The Academy annually awards $1.25 million to more than 200 poets (Academy of American Poets). It also says Poets.org is the world’s largest publicly funded website for poets and poetry (Academy of American Poets).

Its 2016 annual report provides additional activity measures. Poets.org added more than 500 new poems, 100 biographies of poets, and 200 audio and video clips that year (Academy of American Poets annual report). Teach This Poem grew to more than 13,500 educators and exceeded its goal by 30% (Academy of American Poets annual report).

These institutional statistics give context for AI poetry discussions. Generative tools enter an established field with readers, educators, journals, presses, festivals, public programs, and professional organizations. The figures do not measure how many of those participants use AI, but they do quantify the surrounding poetry infrastructure.

Human and AI writing comparisons

The GenAI poetry ownership study recruited 88 college writers (Written Communication / Sage). It reported that human-made poems had significantly greater ownership than AI-made poems (Written Communication / Sage). Human-made poems were also rated as more accurately reflective of lived experience than AI-made poems (Written Communication / Sage).

The same study found that AI-generated poems scored higher for imagery, language, and form than human-made poems (Written Communication / Sage). Half of the students preferred GenAI poems, while less than half preferred human poems (Written Communication / Sage).

Those findings describe several dimensions rather than one overall quality score. Ownership and reflection of lived experience favored human-made poems in the reported results, while imagery, language, and form favored AI-generated poems. Preference was divided between the two categories.

The available figures do not give the exact participant count behind the preference percentages beyond the study’s 88 college writers, and they do not supply scores for each rating dimension. It is therefore more precise to retain the study’s directional findings than to create unsupported averages or rankings.

What writing datasets measure

The CoAuthor dataset included 63 writers and four GPT-3 instances (CoAuthor). It recorded 1,445 writing sessions in English (CoAuthor). Its creative-writing subset contained 830 stories written by 58 writers, while its argumentative-writing subset contained 615 essays written by 49 writers (CoAuthor).

The average story or essay length in CoAuthor was 418 words (CoAuthor). The dataset averaged 11.8 queries per writing session and had a 72.3% acceptance rate for suggestions (CoAuthor).

CoAuthor measureReported figure
Writers63
GPT-3 instances4
English writing sessions1,445
Creative-writing stories830
Writers in creative-writing subset58
Argumentative essays615
Writers in argumentative-writing subset49
Average story or essay length418 words
Average queries per session11.8
Suggestion acceptance rate72.3%

CoAuthor measures interaction between writers and an AI writing system across writing sessions. Its reported acceptance rate is not a measure of poem quality, authorship recognition, or literary value. Likewise, its stories and essays are not a direct count of poems.

Written by

anvilpresspoetry.com Editorial Team

Editorial team

anvilpresspoetry.com publishes practical how-to guides and educational articles with clear steps and useful context.