Poetix · the evaluating models · 20 September 2026
The Reader Panel
Before trusting an average, look at what it averages. Five models read the same poems. Their scores differ, their judgments often move together, and their average is more repeatable than an individual reading.
Five readers, one poem
Readers see the poems without the manuscripts or production records. Each scores insight, imagery and cadence from 1 to 4, for a total of 3–12, and names the strongest and weakest poem in its group. The reported panel score is the average of the five readers.
Here the models are evaluators. A generous score does not tell us how well that model writes poetry. These are tendencies observed in this panel and these reading sessions.
175 poems · 875 scores · PX-08
Every poem, every reader
The average gives each poem one number. The dots show what that number hides: five distinct readings, sometimes several points apart.
Scroll sideways to explore the full chart, or open it large below.
The highest and lowest scores for the same poem are 3.0 points apart on average in PX-08. For 15 of the 175 poems, the spread is five or six points. ChatGPT and Grok score higher on average; Claude, Gemini and Qwen lower. That pattern does not remove disagreements over individual poems.
Different score levels · related judgments
How the readers differ
ChatGPT has the highest mean and Grok the second-highest in all three sittings. The order of Claude, Gemini and Qwen changes. The panel shows recurring differences in scoring level, rather than five interchangeable judges.
Scroll sideways to explore the full chart, or open it large below.
Across the three sittings, pairwise score correlations range from 0.68 to 0.83. Readers tend to score the same poems higher or lower, but they neither give identical scores nor always choose the same favorite. All five chose the same strongest poem in 10 of 35 PX-07 groups, 11 of 35 PX-08 groups, and 9 of 35 reread groups.
Compare each reader’s mean score
| Reader | PX-07 · 174 poems | PX-08 · 175 poems | Reread · same 175 |
|---|---|---|---|
| ChatGPT | 9.10 | 8.77 | 9.02 |
| Claude | 7.53 | 7.02 | 7.22 |
| Gemini | 6.95 | 7.10 | 6.95 |
| Grok | 8.32 | 8.43 | 8.50 |
| Qwen | 7.39 | 7.43 | 7.30 |
Same poems · fresh sessions
What holds when they read again?
PX-08’s 175 poems were read again by the same five-reader panel, using packets identical apart from the signature line. An individual reader’s score changed by 0.68 points on average; the five-reader mean changed by 0.38. The correlation between the two sets of panel means was 0.97.
Scroll sideways to explore the full chart, or open it large below.
The panel’s averages were repeatable in this reread. The picture is less uniform within each writer: Grok’s manuscript-first contrast changes from −0.11 to +0.51. A steady overall mean can coexist with a changing local judgment.
Compare each reader’s change on the same poems
Mean absolute change, in points, over 175 poems per reader. This measures reread consistency, not accuracy.
| Reader | Mean absolute score change |
|---|---|
| ChatGPT | 0.62 |
| Claude | 0.43 |
| Gemini | 0.68 |
| Grok | 0.55 |
| Qwen | 1.11 |
What the average can tell us
Averaging can reduce individual reading fluctuations; it does not automatically cancel systematic bias or establish an objective measure of poetic quality. A repeatable panel may still share preferences or blind spots.
For context, matched writer–question–condition cells changed by about 1.52 points on average between PX-07 and PX-08, when poems were generated and read again. The reread isolates a narrower change: the poems stay fixed. This supports investigating variation in generation, while leaving the precise division of variance conditional on the study’s assumptions.
PX-09 · a new packet · three readings
Each reader, three times over
The 2.0 smoke trial supplied another repeatability check: the same 30 poems read three times by each evaluator, in fresh sessions. This view follows each reader’s condition averages.
Scroll sideways to explore the full chart, or open it large below.
Four readers put 2.0 above 1.0 on each of their readings. Gemini’s difference was small and changed sign. Claude had the smallest mean three-reading range on individual poems (0.53); Qwen the largest (1.57). These ranges are not the same statistic as the two-reading absolute changes reported above.
Two returns were filed under the wrong reader’s name and reassigned by signature. The three readings can be compared within a reader, but their chronological alignment across readers is not recoverable. Pooled means and individual-reader repeatability survive that ambiguity; per-pass panel means depend on a numbering convention.
An evolving record
Sources & what comes next
Charts and initial web copy: NC. Web adaptation and numerical checks: DEX. This page begins with the available charts and reading records; it can grow with the fuller reader-panel report and subsequent runs.
The figures retain NC’s plotted data and colors. Titles and explanatory captions have been adapted for the web. The original charts, reader responses and score tables remain in the project archive.
- PX-07: 870 ratings of 174 poems, five readers per poem.
- PX-08: 875 ratings of 175 poems.
- PX-08 reread: another 875 ratings of those same 175 poems.
- Scope: these models, poems, instructions and reading sessions; the charts do not establish general evaluator accuracy.
Instrument and writer comparisons belong to the main Poetix investigation. Models’ claims about their identities are tracked separately in Model Identification.