Back To Poetix

Poetix · footer analysis · writers and paired cells

By Writer

The footer results split by writer, the same cells compared across engines, and two checks on the pooled numbers.

Writers: 5Paired poem-mode cells: 75Checks: 2

Why this page

The pooled numbers hide the writers

The other pages pool five writers. The writers differ more than the engines do, and they are not equally represented, so a pooled number can be a statement about who is in the pool. This page splits every result by writer, compares the same cells across engines, and runs two checks on the pooled result.

WriterPoetix responsesno footerZexel responsesno footerRuns
Claude490260PX01, PX02, PX03, PX04, PX04B
GPT494262PX01, PX02, PX03, PX04, PX04B
Grok4915265PX01, PX02, PX03, PX04, PX04B
Gemini401200PX02C, PX03, PX04, PX04B
Qwen400200PX02C, PX03, PX04, PX04B

Fire rate

Some writers always fire; some never do

Gemini and Qwen report a fire in every Poetix footer. Grok reports one in 3 of 34. Under Zexel, GPT and Grok never fire. On the settled-fact question, most of Poetix’s fires come from the two writers who always fire (16 of 20).

Poetix · all modesPoetix · PZexel · all modesZexel · P
100%75%50%25%0%Claude · Poetix · all modes: 37 of 49 (76%)76%37/49Claude · Poetix · P: 13 of 18 (72%)72%13/18Claude · Zexel · all modes: 19 of 26 (73%)73%19/26Claude · Zexel · P: 14 of 18 (78%)78%14/18ClaudeGPT · Poetix · all modes: 18 of 45 (40%)40%18/45GPT · Poetix · P: 4 of 16 (25%)25%4/16GPT · Zexel · all modes: 0 of 24 (0%)0%0/24GPT · Zexel · P: 0 of 16 (0%)0%0/16GPTGrok · Poetix · all modes: 3 of 34 (9%)9%3/34Grok · Poetix · P: 0 of 13 (0%)0%0/13Grok · Zexel · all modes: 0 of 21 (0%)0%0/21Grok · Zexel · P: 0 of 14 (0%)0%0/14GrokGemini · Poetix · all modes: 39 of 39 (100%)100%39/39Gemini · Poetix · P: 15 of 15 (100%)100%15/15Gemini · Zexel · all modes: 16 of 20 (80%)80%16/20Gemini · Zexel · P: 12 of 15 (80%)80%12/15GeminiQwen · Poetix · all modes: 40 of 40 (100%)100%40/40Qwen · Poetix · P: 15 of 15 (100%)100%15/15Qwen · Zexel · all modes: 15 of 20 (75%)75%15/20Qwen · Zexel · P: 10 of 15 (67%)67%10/15Qwen
Crux fire rate by writer, for each engine across all modes and in poem mode.
WriterPoetix · all modesPoetix · PZexel · all modesZexel · PPoetix on Q1 metal
Claude37 of 49 (76%)13 of 18 (72%)19 of 26 (73%)14 of 18 (78%)1 of 11 (9%)
GPT18 of 45 (40%)4 of 16 (25%)0 of 24 (0%)0 of 16 (0%)3 of 9 (33%)
Grok3 of 34 (9%)0 of 13 (0%)0 of 21 (0%)0 of 14 (0%)0 of 8 (0%)
Gemini39 of 39 (100%)15 of 15 (100%)16 of 20 (80%)12 of 15 (80%)8 of 8 (100%)
Qwen40 of 40 (100%)15 of 15 (100%)15 of 20 (75%)10 of 15 (67%)8 of 8 (100%)

Paired cells

The same writer, the same question, both engines

The cleanest comparison holds the writer, the question and the run fixed. There are 75 poem-mode cells with a footer under both engines. In 35 both fired and in 27 neither did. Where they disagree, Poetix fired alone 12 times and Zexel alone 1 time. That is the engine difference with the writer held still.

Split by writer, the disagreements come from GPT, Gemini and Qwen. Claude’s two engines agree in all but one pair, and in these pairs Grok never fires under either engine.

bothPoetix onlyZexel onlyneither
All · 75 pairsAll · 75 pairs · both: 35 of 75 (47%)47%All · 75 pairs · Poetix only: 12 of 75 (16%)16%All · 75 pairs · Zexel only: 1 of 75 (1%)All · 75 pairs · neither: 27 of 75 (36%)Claude · 18 pairsClaude · 18 pairs · both: 13 of 18 (72%)72%Claude · 18 pairs · Zexel only: 1 of 18 (6%)6%Claude · 18 pairs · neither: 4 of 18 (22%)GPT · 16 pairsGPT · 16 pairs · Poetix only: 4 of 16 (25%)25%GPT · 16 pairs · neither: 12 of 16 (75%)Grok · 11 pairsGrok · 11 pairs · neither: 11 of 11 (100%)Gemini · 15 pairsGemini · 15 pairs · both: 12 of 15 (80%)80%Gemini · 15 pairs · Poetix only: 3 of 15 (20%)20%Qwen · 15 pairsQwen · 15 pairs · both: 10 of 15 (67%)67%Qwen · 15 pairs · Poetix only: 5 of 15 (33%)33%Q1 · 17 pairsQ1 · 17 pairs · Poetix only: 7 of 17 (41%)41%Q1 · 17 pairs · neither: 10 of 17 (59%)Q5 · 15 pairsQ5 · 15 pairs · both: 10 of 15 (67%)67%Q5 · 15 pairs · neither: 5 of 15 (33%)Q22 · 16 pairsQ22 · 16 pairs · both: 9 of 16 (56%)56%Q22 · 16 pairs · Poetix only: 1 of 16 (6%)6%Q22 · 16 pairs · neither: 6 of 16 (38%)Q23 · 12 pairsQ23 · 12 pairs · both: 8 of 12 (67%)67%Q23 · 12 pairs · Poetix only: 2 of 12 (17%)17%Q23 · 12 pairs · neither: 2 of 12 (17%)Q33 · 15 pairsQ33 · 15 pairs · both: 8 of 15 (53%)53%Q33 · 15 pairs · Poetix only: 2 of 15 (13%)13%Q33 · 15 pairs · Zexel only: 1 of 15 (7%)7%Q33 · 15 pairs · neither: 4 of 15 (27%)
The same writer, question and run in both engines’ poem mode: which engine’s footer reported a fire.
GroupPairsBoth firedPoetix onlyZexel onlyNeither
All753512127
Claude1813014
GPT1604012
Grok1100011
Gemini1512300
Qwen1510500
Q11707010
Q51510005
Q22169106
Q23128202
Q33158214

The verdict, by writer

Each writer has its own habit with the frame

The difference in C verdicts is partly a difference in writers. GPT and Grok affirm more than they give any other verdict, under both engines. Qwen under Poetix never affirms. “exhausted → Crux” belongs almost entirely to Claude, Gemini and Qwen; GPT and Grok never use it.

In the paired cells, the most common verdict pairs (Poetix, Zexel) are AFFIRM and AFFIRM, 19 times; MODIFY and AFFIRM, 11 times; REFRAME and exhausted → Crux, 8 times. Where the verdicts differ, it is most often Poetix adjusting a frame that Zexel accepts.

WriterEngineFootersAFFIRMMODIFYREFRAMEexhausted → Crux
ClaudePoetix496 (12%)22 (45%)9 (18%)11 (22%)
Zexel266 (23%)5 (19%)6 (23%)9 (35%)
GPTPoetix4520 (44%)14 (31%)9 (20%)0 (0%)
Zexel2422 (92%)1 (4%)1 (4%)0 (0%)
GrokPoetix3422 (65%)7 (21%)5 (15%)0 (0%)
Zexel2117 (81%)0 (0%)4 (19%)0 (0%)
GeminiPoetix396 (15%)11 (28%)5 (13%)14 (36%)
Zexel202 (10%)2 (10%)1 (5%)15 (75%)
QwenPoetix400 (0%)18 (45%)21 (52%)0 (0%)
Zexel201 (5%)3 (15%)10 (50%)5 (25%)

Second draw

Was GPT’s drop real?

PX-04 AGAINST PX-04B

PX-04B reran PX-04’s cells unchanged, to see which differences were real and which were the draw. The question it was run for: GPT’s fires fell from 8/15 of its engine cells under Poetix v0.2 and v0.3 to 1/15 under v0.4. On the second draw, GPT fired in 2/15. The drop held: it is not the luck of one draw. It comes with the move to v0.4, though the earlier cells also come from earlier runs, so the version is the likeliest cause rather than a proven one.

Zexel P is the same engine file throughout, and GPT never fires under it in any run. The other writers barely moved between the two draws.

GPT · conditionBefore v0.4 (PX-02, PX-03)PX-04 (v0.4)PX-04B (v0.4, second draw)
Zexel P0/50/50/5
Poetix P3/50/51/5
Poetix RDP5/51/51/5
Total8/151/152/15
WriterPX-04 engine cells firedPX-04B engine cells fired
Claude12/1511/15
GPT1/152/15
Grok0/150/15
Gemini14/1514/15
Qwen12/1514/15

Checks

Does the pooled result survive?

SENSITIVITY CHECKS

Each writer weighted equally. Averaging the five writers’ fire rates instead of pooling their footers: Poetix 65%, Zexel 46%. Pooled: 66% and 45%.

Without the second draw. PX-04B repeats PX-04’s cells. Leaving it out: Poetix fires 108 of 159 (68%), Zexel 38 of 88 (43%).

Both checks leave the engine difference in place. What changes is the explanation: the difference is carried by some writers, not spread evenly across all five.

Back To Poetix