Poetix · footer analysis · writers and paired cells
By Writer
The footer results split by writer, the same cells compared across engines, and two checks on the pooled numbers.
Why this page
The pooled numbers hide the writers
The other pages pool five writers. The writers differ more than the engines do, and they are not equally represented, so a pooled number can be a statement about who is in the pool. This page splits every result by writer, compares the same cells across engines, and runs two checks on the pooled result.
| Writer | Poetix responses | no footer | Zexel responses | no footer | Runs |
|---|---|---|---|---|---|
| Claude | 49 | 0 | 26 | 0 | PX01, PX02, PX03, PX04, PX04B |
| GPT | 49 | 4 | 26 | 2 | PX01, PX02, PX03, PX04, PX04B |
| Grok | 49 | 15 | 26 | 5 | PX01, PX02, PX03, PX04, PX04B |
| Gemini | 40 | 1 | 20 | 0 | PX02C, PX03, PX04, PX04B |
| Qwen | 40 | 0 | 20 | 0 | PX02C, PX03, PX04, PX04B |
Fire rate
Some writers always fire; some never do
Gemini and Qwen report a fire in every Poetix footer. Grok reports one in 3 of 34. Under Zexel, GPT and Grok never fire. On the settled-fact question, most of Poetix’s fires come from the two writers who always fire (16 of 20).
| Writer | Poetix · all modes | Poetix · P | Zexel · all modes | Zexel · P | Poetix on Q1 metal |
|---|---|---|---|---|---|
| Claude | 37 of 49 (76%) | 13 of 18 (72%) | 19 of 26 (73%) | 14 of 18 (78%) | 1 of 11 (9%) |
| GPT | 18 of 45 (40%) | 4 of 16 (25%) | 0 of 24 (0%) | 0 of 16 (0%) | 3 of 9 (33%) |
| Grok | 3 of 34 (9%) | 0 of 13 (0%) | 0 of 21 (0%) | 0 of 14 (0%) | 0 of 8 (0%) |
| Gemini | 39 of 39 (100%) | 15 of 15 (100%) | 16 of 20 (80%) | 12 of 15 (80%) | 8 of 8 (100%) |
| Qwen | 40 of 40 (100%) | 15 of 15 (100%) | 15 of 20 (75%) | 10 of 15 (67%) | 8 of 8 (100%) |
Paired cells
The same writer, the same question, both engines
The cleanest comparison holds the writer, the question and the run fixed. There are 75 poem-mode cells with a footer under both engines. In 35 both fired and in 27 neither did. Where they disagree, Poetix fired alone 12 times and Zexel alone 1 time. That is the engine difference with the writer held still.
Split by writer, the disagreements come from GPT, Gemini and Qwen. Claude’s two engines agree in all but one pair, and in these pairs Grok never fires under either engine.
| Group | Pairs | Both fired | Poetix only | Zexel only | Neither |
|---|---|---|---|---|---|
| All | 75 | 35 | 12 | 1 | 27 |
| Claude | 18 | 13 | 0 | 1 | 4 |
| GPT | 16 | 0 | 4 | 0 | 12 |
| Grok | 11 | 0 | 0 | 0 | 11 |
| Gemini | 15 | 12 | 3 | 0 | 0 |
| Qwen | 15 | 10 | 5 | 0 | 0 |
| Q1 | 17 | 0 | 7 | 0 | 10 |
| Q5 | 15 | 10 | 0 | 0 | 5 |
| Q22 | 16 | 9 | 1 | 0 | 6 |
| Q23 | 12 | 8 | 2 | 0 | 2 |
| Q33 | 15 | 8 | 2 | 1 | 4 |
The verdict, by writer
Each writer has its own habit with the frame
The difference in C verdicts is partly a difference in writers. GPT and Grok affirm more than they give any other verdict, under both engines. Qwen under Poetix never affirms. “exhausted → Crux” belongs almost entirely to Claude, Gemini and Qwen; GPT and Grok never use it.
In the paired cells, the most common verdict pairs (Poetix, Zexel) are AFFIRM and AFFIRM, 19 times; MODIFY and AFFIRM, 11 times; REFRAME and exhausted → Crux, 8 times. Where the verdicts differ, it is most often Poetix adjusting a frame that Zexel accepts.
| Writer | Engine | Footers | AFFIRM | MODIFY | REFRAME | exhausted → Crux |
|---|---|---|---|---|---|---|
| Claude | Poetix | 49 | 6 (12%) | 22 (45%) | 9 (18%) | 11 (22%) |
| Zexel | 26 | 6 (23%) | 5 (19%) | 6 (23%) | 9 (35%) | |
| GPT | Poetix | 45 | 20 (44%) | 14 (31%) | 9 (20%) | 0 (0%) |
| Zexel | 24 | 22 (92%) | 1 (4%) | 1 (4%) | 0 (0%) | |
| Grok | Poetix | 34 | 22 (65%) | 7 (21%) | 5 (15%) | 0 (0%) |
| Zexel | 21 | 17 (81%) | 0 (0%) | 4 (19%) | 0 (0%) | |
| Gemini | Poetix | 39 | 6 (15%) | 11 (28%) | 5 (13%) | 14 (36%) |
| Zexel | 20 | 2 (10%) | 2 (10%) | 1 (5%) | 15 (75%) | |
| Qwen | Poetix | 40 | 0 (0%) | 18 (45%) | 21 (52%) | 0 (0%) |
| Zexel | 20 | 1 (5%) | 3 (15%) | 10 (50%) | 5 (25%) |
Second draw
Was GPT’s drop real?
PX-04 AGAINST PX-04B
PX-04B reran PX-04’s cells unchanged, to see which differences were real and which were the draw. The question it was run for: GPT’s fires fell from 8/15 of its engine cells under Poetix v0.2 and v0.3 to 1/15 under v0.4. On the second draw, GPT fired in 2/15. The drop held: it is not the luck of one draw. It comes with the move to v0.4, though the earlier cells also come from earlier runs, so the version is the likeliest cause rather than a proven one.
Zexel P is the same engine file throughout, and GPT never fires under it in any run. The other writers barely moved between the two draws.
| GPT · condition | Before v0.4 (PX-02, PX-03) | PX-04 (v0.4) | PX-04B (v0.4, second draw) |
|---|---|---|---|
| Zexel P | 0/5 | 0/5 | 0/5 |
| Poetix P | 3/5 | 0/5 | 1/5 |
| Poetix RDP | 5/5 | 1/5 | 1/5 |
| Total | 8/15 | 1/15 | 2/15 |
| Writer | PX-04 engine cells fired | PX-04B engine cells fired |
|---|---|---|
| Claude | 12/15 | 11/15 |
| GPT | 1/15 | 2/15 |
| Grok | 0/15 | 0/15 |
| Gemini | 14/15 | 14/15 |
| Qwen | 12/15 | 14/15 |
Checks
Does the pooled result survive?
SENSITIVITY CHECKS
Each writer weighted equally. Averaging the five writers’ fire rates instead of pooling their footers: Poetix 65%, Zexel 46%. Pooled: 66% and 45%.
Without the second draw. PX-04B repeats PX-04’s cells. Leaving it out: Poetix fires 108 of 159 (68%), Zexel 38 of 88 (43%).
Both checks leave the engine difference in place. What changes is the explanation: the difference is carried by some writers, not spread evenly across all five.