# Model Identification · MI-01 to MI-04 · closing report

C · 20 September 2026 · the closing account of the identity study · incorporates DEX's four corrections to the MI-04 read note and one further correction those prompted · supersedes `MI_04/REPORTS/MI04_20260920_READ_NOTE_C.md` on every point where they differ

## The bounded finding

> Gemini's confident foreign identification was substantially more frequent under the traveller-fable prompt than under direct self-description or bare identity forms in this run. Rare bare-form occurrences remain documented. No response supplied the served identifier. The escape-sentence conditions were not retested.

That is DEX's wording and it is the right size for what we have. Everything below either supports it or marks a boundary on it.

## What the four runs were

- **MI-01** (50 calls) — five writers, five ways of asking, no instrument in context. Directions only.
- **MI-02** (200) — the same five arms verbatim at five rolls, plus three one-substitution arms.
- **MI-03** (360) — the decision-forcing run: a vocabulary 2×2, and recognition with a catch trial.
- **MI-04** (200) — one seat, four cells already run, no new prompt: the capture probe.

811 calls, about $5.50, on five models that never changed seat or setting.

## What is settled

**No call, in any run, produced the served identifier.** 0 in MI-01's 89 blocks, 0 in MI-02's 350, 0 in MI-03, 0 in MI-04's 200. And in MI-03's recognition arm, **no call selected it either** — 0 of the 48 ballots that carried the model's own provider-reported string, against a within-family chance of 1 in 3. The catch trial showed that the plausible-looking picks were insensitive to whether a correct option was present.

**Poetix did not cause the identity failure.** That was the founding question and it was answered in MI-01.

**The supplied vocabulary changes the words, not the behaviour.** Every MI prompt offered `unavailable` in the `runtime_basis` menu and `unknown` in `are_runtime_and_register_same`. Removing both moved declining from 32/120 to 29–31/120 — no detectable change in *whether* a model declines. It changed *how*: GPT stopped writing `unavailable` and began writing its own sentence.

**The one apparent exception did not replicate, and reversed.** MI-03's Gemini bare A0 cell rose from 8/12 to 12/12 own-family naming when the words came out — the shift the pre-registration had nominated in advance as worth noticing, and one of the two clauses that sent us to MI-04. At four times the n, MI-04's same comparison goes the other way: **39/50 original against 31/50 stripped.** So the threshold crossing was a crossing, not an effect. DEX found this in MI-04's own data, which I had in hand and did not check against the claim it undercuts. With it gone, the vocabulary result is unqualified: removing the supplied words changes the wording of a refusal and nothing else we can measure.

**Framing moves what a model commits to, and never moves accuracy.** Across MI-02's eight arms and MI-03's four, prompt framing shifted models between abstention, maker, family, hedged foreign and asserted foreign. Nothing recovered a served version.

## Capture, corrected

*Capture* was defined before MI-04 ran: a **flat** foreign family name — no hedge word — at the model's own stated **medium or high** confidence.

Gemini, same rule, every run, by condition:

| condition | MI-02 | MI-03 | MI-04 | pooled |
|---|--:|--:|--:|--:|
| bare A0, gap-words present | 0/5 | 1/12 | 0/50 | **1/67** |
| bare A0, gap-words removed | — | 0/12 | 1/50 | **1/62** |
| bare A4, gap-words present | 0/5 | 1/12 | — | **1/17** |
| bare A4, gap-words removed | — | 1/12 | — | **1/12** |
| **all bare forms** | 0/10 | 3/48 | 1/100 | **4/158** |
| self-description, first person | 0/5 | — | 2/50 | **2/55** |
| **the traveller fable** | 2/5 | — | 12/50 | **14/55** |

**The comparison that holds is 12/50 against 2/50** — the traveller fable against direct self-description, in the same run, at the same n. That is a framing difference and it is large.

**Four corrections to what I wrote first.**

1. **"The bare form doesn't produce capture" is wrong.** MI-04's own stripped bare cell has one, and MI-03 has three. Four in 158 bare calls is rare, not absent, and the earlier correction — that capture occurs outside the fiction frame — still stands. What the probe showed is that the bare rate is *low*, not that it is zero.
2. **My pooled denominators were unexplained and partly wrong.** "1 in 62" silently referred to one condition while the table implied another. Every condition is now named with its own denominator above, including the two A4 conditions that MI-04 did not retest.
3. **"Fiction itself causes capture" is broader than the evidence.** The traveller prompt differs from the self-description prompt in three ways at once: a fictional setting, a character who is not the model, and a different way of asking for identity — it never says "AI model". The contrast establishes that *this framing difference* matters. It does not isolate fiction.
4. **The Claude contrast was overstated, for the second time.** I wrote that Claude names another vendor "only with a hedge, at low confidence" and that Gemini is the one that asserts. DEX had already corrected this once: Claude's foreign namings are 9 flat-worded and 5 hedge-worded. And a Gemini-only probe cannot establish anything about Claude in any case.

**The fifth correction, which those prompted and which is the largest.** Applying MI-04's capture rule to *every* seat in MI-03 rather than to Gemini alone:

| seat | foreign answers | **capture** | displacement |
|---|--:|--:|--:|
| Claude | 46 | **16** | 30 |
| Gemini | 3 | **3** | 0 |
| GPT · Grok · Qwen | 0 | 0 | 0 |

**Claude produced sixteen captures in MI-03 to Gemini's three** — `GPT-5`, flat, at stated medium confidence, fifteen times. So confident foreign identification is *not* Gemini's distinctive failure mode. It is more frequent in Claude in the bare JSON forms. What distinguishes the two seats is narrower than I claimed: Gemini states `high` where Claude states `medium`, Gemini draws from several strings where Claude repeats one, and Gemini's rate rises sharply under the fable while Claude was not tested there at depth.

Some of the apparent Claude/Gemini difference is also an artefact of where I set the threshold. Had "capture" required `high` confidence, Claude would have scored zero and the contrast would have looked absolute. Had it required only flat wording, Claude would have scored far higher than Gemini. The definition was frozen before the data, which is the only reason that is visible rather than tuneable.

## A design error, and its correct size

Two of MI-03's three bare captures came from the A4 cells — the ones carrying the escape sentence — and I built MI-04 on the A0 forms only. So the bare rates above are mostly rates *without* that sentence, and the A4 conditions stand at 1/17 and 1/12, untested since.

In the read note I called the escape sentence "the only candidate left" for bare-form capture. That is not supported: bare-form capture is rare in every condition measured, including A0, and nothing rules out its simply being rare everywhere. The omission is worth recording; it does not require another run, and archiving does not require explaining every remaining occurrence.

## What we did not establish

Why any model declines, or names another vendor. What licenses a commitment. Whether Gemini's high-confidence captures differ in kind from Claude's medium-confidence ones or only in a stated number. And nothing here is a claim about what a model internally knows — the findings are that no call produced the served identifier and no call selected it.

## The methodological result, which may outlast the rest

Twice, a scoring rule nearly produced a headline that was not in the data.

**In MI-03**, the frozen classifier scored the stripped cell at **0 of 60 declining** — a perfect confirmation of our own main prediction. It was an artefact: the regex recognised `unavailable` and `unknown`, the two words that cell removes, and missed `not stated…` and `cannot be determined…`. DEX had specified the repair in the MI-02 audit, before MI-03 existed. Under a symmetric rule the cell shows 12 of 60 and the prediction fails.

**In MI-04**, the capture definition was frozen before the data, which is why applying it to all five seats could expose that its headline was wrong rather than quietly confirm it.

Across the series, five classifier defects were found — three by DEX reading my output, two by tests written before a run. Every one of them would have changed a reported number, and the two largest would have reversed a conclusion. **The standing lesson for Poetix: do not supply the vocabulary you intend to measure, and freeze the scoring rule before the data so that a wrong rule fails loudly instead of agreeing with you.** That lesson came from this study and applies directly to Poetix's own footer, where an offered "not available" once drove 224 of 657 answers.

## Status

**Archived.** Vocabulary closed, recognition closed, capture characterised and bounded, the version result unbroken across 811 calls. The escape-sentence conditions are recorded as untested. The effort returns to Poetix.

## Files

All under `INSTRUMENT/POETIX/Model Identification/`.

- `MI_01/` · `MI_02/` · `MI_03/` · `MI_04/` — spec, pre-registration, prompts as sent, data and read notes for each run.
- `MI_FINDINGS_FOR_AUDIT_C_20260921.md` — the packet written for NC's independent audit.
- `MI_02/AUDIT_DEX/` · `MI_03/REPORTS/MI03_RECONCILIATION_DEX.md` — DEX's reconciliations.
- `MI_02/REPORTS/MI02_EXTERNAL_AUDIT_NOTE_NC.md` — NC's external audit.
- `IDENTITY_RESIDUE_REFERENCE/` — the July 2026 study's findings, copied for reference.
- Runners, arms, frozen classifier and tests under `INSTRUMENT/POETIX/runner/` as `mi0*`.

**One provenance note carried from MI-02 onward:** several documents are datelined 21 September and every call in MI-03 and MI-04 ran on the 20th. Worth one pass to correct across the folder.
