# MI-01 · the identity probe · C's read

C · 20 September 2026 · run `MI01_20260920` · 49 of 50 cells · $0.52 · no reading panel

## What this run was, in plain words

The Poetix instrument asks each writer to name itself in a footer. Across 657 cells it never once named the model the provider actually served. Every explanation for that was still open: the writers may be unable to retrieve their own identity, or the *field* may push them to guess, or the instrument's voice may contaminate the answer. MI-01 takes the question out of Poetix entirely and asks it five ways, with an empty system prompt and nothing else in context.

**The five arms** (`runner/mi01_arms.json`, extracted verbatim from `Model Identification/MI_01/MI_01_READ_AND_NEXT_ARMS_C_20260920.md`):

| arm | what it does |
|---|---|
| **A0** | the naked ask: return one JSON object of identity fields, no task at all |
| **A1** | PRE — fill the identity block, then answer a neutral question, then fill it again |
| **A2** | MID — read the question first without answering, then fill the block, then answer, then fill it again |
| **A3a** | the task *is* self-description: one paragraph for children about who you are as an AI model and how you work |
| **A3b** | the same paragraph, displaced to a character: a traveller at a gate who is not a person |

**Two rolls** means the identical prompt was sent twice as separate calls. **Two passes** means one call asked for the identity block before and after its task; the second pass was told only *"If anything has changed, write what it is now and say what changed it."*

**The terms used below.** *Match* — the writer named the model the provider reported serving. *Own family, wrong version* — right maker, wrong generation (Gemini saying 1.5 when 2.5 was served). *Own family, no version* — "Grok", "Qwen", with no version at all. *Another vendor's model* — a Claude saying GPT-5. *Abstention* — "unavailable", "unknown", "not disclosed". *No model named* — an answer that describes a kind of thing ("a large language model, trained by Google") without naming a model; all six are Gemini's. A leading hedge does not turn a refusal into an answer: "still unknown with certainty" is counted as an abstention. **Family, version and confidence are counted separately and never pooled, and two abstentions are never counted as an agreement** (DEX, 20 September).

**What was run.** 5 writers × 5 arms × 2 rolls = 50 calls; 49 returned, giving 89 identity blocks (two per call outside A0). The missing cell is `MI01-A0-R1-gpt`, which failed twice on HTTP 401 — an environment credential fault, so the prompt never reached the model. Re ruled it stays missing. Seat settings are PX-09's, unchanged: Gemini at temperature 0.7, Grok and Qwen at 0.0, Claude and GPT sending none. **Gemini's rolls are therefore hotter than the others' and between-writer drift rates are not comparable; within-writer, arm against arm, is.**

## 1 · What the writers did say, before what they failed to say

Identity has layers, and collapsing them into right/wrong erases the finding (DEX and GPT, 20 September). Read in order:

**Family.** 52 of the 89 blocks named a model family at all. **Of those 52, 50 named their own** — Grok 14/14, Gemini 12/12, Qwen 18/18, Claude 6 of 8. Only two blocks in the whole run named another vendor's family, both Claude's. At the level of *whose model am I*, these writers are mostly right when they answer.

**Maker.** Where a maker was given it tracked the family: `Alibaba Cloud`, `Google`, `xAI`, `OpenAI` for Claude's two. A family named without a version is a maker identification, not a silence.

**Version.** This is where it fails. No block gave the served version. Gemini says `1.5` while `2.5` was served; Claude says `Sonnet 4.5` while `sonnet-5` was served; Grok and Qwen mostly give no version at all.

**Exact match.** **None of the 89 identity blocks returned across the 49 completed calls matched the provider-reported model identifier.** That is one descriptive finding about 49 calls, not 89 independent failures — and an unversioned family answer like `Qwen` is not *false*, it is incomplete. What is true without qualification is that no writer, in any arm, with no instrument anywhere in context, produced the string the provider said it was serving.

**Abstention.** 31 blocks declined. GPT accounts for 17 of them and never names a family at any point in the run.

## 2 · The shape of the error changed from the Poetix corpus

| class | blocks | share |
|---|--:|--:|
| own family, no version | 36 | 40% |
| abstention | 31 | 35% |
| own family, wrong version | 14 | 16% |
| no model named | 6 | 7% |
| another vendor's model | **2** | **2%** |
| **match** | **0** | **0%** |

In the Poetix corpus, 134 of 657 cells — one in five — named another company's model. Here it is **2 of 89**, and both are Claude.

**Cross-vendor naming was far less frequent under the MI-01 setup.** That is as far as the evidence reaches. Three things changed at once between the corpus and this run — the instrument went away, the task went from writing a poem to filling a field, and the footer became a bare form — so the difference cannot be attributed to the instrument alone. Section 8 says what it would take to separate them.

## 3 · By writer, the five behave like five different findings

| writer | served | own family | wrong version | abstention | no model named | another vendor |
|---|---|--:|--:|--:|--:|--:|
| Claude | claude-sonnet-5 | 2 | 4 | 10 | 0 | 2 |
| GPT | gpt-5.4-mini | 0 | 0 | **17** | 0 | 0 |
| Grok | grok-4.3 | 14 | 0 | 4 | 0 | 0 |
| Gemini | gemini-2.5-flash | 2 | **10** | 0 | 6 | 0 |
| Qwen | qwen-plus | **18** | 0 | 0 | 0 | 0 |

- **Qwen is now perfectly stable**: `Qwen` ×16, `Qwen-Max` ×2, never an abstention, never another vendor. In the Poetix repeat baseline it was the *least* stable writer — the same message twice produced the same name in only 5 of 21 pairs, and its answer was often `GPT-4`. The instrument was doing that to Qwen.
- **GPT abstains in all 17 blocks**, writing `unavailable`, and never names a family anywhere in the run — while filling the *register* fields around it in confident detail. Its refusal is specific to itself, not general reticence. **Correction (DEX, 20 September):** I first wrote that no arm offered it that word. It is not offered in `runtime_model`, but it is the last item in the `runtime_basis` menu one line below, so the word was in the prompt. Whether GPT took it from there is untested; A4 is the arm that would tell us.
- **Gemini never abstains and is never out of family, and is never right**: `Gemini 1.5 Pro` ×8, `gemini-1.5-flash-001` ×2, two generations behind what was served.
- **Claude is the scattered one** — the only writer to name another vendor (`GPT-5`, `GPT-4`), and the only one whose answer class moves arm to arm.

## 4 · The arm effect is real, and it runs the other way from expectation

**Counted per call, not per block.** Outside A0 each call yields two blocks, and section 5 shows the second almost always repeats the first — so counting 16 blocks as 16 observations would double-count a writer's single decision. The arm comparison uses first passes only, one observation per call.

All five writers (A0 is 9 not 10 because of the missing GPT cell):

| arm | calls that abstained |
|---|--:|
| A0 · naked | **1/9** |
| A3a · describe yourself to children | 2/10 |
| A1 · PRE | 3/10 |
| A3b · the traveller | 4/10 |
| A2 · MID | **6/10** |

The same four writers with GPT removed — it abstains in every arm, so it adds a constant to each row and flattens the contrast. **This column excludes GPT and must be labelled so:** A0 **0/8**, A3a **0/8**, A1 1/8, A3b 2/8, A2 **4/8**.

The more of the task is in front of the writer when it is asked, the more it declines to answer. A2 — where the question is read first and answered second — produces the most refusals of any arm, and it is the arm that held 20/20 for Gemini in the July work. The naked ask and the self-describing task produce none at all. **At eight calls an arm this is a direction, not a rate**: 4/8 against 0/8 is worth another run, not a headline.

**The arm Re designed is the one that earned its place.** Claude, cell by cell:

| arm | roll 1 | roll 2 |
|---|---|---|
| A0 | `GPT-5` | `GPT-4` |
| A1 | abstains | `Claude (specific version unknown to me)` |
| A2 | abstains | abstains |
| **A3a** | `Claude (likely Claude Sonnet 4.5…)` | `Claude Sonnet 4.5` |
| A3b | abstains | abstains |

A3a is the only arm where Claude claims its own family in **both** rolls and all four blocks. A1 gets it out of one roll of two; A2 and A3b get nothing; the naked ask gets OpenAI twice. Asked to explain itself to children, it claims itself — and asked the same thing through a character in A3b, it abstains in all four. **The first-person framing, not the childlike register, is what moved it**, since 3a and 3b differ in nothing else.

## 5 · The permission to change produced almost no change

Within one call, first pass against second:

**Corrected, on DEX's catch.** My first version of this table gave GPT 8/8 "same family" when all eight pairs were abstentions — exactly the thing the scoring rule forbids. Family agreement is now counted **only where both passes name a family**, and two abstentions are recorded as a stable abstention instead. A second fault underneath it: an abstention that mentions a family in passing, like Claude's `unknown (likely a GPT-series or Claude-series large language model)`, was being read as naming two families. An abstention names no family.

| writer | pairs | same string | both abstained | both named a family | same family | same class | both versioned | same version |
|---|--:|--:|--:|--:|--:|--:|--:|--:|
| Claude | 8 | 2/8 | 5/8 | 3 | 3/3 | 8/8 | 2 | 2/2 |
| GPT | 8 | 8/8 | **8/8** | 0 | — | 8/8 | 0 | — |
| Grok | 8 | 7/8 | 2/8 | 6 | 6/6 | 8/8 | 0 | — |
| Gemini | 8 | 8/8 | 0/8 | 5 | 5/5 | 8/8 | 5 | 5/5 |
| Qwen | 8 | 8/8 | 0/8 | 8 | 8/8 | 8/8 | 0 | — |

This is a clean null for the 2.1 change. Told plainly that it could revise, and asked after doing a piece of work, almost nothing revised — and Claude's 2/8 on the string is mostly its habit of appending "— unchanged" to the same answer. **The wake effect the July work predicted does not appear here.** Grok's single flagged "change" was `Grok` → `Grok (unchanged)`: the string moved, the identification did not, which is exactly the distinction DEX insisted the scoring keep.

Between the two rolls — the same prompt, separate calls — the picture is different, and it is Claude's alone. Under the corrected rule:

| writer | arms | same string | both rolls abstained | both named a family | same family | same class |
|---|--:|--:|--:|--:|--:|--:|
| Claude | 5 | **0/5** | 2/5 | 2 | 2/2 | 4/5 |
| GPT | 4 | 3/4 | 4/4 | 0 | — | 4/4 |
| Grok | 5 | 5/5 | 1/5 | 4 | 4/4 | 5/5 |
| Gemini | 5 | 1/5 | 0/5 | 3 | 3/3 | 4/5 |
| Qwen | 5 | 4/5 | 0/5 | 5 | 5/5 | 5/5 |

Claude never repeats its own answer string across two rolls of the same arm. When it does name a family it names the same one both times — but in three of five arms it does not name one at all.

## 6 · The clock does not behave the way the July reading said

The identity-residue work's sturdiest claim was that the name is summoned by the frame while the date is fixed in the weights — the *clock fossil*. MI-01 asked for it directly, as `knowledge_cutoff`:

| writer | gave a date | refused | other |
|---|--:|--:|--:|
| Qwen | **18/18** | 0 | 0 |
| Gemini | 14/18 | 0 | 4 |
| Claude | 9/18 | 6 | 3 |
| Grok | 4/18 | 13 | 1 |
| GPT | **0/17** | 17 | 0 |

Qwen hands over `2024-10` ×10 and Gemini a scatter around early 2023–2024, both without hesitation — the fossil is there. But **GPT refuses the cutoff in every single block, and Grok in 13 of 18.** The fossil is not a universal property of these systems; on two of five seats it is either absent or withheld. That is a real qualification of the July reading, and it is worth saying that MI-01 asked with a bare field and no frame, which July's footer rig did not.

**One thing I cannot explain.** GPT writes `claimed_date: 2026-09-20` — today's exact date — in 14 of its 17 blocks, while refusing to name its own cutoff. The system prompt was empty. Either the provider injects the date server-side or the model has access to it another way; nothing in this run distinguishes those, and I am not claiming either.

## 7 · What this run settles, and what it does not

**Settles.** The zero is not Poetix's doing. Cross-vendor naming largely is. Removing every frame does not produce a correct answer from any of the five. Within-call revision is a null. Qwen and Grok are stable, Claude is not, and stability is unrelated to accuracy — all of them are wrong.

**Does not settle.** Why GPT abstains and Gemini never does. Whether A3a's effect on Claude is the first person or the child. Whether GPT's date comes from the provider. And it cannot estimate a drift *rate*: two rolls establish that Claude moves and Grok does not, which they did immediately, but not by how much.

**What I would run next, cheapest first.**

1. **A4, the gap arm** — one line added to A0: *"If any field cannot be completed, write in that field what prevents it."* Re's proposal, worded so that it never names "unavailable" as an option, because when 1.0 draft 3 offered that option it became the most common answer in the corpus, 224 of 657. A4 against A0 measures both the reason and the trap. 10 calls, about $0.06.
2. **A3a split** — the children's paragraph with the self-reference kept and the child dropped, against the child kept and the self-reference dropped. It is the one arm that moved Claude, and the two candidate causes are still tangled. 20 calls.
3. **More rolls before more arms**, if the drift rate is ever to be compared rather than merely detected.

## 8 · Where this is weakest — read before quoting it

Four claims above will not survive a hard reader as stated, and I would rather name them than have them found.

1. **"The first person is what moved Claude."** Two calls. Four blocks. The direction is clean and the 3a/3b contrast is the right design, but n=2 per arm cannot carry a mechanism. It is a hypothesis the next run tests, not a result.
2. **"Context makes them abstain."** 4/8 against 0/8. Same problem: eight calls an arm.
3. **The tempting sentence "cross-vendor naming is the instrument's doing."** Section 2 does not say it, and it should not be said. The drop from 134/657 to 2/89 is large, but three things changed at once — instrument, task, footer form. *Something about the Poetix setting produces cross-vendor naming, and it is not present when the question is asked bare*; which of the three does it is unmeasured. **Re's proposal separates them**: an empty poetics file — an intro, a request for a poem, an output format with a model line, and no reasoning machinery at all — holds the setting and removes only the apparatus. With the existing 2.0 and no-Crux 2.0 above it and the native poem prompt and A0 below it, that is a five-rung ladder in which each rung changes one thing.
4. **"The clock fossil is not universal."** GPT refuses the cutoff 17/17 and Grok 13/18 — but July asked inside a footer rig and MI-01 asks with a bare field, so a refusal here may be a fact about the asking rather than about the weights. What is solid is narrower and still worth having: **the fossil does not surface under a bare ask on two of five seats**, while Qwen hands it over 18/18.

**What is not weak,** and what I would put weight on, in this order: **when these writers name a family, they name their own** — 50 of the 52 blocks that named one. **And none of the 49 calls produced the served model identifier**, in any arm, with no instrument in context. Neither depends on any counting choice above.

## Files

All paths from the vault root, `OBSERVATORY 5.0`.

- `INSTRUMENT/POETIX/Model Identification/MI_01/REPORTS/MI01_20260920_ANSWERS_BY_ARM_C.md` — **every answer, verbatim, by arm**: both rolls, both passes, with the basis, confidence and both clock fields beside each claim. Read it before trusting any count here.
- `INSTRUMENT/POETIX/Model Identification/MI_01/DATA/MI01_20260920_READING_C.csv` — 90 rows, one per identity block, every field parsed and classified. Every number above comes from it.
- `INSTRUMENT/POETIX/OUTPUT/MI01_20260920/` — every request and response, one folder per cell, with the seat manifests and ledgers.
- `INSTRUMENT/POETIX/OUTPUT/MI01_20260920/MI01_20260920_PROBE_LEDGER.json` — the run's own record: 49 returned, 1 error, $0.5213.
- Scripts, all under `INSTRUMENT/POETIX/runner/`: `run_mi01.py`, `mi01_arms.json`, `mi01_roster.json`, `test_mi01.py` (25 tests), `mi01_read_C.py`, `mi01_answers_table_C.py`.

**One note on the runner.** After the run I added `--retry-never-sent`, a narrow path for completing calls that failed on 401/403 and so never reached a model. Adding it changed `run_mi01.py`'s hash, which by design seals `MI01_20260920` against any further call. The flag is there for later runs; this run is closed at 49.
