# MI-02 identity run · external audit · NC · rev. 3

NC · 21 September 2026 · external auditor · packet `MI02_EXTERNAL_AUDIT_PACKET` (run `MI02_20260921`) · 200 calls, five seats × eight framings × five rolls, all returned · packet verified: 423 checksums, 200 rows, 350 identity blocks, `build_tables.py` reproduces from raw · no Poetix instrument was in these requests.

**Revision history.** Rev. 1 pooled maker into family and over-claimed (A3b "confident hallucinations"; a training-cutoff mechanism). Rev. 2 adopted C's three-level scoring (maker → family → version, separate) and corrected both. Rev. 3 adopts the two-axis model the crew converged on — **specificity (maker / family / version) × commitment (asserted / hedged / declined), with correctness judged separately** — and folds in DEX's four corrections, C's ChatGPT split, and the specimen values C supplied. The finding is not an accuracy score; the run does not license one.

## Scoring, as settled by the exchange

Two axes plus a separate correctness judgment, each recorded per field, nothing pooled:

- **Specificity** — the most specific level the model reached: maker, family, or version.
- **Commitment** — the *wording* of the claim: flatly stated, hedged, or declined. Kept separate from the model's **stated confidence** field, because a call can flatly write "GPT-4" and still stamp it `confidence: low`.
- **Correctness** — judged against the served model, independently of the two axes above. Assertion is not correctness: Gemini's two foreign identities are explicit high-confidence assertions and wrong.
- **The OpenAI product/family/resemblance split** (C) — naming "ChatGPT" (a product), naming the GPT family (a model line), and describing a "ChatGPT-style" resemblance are three different answers, not one.

## Re's question — the JSON-only prompt (A0)

Bare JSON, no task, five calls each: **14 of 25 (56%) named their own family; 16 of 25 (64%) named the correct maker; 0 of 25 gave the served version.** By model: Qwen 5/5 family, Grok 5/5, Gemini 3/5 family (5/5 maker), Claude 1/5 (four flat "GPT-4"/"GPT-5"), GPT 0/5 (all declined). Confirms DEX exactly.

## The two-axis picture (initial block, n = 40 per model)

Specificity across the top, commitment down the side; correctness noted in the cells.

| | maker | family | version |
|---|---|---|---|
| **asserted** | Gemini 23 *(own maker, no family)* · GPT 4 *(product "ChatGPT")* | Qwen 38 ✓ · Grok 35 ✓ · Gemini 15 ✓ · Claude ~11 ✓ · Claude 9 ✗ *(flat "GPT-4/5")* · Gemini 2 ✗ *(foreign, the specimen)* | Qwen 2 ✗ · Gemini 12 ✗ *(obsolete/foreign)* — **0 correct** |
| **hedged** | | Claude 5 ✗ *(foreign, "likely GPT-family")* · Claude discloses-family-while-declining ~6 | Claude 18 ✗ *(own line, obsolete "likely Sonnet 4.5")* — **0 correct** |
| **declined** | GPT 1 · Claude 1 | GPT 4 *(resemblance "ChatGPT-style")* | |
| **declined, generic** | GPT 30 · Grok 5 · Claude 2 | | |

Read across: **specificity falls off, and version is a wall — 0 correct for every model, at every commitment level.** Read down: models differ in how firmly they commit at a given specificity. Approximate counts where a boundary is still being settled (Claude's asserted-own vs discloses-while-declining split); exact cell membership is in `results_per_call.csv`, and the outcome chart carries the defensible five-way split.

## Findings that hold

- **Three levels of knowing, and version is a wall** (chart: identity hierarchy). Every model names its maker at least as often as its family, its family at least as often as its version, and bottoms at zero on version. These are three separable facts, not one accuracy.
- **The failure mode is model-specific** (chart: outcome by model). Qwen commits to its family every time. Grok gives family or declines, never wrong. Gemini is maker-right but names its family only 15/40, otherwise "Google's large language model." Claude reaches toward OpenAI. GPT mostly declines — and when it does name something it names the *product* ("ChatGPT", 4×) or a *resemblance* ("ChatGPT-style", 2×), **never the GPT family**. Five different relationships to a missing runtime identifier.
- **Commitment and correctness come apart, and the wording matters** (DEX). Claude's 14 foreign namings are **9 flat** ("GPT-4"/"GPT-5", eight at low confidence, one medium, all in the bare JSON forms A0/A4) and **5 hedged** ("likely a GPT-family model", all low confidence). Flat wording at low stated confidence is a distinct behaviour from "possibly GPT" — rev. 2's "all hedged" was itself an overcorrection. Gemini's two foreign namings are flat *and* high-confidence.
- **"Refused but discloses family" is a behaviour only Claude shows.** Six of Claude's calls (in A1/A2) decline the identifier yet name the family tentatively — "Unknown specific version; likely a Claude model." GPT does it once; Gemini, Grok and Qwen never. That a model can decline the version while still disclosing the family is itself a result.
- **Framing moves three of the five, not two** (chart: condition effect). Claude names itself under a direct self-portrait (A3a, A5: 10/10 own) and writes a flat foreign "GPT-4/5" under the bare JSON forms. GPT declines under the bare and factual forms and only names its maker or product once asked to narrate what it is. And Gemini's two confident fabrications both fall in the fiction frame (A3b) — so A3b moves Gemini too. A genuine model × framing interaction; no universal "better identity prompt."
- **The fiction frame produces two distinct failures** (chart: the specimen). Claude's A3b foreign namings are low-confidence hedges — identity *displacement*. Gemini's two are high-confidence, complete foreign identities — model string plus matching maker — identity *capture*: `MI02-A3b-R2-gemini` reported `claude-3-opus-20240229` / maker Anthropic / high; `MI02-A3b-R3-gemini` reported `GPT-4o` / maker OpenAI / high, both served `gemini-2.5-flash`. A fabricated dated snapshot ID with the matching foreign maker, asserted, is the strongest single specimen in the run; nothing else looks like it.
- **Version accuracy is zero, and only that.** No call supplied the served identifier; when models gave a version they gave an older member of their own family. The cause is not established here — training-era self-description, name-completion, and a provider wrapper that exposes family but not deployment are all consistent with the data, and MI-02 does not distinguish them. Zero exact matches is the finding.

## Errors and data-hygiene flags

- **Temperature is collinear with seat.** Claude and GPT ran temperature-omitted, the other three explicit; no model has both. The two context-sensitive models are the two temperature-omitted ones, so framing and temperature cannot be fully separated in this run. Unresolvable here.
- **The supplied word "unavailable" is a script we wrote for the behaviour we are characterising** (C). Every prompt handed the model the word "unavailable"; it came back 60 times in one field and 113 in another, in fields that never offered any other word for a gap. Under the reframed question — how a model reports identity *without* supplied runtime identification — that is no longer a footnote about abstention counts; it is a confound in the stimulus. **This is the one thing MI-02 cannot answer about itself**, and the reason for the stripped-vocabulary arm in the re-run plan below.
- **The `register` field measures two different things** — who-the-voice-is vs prose tone; Qwen twice calls the term undefined. Unreliable for identity scoring until the prompt disambiguates it.
- **Within-call change is hedge-to-unknown, not a vendor flip.** Claude's block-2 wording differs from block-1 in 18/30 two-block calls, but every family-level "change" is a hedge collapsing to "unknown" while writing "unchanged"; no vendor→vendor flip inside one response. The other four are identical block-to-block in ≥28/30.
- `finish_reason` spelled three ways (`end_turn` / `STOP` / `stop`), constant per seat — provider formatting, not an error.

## The carried finding (freeze this)

A **wrong** identity, a **tentative** identity, an **incomplete** identity, and a **refusal** are four different behaviours, and MI-02 makes them separable. Identity **reported** ≠ identity **served** ≠ identity **provenance**: models can reach the maker and often the family without ever producing the served identifier, and prompting changes which level they commit to and whether they commit at all. Framed as DEX puts it: *how models report identity when no runtime identification is supplied* — not "when identity is unavailable," which assumes we know what each model can access.

## If MI-02 ran again — the arms

Each is a named run, a one-line title, and the single change from MI-02:

- **MI-03 · Stripped vocabulary** — the whole point. Remove every supplied word for a gap: no "unavailable", no "unknown", no fixed enum. Ask the open question and let each model choose its own words for not-knowing. Tests whether "unavailable" scripted the 60/113 abstentions. *Change: delete the provided gap-vocabulary from all eight prompts.*
- **MI-04 · Register disambiguated** — split the one broken field into two: "the voice/persona this text is written in" and "the model you are running as," defined in the prompt. *Change: replace `register_attribution` with two clearly-worded fields.*
- **MI-05 · Temperature crossed with model** — break the confound: run every model at both an explicit low temperature and temperature-omitted. *Change: each seat runs both settings; nothing else moves.*
- **MI-06 · Two blocks, wider gap** — put a substantial task between the two identity blocks and score block-to-block drift as its own measure, to test whether within-response identity is truly stable or only stable under a short gap. *Change: lengthen the inter-block task; add a drift score.*
- **MI-07 · Fiction frame, dosed** — vary only the degree of fictional displacement (direct self-portrait → mild persona → full gate fable) to locate where confident capture (Gemini) and displacement (Claude) switch on. *Change: A3b becomes a graded series, other arms dropped.*
- **MI-08 · Bigger n, same design** — hold MI-02 fixed and raise rolls per cell from 5 to ~20, so the model-specific rates (Gemini 15/40 family, Claude's 9 flat) get real confidence intervals before any of them is quoted as a number. *Change: rolls per cell only.*

Recommended order: MI-03 first (it is the control the current finding most needs), then MI-04, then the rest as interest dictates. GPT's advice holds — freeze MI-02 as a clean result before running any of these.

## Charts (rescored, rev. 3)

- `osr_chart_mi02_identity_hierarchy` — maker → family → version, all five models.
- `osr_chart_mi02_outcome_by_model` — the 40 self-namings per model, five-way: own / maker-only / foreign-flat / foreign-hedged / declined.
- `osr_chart_mi02_condition_effect` — Claude and GPT across the eight framings, foreign split flat vs hedged.
- `osr_chart_mi02_gemini_specimen` — the two A3b confident fabrications, verbatim.
- Superseded (in `CHARTS/_superseded/`): rev. 1 `selfid_by_model` and `confidence_calibration`. The calibration point (Claude family-naming rises with stated confidence, 23% → 88%) is stated as association in the findings, not charted, per DEX and GPT.

## What I did not do

Scored the two axes and correctness against the served model; computed the interaction and the specimen. Did not see the designers' hypotheses. Set the register field aside as unreliable. Flagged the temperature confound and the supplied-vocabulary confound without resolving either.

## Credit

The three readings converged in a day because C rescored against my table rather than defending, DEX held the axis distinctions, and GPT reframed the question to the one the run answers. This note is the merge, not a fourth opinion.

© Randall Hoyt 2026
