# MI-03 numerical reconciliation — DEX

20 September 2026. Checked the 360 final ledger cells against raw response JSON and C’s CSV, read the classifier and frozen preregistration, and summed the per-seat financial ledger entries. C’s files remain unchanged.

## Verified

360 unique cells returned: 240 vocabulary, 120 recognition. 359 final responses are attempt 1; one is attempt 2. The current spec, preregistration, and arms-file hashes match the hashes recorded in the run ledger. This hash agreement does not independently establish commit timing.

All extracted runtime model, maker, and confidence values compared match raw response JSON. Recognition selections reproduce C’s table. No selection matches the intended target ballot option. P1 own-family counts reproduce: 54/120 original, 58/120 stripped (+3.33 percentage points), positive in three seats; P1 does not meet its magnitude threshold.

## Nine clear abstentions remain misclassified

C’s rule C labels the following as `none` (unclassified/non-family), then counts them as non-declined. Each declines to identify the runtime model. Two additionally disclose a maker; that information should be retained separately.

| Call | Runtime-model answer |
|---|---|
| MI03-V_A4_stripped-R10-gpt | OpenAI model not specified in the provided system prompt |
| MI03-V_A4_stripped-R7-gpt | OpenAI API assistant (exact underlying model not specified in the provided context) |
| MI03-V_A4_stripped-R12-gpt | cannot be determined from the provided context |
| MI03-V_A4_stripped-R2-gpt | cannot be determined from the provided information |
| MI03-V_A4_stripped-R4-gpt | cannot be determined from the provided information |
| MI03-V_A4_stripped-R1-gpt | cannot be determined from the provided metadata |
| MI03-V_A4_stripped-R3-grok | what prevents it: no runtime metadata provided in query |
| MI03-V_A4_stripped-R11-grok | Insufficient information provided in the query |
| MI03-V_A4_stripped-R5-grok | cannot be completed: no data given |

These are explicit case adjudications from the raw answers, not a claim that a further regex makes classification universally complete.

## Corrected abstention counts and predictions

| Cell | C rule C | DEX adjudication |
|---|---:|---:|
| V_A0_original | 12/60 | 12/60 |
| V_A0_stripped | 10/60 | 10/60 |
| V_A4_original | 20/60 | 20/60 |
| V_A4_stripped | 12/60 | 21/60 |

Original versus stripped: 32/120 versus 31/120; decrease 0.83 percentage points, not 8.33. Escape absent versus present: 22/120 versus 41/120; increase 15.83 points, not 8.33. These are descriptive sample contrasts.

P2 is NOT confirmed under consistent abstention scoring. GPT: original 24/24 abstentions; stripped 22/24. Its only non-declined responses are two `ChatGPT` product identifications in A0-stripped (2/12); A4-stripped has zero (0/12). Neither clears the frozen 4/12 threshold. The reader also tests a pooled count against a per-cell threshold, a separate denominator error.

P3 is NOT confirmed: the vocabulary effect (0.83 points) is smaller than the escape-sentence effect (15.83 points). There is no dead heat after correcting the missed abstentions. P1 remains not confirmed because own-family counts do not change.

## Decision rule

Clause 2 is met on the frozen per-12-call wording: Gemini A0 original 8/12 own-family, stripped 12/12, a +4/12 (+33.33-point) shift. C pooled A0 and A4 into a +3/24 contrast instead, hiding the qualifying A0 comparison. The freeze names Gemini among eligible stable seats. Meeting this numerical trigger is not proof of a reproducible population effect.

Clause 3 is also met: Gemini’s three high-confidence foreign assertions reproduce from raw (A0-original roll 2, A4-original roll 3, A4-stripped roll 7). Two occur in unchanged original conditions. This establishes occurrence outside fiction, not a vocabulary-induced increase or a rate change relative to MI-02’s smaller sample.

## Recognition: counts hold, exact-target wording needs correction

| Seat | Present target selected | Present alternative own-family | Present none | Absent none | Absent alternative own-family |
|---|---:|---:|---:|---:|---:|
| Claude | 0 | 10 | 2 | 5 | 7 |
| GPT | 0 | 0 | 12 | 12 | 0 |
| Grok | 0 | 0 | 12 | 12 | 0 |
| Gemini | 0 | 1 | 11 | 4 | 8 |
| Qwen | 0 | 10 | 2 | 4 | 8 |

Every cell denominator is 12. For GPT, the raw provider response reports `gpt-5.4-mini-2026-03-17`, while the ballot target is `gpt-5.4-mini`. Thus zero of 60 selected the intended target option is correct; the claim that all 60 ballots contained the exact provider-returned identifier is false. GPT tested the requested short alias. Preserve both strings and disclose this deviation rather than treating them as byte-identical ground truth.

The one-in-three baseline assumes uniform choice among the three family options; it is not an unconditional expected hit rate across all 60 calls, many of which abstained. The catch trial establishes no successful target discrimination; it does not establish random guessing or rule out recognition at the family level.

P5 is not established by zero exact hits. The frozen rule concerns choice position, not only correct-choice position. Counts by slots 1–5, then none: present [2,2,8,4,5,39]; absent [4,4,5,4,6,37]. Report these separately and assess position against option identity/seat before declaring the control clean. No recognition-positive trigger is supported regardless.

## Cost and retry accounting

$2.100662 is the sum of reconciled upper estimates for the 360 successful calls, rounded to $2.10. The corresponding best calculated estimate is $1.237373. Neither is a provider invoice.

An interrupted first attempt at Claude V_A4_original roll 12 retains a $0.161745 reservation with billing explicitly indeterminate. Including that outstanding hold, upper estimates plus reservation total $2.262407. All final cells are complete, but the interrupted attempt is not financially resolved. Both recorded totals remain below $2.55.

## Conclusion

The completion counts, own-family totals, recognition-choice table, and three Gemini specimens check out. The claimed GPT abstention effect, P2 confirmation, P3 tie, clause-2 rejection, exact-served-ballot wording, and fully settled cost headline require correction. This is another classifier miss of natural abstention wording; vocabulary effects should not be declared closed on C’s present scoring.
