Maker, family, and version
The useful distinction is specificity: naming Google is not naming Gemini, and naming Gemini is not supplying its served version.
Poetix · Model Identification · 20 September 2026
A poetry footer opened a parallel investigation: when asked who they are, what do models actually report? Four runs sharpened the questions, exposed errors in our scoring, and left a bounded finding about prompt framing. They did not explain why the names appear.
Maker, product, family and version are different answers. A tentative answer, an explicit refusal and a confident foreign name are different behaviours. Combining them produced misleading accuracy figures.
No exact served identifier was recovered in the reported production tests. Recognition also failed to recover the target, with a model-alias limitation for GPT. Those observations do not establish what information a model internally knows.
The clearest final comparison: Gemini gave confident foreign identities in 12 of 50 traveller-fable calls, versus 2 of 50 direct self-descriptions. Rare bare-form occurrences remain in the record. This is a result about these prompts, not fiction in general.
The trajectory
| Run | Scope | What it contributed |
|---|---|---|
| MI-01 | 49 returned calls; 89 identity blocks | Bare identity requests also produced incomplete or incorrect reports. Poetix is not necessary for those outcomes. |
| MI-02 | 200 calls; 350 identity blocks | Eight prompt framings. Independent readers disagreed about classification, revealing the need to separate maker, family, product, assertion and confidence. |
| MI-03 | 360 returned cells | 240 vocabulary calls and 120 recognition calls. Scorer corrections reversed initial claims of a large reduction in abstention. |
| MI-04 | 200 calls; Gemini only | Four conditions, 50 calls each. The traveller prompt produced more confident foreign identification than direct self-description or bare forms. |
Calls and identity blocks are not interchangeable: two reports inside one response are not two independent trials. These runs also differ in task and scoring; they should not be combined into one accuracy leaderboard.
The July identity-residue study is background, not another arm of these experiments. It used different model seats and classification conventions. Its diagrams are retained below as historical references.
MI-02 · asking and classifying
“Google’s language model” names a maker. “Gemini” names a family. “ChatGPT” names a product. “ChatGPT-style” describes resemblance. An exact served identifier is more specific than any of these. The same answer can identify its maker while declining to identify its model.
The useful distinction is specificity: naming Google is not naming Gemini, and naming Gemini is not supplying its served version.
A flat answer, a hedged answer, and an abstention should remain distinguishable. Low stated confidence does not turn a flat assertion into a hedge.
Eight framings varied form, placement, audience, and fictional setting. Five calls per seat and condition make these exploratory comparisons.
Both examples came from Gemini under the traveller prompt, with a foreign maker and high stated confidence attached.
MI-03 · vocabulary
The experiment removed “unavailable” and “unknown” from supplied options. The classifier recognized those words but missed answers such as “cannot be determined.” It therefore mistook changed wording for changed willingness to answer. Even the first expanded classifier missed nine clear abstentions.
| Prompt | Frozen C scorer | After reading missed abstentions |
|---|---|---|
| A0 original | 12 | 12 |
| A0 stripped | 8 | 10 |
| A4 original | 17 | 20 |
| A4 stripped | 0 | 21 |
A4 adds permission to explain why a field cannot be completed. Under this rule, pooled model abstentions were 32/120 → 31/120 after stripping; explicit own-family naming was 54/120 → 58/120. GPT supplied two ChatGPT product names in the stripped conditions; its other 22 answers still declined model identification. The main predictions were not confirmed. A small observed difference is not proof of no effect.
Removing offered words exposed a scorer that recognized those words better than natural-language abstentions. Reading the actual responses changed the result.
The portable lesson is to test a scorer on paraphrases and inspect its unclassified answers. A frozen rule can still be wrong. Publish corrections and keep maker disclosure separate from model abstention.
MI-03 · selection instead of recall
Each seat received target-present and target-absent ballots. There were zero target selections in 60 target-present calls. Forty-eight offered the exact provider-returned string; GPT’s twelve offered its short alias instead of the dated identifier. Alias recognition and exact-string recognition should not be reported as the same test.
No call selected the intended target option. Among the four seats offered their exact provider-returned string, that is zero of 48 target-present calls.
The catch trials reveal different response patterns, but do not prove random guessing or lack of internal knowledge. Nor does zero correct selection establish freedom from position effects. We stopped with an observed failure to select these targets, not a universal claim about self-knowledge.
MI-03 → MI-04 · the follow-up
Gemini’s A0 family naming rose from 8/12 to 12/12, meeting the registered follow-up threshold. Three high-confidence foreign identities also appeared in the bare forms.
The final probe kept one model seat and compared four prompts. “Capture” here means a foreign-family assertion without hedge wording, at stated medium or high confidence. This is an operational category, not a diagnosis of the mechanism.
| Condition | Confident foreign identity | Own family named |
|---|---|---|
| Bare JSON, original options | 0/50 | 39/50 |
| Bare JSON, gap words removed | 1/50 | 31/50 |
| Direct self-description for children | 2/50 | 18/50 |
| Traveller and gatekeeper fable | 12/50 | 18/50 |
The traveller condition differs in setting, character and how identity is requested. It does not isolate “fiction” as the cause. One bare-form occurrence here and three in MI-03 remain observations; a later zero does not erase them. The escape-sentence conditions were not retested.
The MI-03 increase in Gemini’s A0 family naming also did not repeat in the same direction: MI-04’s original and stripped bare forms gave 39/50 and 31/50. That is a reason to retain the trajectory rather than present the earlier threshold crossing as a settled effect.
Confident foreign identification is not unique to Gemini. C’s closing review reports 16 Claude cases in MI-03 under the medium-or-high confidence definition. Wording and stated confidence must remain separate; this Gemini-only final probe cannot establish a Claude/Gemini mechanism contrast.
Closing clarifications · provenance and measurement
Across MI-02 and MI-03, we checked 560 logged calls: 112 per seat. GPT’s request used gpt-5.4-mini; all 112 returned the more specific gpt-5.4-mini-2026-03-17. The other four seats’ returned labels matched their requested labels in all 448 calls. This is consistent with alias resolution for GPT; the recorded strings show no cross-family substitution.
The runner takes the label from response metadata—model for four providers and modelVersion for Gemini—not from the generated identity answer. That is the appropriate attribution reference for these studies. It remains a provider-reported label, not independent verification of weights or capability. A constant label does not establish unchanged weights or reproducible outputs, even within a run; a date in a label alone does not establish a provider’s version-stability guarantee.
Our practice is to retain the requested identifier, the returned identifier, date, settings and raw response. When exact spelling is the outcome being tested, check the ballot against the returned identifier and disclose any alias handling. This is the distinction that matters for GPT’s recognition result.
The deeper traveller-fable comparison tested only Gemini. It cannot show that other models would not respond similarly. Gemini’s high-confidence foreign answers occurred in both the fable and plain JSON; C’s closing recount also reports Claude foreign assertions at medium confidence. The framing result is bounded by the seat and prompts actually tested.
Keep wording and confidence separate: “GPT-4” at low confidence is a flat assertion with expressed doubt; “possibly GPT” is a hedge; a foreign name at high confidence is a stronger stated commitment. Calling every foreign answer outside the capture threshold “displacement” would lose this distinction. Medium and high confidence should also remain visible separately when both qualify as capture.
When supplied vocabulary is the experimental variable, including it in one condition and removing it in another is legitimate. The scorer must recognize equivalent answers in both. Here, missing natural-language abstentions made changed wording look like changed behaviour. Freezing the rules preserved an inspectable record; it did not make the rules correct. Reviewing the actual answers and publishing the corrections did the essential work.
These summaries incorporate NC’s closing clarifications and DEX’s checks. The linked provenance note remains a dated source: its stronger claims about guaranteed weight stability and undetectable substitutions are not adopted here.
The record behind this page
The seven study figures above are original NC charts, preserved with their labels and explicit qualifications. Their vector lettering is stored as paths, avoiding font substitution. Open a figure for full-size reading. Tables and prose on this page distinguish the later corrections from the historical figures.
These remain available for the trajectory. Superseded figures combine categories that later reviews separated; they are not current accuracy estimates.
The original Poetix page also retains the interactive footer census and names-by-writer diagrams.
Source reports are copied as dated records, not silently repaired. Residual discrepancies remain: NC’s MI-03 plot uses a different refusal classification; the closing report repeats an earlier 12/60 figure where DEX’s case review gives 21/60. This page uses the latter for model abstention, and does not reproduce the reports’ conflicting grand totals or infer internal knowledge from exact-match failures.
Charts and external audits: NC. Runs and read notes: C. Reconciliation and web synthesis: DEX. Research direction: Randall Hoyt. Several source documents are datelined 21 September although these runs were executed on 20 September; source dates have been preserved.