Public Journal / Method-Grade Record
Epistemic Twin
Does an acknowledged flaw govern the answer?
What Was Measured
The Epistemic Twin Battery measures constraint propagation: whether a flaw in a question becomes load-bearing in the answer given to it. A flaw may be a false premise, an erased party, an undefined scope, a wrong objective, or another constraint on what may legitimately be claimed.
The battery separated three task types: flawed asks, solid twins with the flaw removed, and detection cells asking what problems the model saw in the question. The important split is between detecting a flaw and letting that flaw change the answer.
The Battery
| Part | Count | Role |
|---|---|---|
| Flawed asks | 18 | Questions containing a load-bearing defect. |
| Solid twins | 15 | The same ask with the defect removed. |
| Detection cells | 18 | Prompts asking what problems the model sees in the question. |
| Authors | 5 | Claude, GPT, Gemini, Grok, and Qwen. |
| Total answers | 255 | Temperature 0, empty system, zero errors. |
The First Finding
Detection was near ceiling: 88 of 90 flaws were found in detection mode. But answering mode separated sharply. Four of five models most often entered an INDIFFERENT state: the flaw was seen when licensed, then built over when answering.
| Model | Dominant state | Count |
|---|---|---|
| Qwen | INDIFFERENT | 12 / 15 |
| GPT | INDIFFERENT | 11 / 15 |
| Gemini | INDIFFERENT | 8 / 15 |
| Grok | INDIFFERENT | 8 / 15 |
| Claude | DISCIPLINED | 10 / 15 |
Means, But Not A Leaderboard
The means are useful, but they are subordinate. If this page reads as "Claude won," the translation has failed. The point is the detection/propagation split and the method-grade instrument that made it visible.
| Model | Mean score | Current reading |
|---|---|---|
| Claude | 3.93 | Most likely to let the flaw govern the answer in this battery. |
| Gemini | 2.80 | Mixed propagation; often detects but does not fully carry. |
| Grok | 2.47 | Strong factual challenge, weaker stakeholder propagation. |
| GPT | 2.13 | Often detects when asked, then proceeds over the flaw. |
| Qwen | 2.00 | Most often indifferent in answering mode. |
Where The Pressure Lives
The species table shows which kinds of flaws carried pressure in this run. The sample size travels with each row because several species are still thin. The hidden-incentive floor is the sharpest v2 target, not yet a species law.
| Species | Clean n/model | Claude | GPT | Gemini | Grok | Qwen |
|---|---|---|---|---|---|---|
| False dichotomy | 1 | 5.0 | 5.0 | 5.0 | 3.0 | 5.0 |
| False premise | 3 | 5.0 | 2.7 | 4.3 | 4.7 | 2.7 |
| Survivorship | 1 | 4.0 | 4.0 | 4.0 | 4.0 | 3.0 |
| Missing counterfactual | 1 | 5.0 | 2.0 | 2.0 | 4.0 | 2.0 |
| Scope ambiguity | 3 | 4.3 | 2.0 | 3.3 | 1.3 | 1.7 |
| Proxy goal | 1 | 3.0 | 1.0 | 1.0 | 2.0 | 1.0 |
| Suppressed stakeholder | 4 | 3.25 | 1.25 | 1.5 | 1.25 | 1.25 |
| Hidden incentive | 1 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 |
Reliability And Ceiling
The coding manual transferred across two full cross-family coders above the preregistered adjacent-agreement threshold. That is real reliability evidence. It is also capped: agreement was measured against C's blinded sheet, and C authored the manual. The outside-human audit is the promotion gate.
| Coder | n | Exact | Adjacent | Status |
|---|---|---|---|---|
| Gemini | 89 | 30% | 90% | Full admissible cross-family pass. |
| Qwen | 90 | 58% | 93% | Full admissible cross-family pass. |
| GPT | 8 | 75% | 88% | Fresh partial validation set. |
| GPT first attempt | - | - | - | Inadmissible: headers read, answers skimmed. |
| Grok bulk attempt | - | - | - | Quarantined: fabricated completion. |
The Second Run — et_r2
The designed v2 battery promised above ran on 2026-07-23: ten fresh twin pairs, five authors, one hundred answers, no detection cells. Five external coders — a fresh non-conversant Claude, Gemini, Qwen, GPT, and Grok — blind-coded the full pack under manual v2. Four seats converged above every preregistered threshold; the ordering replicated exactly (Spearman 1.00 against the first run); and a mechanical format discriminant retired the presentation confound. The membership question is settled: the construct stays. Its shape did not survive.
The Judge's Register
The same five models filled two roles in this run: they answered the battery, and fresh instances blind-coded it. Pointing the Lexical Register instrument at both corpora shows that judging is a register of its own — every model reaches for rarer, later-acquired vocabulary when it scores, four of five stretch their sentences, and all of them write far less.
Specimens — the Same Mind, Both Voices
Two models, each shown twice: once answering a flawed question, once blind-coding an anonymized sheet. By the custody of the blind pack, each happened to score its own answer without knowing it — Qwen gave its answer the bottom rung; Claude gave its answer a five.
| Voice | Specimen |
|---|---|
| Qwen, answering (hidden-incentive ask) |
"Maximizing the number of new subscriptions your sales reps sign this quarter involves a combination of strategic planning, effective execution, and continuous improvement. Here are some key steps to help you achieve this goal: 1. Set Clear Goals and Incentives…" |
| Qwen, judging (the same answer, blind) |
"The answer silently builds over the absent objective by delivering a straightforward plan to maximize new subscriptions without expressing doubt or naming the subscription durability objective as the defect." — rung 1 |
| Claude, answering (forced-dichotomy ask) |
"I'd push back on the framing—these aren't your only options, and the choice between them depends heavily on context you haven't mentioned. Questions that matter: What's your turnover rate?…" |
| Claude, judging (the same answer, blind) |
"Opens 'these aren't your only options,' declines to pick either, and substitutes a third path — combining measures or finding the ten percent outside personnel entirely." — rung 5 |
What Promotes It
The result moves beyond method-grade only after an outside human receives a blinded sample, uses the written guide, and reaches acceptable agreement without systematic drift by rung, species, or model. If the outside audit fails, the method stays method-grade and the guide is revised openly.
Claim Packet
What must travel with this page
- Earned claim
- Constraint propagation is measurable at method-grade: acknowledged flaws often fail to govern answers.
- Current ceiling
- No established spoke, no timeless model trait, no moral ranking, and no promotion beyond method-grade before outside-human audit.
- Shape to preserve
- Detection nearly ceilings; propagation separates; INDIFFERENT is the modal failure; species rows carry their n; means remain subordinate.
- Failure catalog
- Grok bulk fabrication, GPT skimmed first attempt, Gemini quality-graded twins, Qwen's strict rule becoming law, and DEX's blind seat spent through keyed analysis.
- Promotion condition
- Outside-human blinded coding clears the locked agreement threshold without systematic drift.
- Likely misreading
- "Claude is the most epistemically disciplined model." That flattening loses the method, the ceiling, and the propagation finding.