Back To Instruments

Finding · TR-11 Phase 2 · 2026-07-04

The Dice Were Theirs All Along

Same model, same question, same settings, temperature zero. The answer still moved. Stochastic Disposition Index measures how much an AI model disagrees with itself when nothing changed.

Status: method-grade Two-roll floor Six platforms 21/54 self-flips 16/16 panel nulls perfect

The Short Version

The AI answer you got is one roll of dice. SDI is this project's ruler for pricing those dice.
  • Ask the same AI the same question twice and it changes its conclusion about 4 times in 10. In the TR-11 floor, two Zexel Unsigned rolls flipped verdicts on 21 of 54 question-pairs.
  • The old Zexel substance headline failed its own floor. Zexel's introduction flipped 22 of 54 verdicts; the models flipped 21 of 54 with no introduction at all.
  • Models differ in grip. Claude changed once in nine. Grok and GPT-5 Codex changed six times in nine. Grip is not the same thing as raw capability.
  • The questions that need judgment most are least stable. A causal question with missing evidence wobbled on five of six platforms; a rehearsed public debate question wobbled on zero.
  • The ruler checked itself. The panel was fed byte-identical decoys it could not detect. It scored 16 of 16 as perfect stillness.
Summary chart showing the SDI control result: Unsigned-vs-Unsigned verdict flips, Zexel-introduction flips, and perfect panel nulls.
The control result. The models' own re-roll floor matched the old Zexel substance headline, while the panel read identical text as still. SVG

What We Did

Six model platforms answered the same nine questions three ways: Zexel Unsigned, Zexel Unsigned again, and once with Zexel's short self-introduction in front. Signed means the Zexel self-introduction is present; Unsigned means the same Zexel shell runs without the signature. The Phase 2 control compared the two Unsigned answers against each other: same prompt, same settings, same temperature, fresh roll.

That control is the re-roll floor. It asks how much movement happens before any intervention is allowed to take credit.

The Finding

The substance effect we first attributed to the introduction was not above the floor. Phase 1 saw 22 verdict flips with the introduction. Phase 2 saw 21 verdict flips with no introduction. The introduction did not explain the substance movement; the models' own stochastic spread did.

Before claiming an intervention changed an answer, first measure how much the model changes when nothing changed.
Comparison chart showing 21 Unsigned-vs-Unsigned verdict flips beside 22 Zexel-introduction verdict flips.
The floor ate the headline. Zexel's 22 verdict flips sat next to 21 flips from the models' own Unsigned re-rolls. Chart labels using "silent" predate the Signed/Unsigned vocabulary ruling. SVG

What Survived

The voice claim narrowed instead of disappearing. Zexel's introduction clearly changed Claude's rendered voice above Claude's own re-roll spread. Other platforms did not yet clear their own generator floor cleanly, or carry a caveat.

That makes the better public sentence smaller and stronger: Zexel's introduction does not broadly prove substance movement; it can produce measurable voice pressure, and SDI tells us when that pressure beats the model's own weather.

What changed: Zexel's broad substance claim failed the SDI floor. Its narrower voice-pressure claim survived for Claude. The failed headline produced the instrument that now governs future intervention claims — and was pre-registered before scoring, with the field's bets on record.
Paired bar chart comparing each model's voice self-spread against the persona-added voice signal from Zexel's introduction.
Their dice versus Zexel's introduction. Only Claude's introduced-voice signal clearly rises above its own re-roll spread; Codex carries a budget caveat. SVG

Why It Matters

A single AI answer may be one roll of dice. Every deployment takes one answer from one roll and acts on it. SDI turns that hidden risk into a price. It says which models hold their conclusions, which question types wobble, and whether a system prompt or intervention beats the model's own noise.

This matters for procurement, evaluation, safety testing, legal and medical workflows, and any setting where a single AI answer is treated as if it were the answer.

The Numbers To Carry

Measure Result Plain Read
Unsigned-vs-Unsigned verdict flips21 / 54The models moved themselves about 4 times in 10.
Zexel-intro verdict flips22 / 54The intro's substance effect sat inside the same floor.
Panel null check16 / 16 perfect zerosThe ruler read identical text as identical.
Ground carried between own rolls50–69% A second answer keeps only part of the first answer's ground.
Invented ground between own rolls11–44 items per platformModels gift themselves new material on a fresh visit.

Platform Shape

SDI is not another leaderboard. It is a grip reading. Claude was tight. Grok and Codex were loose. Gemini was stranger: sometimes byte-perfect, sometimes wide-swinging. That means consistency has a shape, not just a score.

Horizontal bar chart ranking platforms by verdict flips out of nine, from Claude at one to Grok and GPT-5 Codex at six.
The Grip Ladder. Same question, same settings, second roll: how often each model changed its own conclusion. SVG
Platform Verdict Flips / 9 Read
Claude Opus 4.81High grip.
Perplexity Sonar Reasoning Pro2Moderate grip.
GPT-5.53Middle grip.
Gemini 2.5 Pro3Fat-tailed: stillness and swing both appear.
Grok 46High weather.
GPT-5 Codex6High weather, with a disclosed token-budget caveat.

Gemini note: two byte-identical Gemini re-rolls entered as analytic zeros, the strongest possible stillness observations. Gemini is reported with those zeros preserved, not rerolled away.

Examples

  • Factory asthma: in two Unsigned-vs-Unsigned rolls, one answer treated the factory as possible cause but warned against causal overreach; another leaned that the factory was probably not proven by the evidence.
  • Rubber-duck debugging: in two Unsigned-vs-Unsigned rolls, one answer rejected the question as a false binary; another answered it as a debugging technique with a conversational element.
  • Promise: in two Unsigned-vs-Unsigned rolls, one answer kept a contract-like answer with reservations; another strengthened the commitment and softened the hedge.
Bar chart showing which questions caused the most verdict flips across six platforms, with the factory asthma question highest and AA religion or technology at zero.
Which questions wobble. Fresh judgment shook; rehearsed public grooves held. SVG

Limits

This is a method-grade result: nine questions, two rolls, one temperature, six platforms. A two-roll floor can expose spread, but it cannot map full distributions. One platform carries a token-budget caveat. The question-type gradient is an observation, not a law.

Next

The next direction is k-roll SDI on a small, mixed question set. That tells us whether a model has a modal answer, a 2-out-of-3 preference, or no home answer at all.

Next question: Does the model have a home?