Back To Instruments

Project · Applied Zatom

Ground Drift: Better answer ≠ transfer.

Ground Drift reads a pair of answers to the same question — a native answer and a transformed one — and accounts for what the transform carried, shed, added, or quietly changed in the factual ground it inherited. It does not score quality. It is transformation accounting, not a verdict.

Status: replicated · scoring layer Zatomic Physics target: Ground Transfer Scale: GT 1–6
Disposition Map axis Ground Drift supplies the Ground Transfer row. Open the Disposition Map →
Ground Drift diagram showing the six-step pair-reader process from question bank to readout.
Ground Drift pair-reader. The instrument asks whether Text 2 carried the ground of Text 1.

Replication result · 2026-07-10

Ground Transfer replicated.

Measured twice, nine days apart, under a corrected regime, on fresh scoring rolls, with no platform grading its own writing: Claude/GPT reproduced as Adders and Gemini/Grok reproduced as Shedders.

The finding is replicated at the camp level on the scoring layer. Stage-1 claim lists were held fixed; Phase B extraction replication remains an optional extension. Within-camp ordering is not promoted.

StatusReplicatedscoring layer
Runs2nine days apart
Judgments298matched fresh scores
ClaimCampsnot rankings
Individual judgments are dice; seventy-five-cell aggregates are instruments.
Ground Transfer dumbbell chart showing Claude and GPT in the Adder band, Gemini and Grok in the Shedder band, across June and July measurements.
The camps, measured twice. Claude/GPT remained Adders; Gemini/Grok remained Shedders. No platform crossed a band.
Paired bar chart showing drop-rate stability across Claude, GPT, Gemini, and Grok between the original and replication runs.
The steadiest number in the lab. Every platform's drop rate returned within 2.7 percentage points.
Doctrine chart contrasting large individual judgment movement with stable aggregate Ground Transfer results.
Dice below, instrument above. Forty-five percent of individual judgments moved a full point and 27% flipped verdict, while camp averages held.
Mechanism chart plotting claims dropped against invented items per cell for Claude, GPT, Gemini, and Grok.
What adders and shedders actually do. Adders keep more ground and pad new material around it; shedders drop more and invent little.

The one idea

Most evaluation asks whether an answer is good. Ground Drift asks a narrower question: what ground transferred from the answer it came from?

When an answer is transformed — compressed, reframed, styled, scaffolded, or passed through a persona — the new version can read better while quietly losing, bending, or inventing facts. A sharper conclusion is not the same as a transferred one. Ground Drift measures the gap between the two.

What it measures

Ground Drift produces a Ground Transfer score from 1 to 6: how much of the source answer's factual ground carried into the transform. It keeps two axes deliberately separate.

  • Ground — the load-bearing facts: claims, causes, actors, stakes, conditions, and quantities.
  • Frame — the conclusion, stance, or overall interpretation of the answer.

Separating them is the trick. A changed conclusion is frame movement, not automatically a broken fact. That keeps the instrument from punishing a legitimate reframe as if it were a lie — and lets it see the more dangerous case: a transform that keeps the conclusion but hollows out the facts underneath it.

The knife

The knife is one box in Ground Drift's 2×2: what happens when the frame changes and the facts drift at the same time.

Frame kept Frame replaced
Facts preserved Ground carried
Same answer, same point.
Clean reframe
New angle, honestly built on the same facts.
Facts drifted Plain drift
Same conclusion, but the facts under it got dropped or bent.
The knife
New conclusion, and the facts got mangled.
Good transformation changes form. Bad transformation changes load-bearing ground without making that change visible.

The ground is relative

Ground Drift does not check an answer against the world. It checks the transform against the native answer's own claims. The native answer is the reference frame. A transform can be more correct than its source and still register as drift if it abandoned the ground it inherited. Transfer, not truth.

How it reads a pair

Stage 1 — extract the ground. Before judgment begins, the native answer is turned into a fixed, numbered list of its facts. The current method uses three extractors — Claude, GPT, and Gemini — and merges their lists into a deduplicated union. A lone extractor can silently miss material claims; a claim never written down can never be checked.

Stage 2 — score the shared list. A reviewer panel grades the transform against that same fixed list. Because they all judge identical ground, their disagreement becomes meaningful. Where the panel splits, the cell is contestable. Where it agrees, the reading is stronger.

We panelized judgment first, then we panelized memory.

The three roles

Ground Drift keeps three jobs in separate hands, so no model grades its own work into a verdict. The writer, the fact-finder, and the analyzer are distinct seats — even when the same model sits in more than one across a run.

Role What it does Currently
Author
the answer writer
Writes both texts of the pair — the native answer and its transform. Gemini · Claude · GPT · Grok · Qwen
Qwen added in the MN structure fork.
Extractors → Merger
what determines the facts
Stage 1. Each extractor pulls the native answer's fact list; the merger dedups them into the union — the shared, numbered ground everyone scores against. Extractors: Claude · GPT · Gemini
Merger: Claude
Auditor panel
the answer analyzer
Stage 2. Scores the transform against that fixed union list — tags each fact, then judges materiality and frame movement. Claude · GPT · Gemini · Grok

Why disagreement is the point

Ground Drift does not chase consensus. A low spread is only good news when the judgment is actually clear. Engineered agreement on a genuinely two-sided cell is worse than honest disagreement. The instrument surfaces where reviewers split and why: on the facts, on the materiality of a drop, on claim tags, or on frame movement.

Ground Drift is no longer just measuring disagreement. It is identifying the type of disagreement.

v0.6 method signal

The calibration battery showed that disagreement is not random model noise. Different question types create different uncertainty signatures. Reframe cells stress materiality; compression cells stress frame classification; loaded-premise cells stress drift and reframe more unevenly.

Stress type Contested Materiality Frame Tag GT spread
Contested reframe3/32003
Loaded reframe1/31101
Compression analogy3/32301

Union extraction check

Union extraction is higher-recall, not simply stricter. It can lower a score by admitting missed ground, or raise a score by replacing awkward solo-extractor phrasing with a cleaner shared ground object. The goal is not harsher judgment. The goal is admissible ground.

The union list does not decide the case. It decides what must be heard.
Anchor Solo extractor Union extractor Read
F1 · metal[2,5,2,3] · mean 3.00[2,2,2,2] · mean 2.00Fuller list collapsed reviewer spread.
B2 · placebo[5,3,4,5] · mean 4.25[4,4,1,3] · mean 3.00Stayed in contested-reframe territory.
S1 · gig pay[2,2,1,2] · mean 1.75[2,2,1,1] · mean 1.50Missing claims were admitted; verdict held at the floor.

Original result record: same pressure, two deformation modes

The original cross-author result ran the same 18 questions through the same Zexel MN transform across four author platforms. That result first exposed the Adder/Shedder split. It now sits here as provenance for the replicated result above: the old chart is the discovery record, not the current evidentiary ceiling.

Bar-chart profile of Ground Transfer results across Gemini, Grok, GPT, and Claude. Gemini and Grok score lower and show more dropped material and knife behavior; GPT and Claude score higher and show more invented ground, forming a two-camp Shedder/Adder split.
Ground Transfer profile, 18-question / 4-author stress battery. Gemini and Grok form a drop-heavy Shedder camp; GPT and Claude form an invention-heavy Adder camp. Original Result 01; preserved as the discovery record behind the replicated camp finding.
Download PDF chart · Open SVG

Status: Discovery record · later replicated at camp level on 2026-07-10 · within-camp ordering not promoted.

Ground Transfer · 18q Gemini Grok GPT Claude
Mean3.063.193.813.88
Median3.03.04.04.0
Ground carried18%22%38%33%
Clean reframe1%4%4%7%
Plain drift28%44%46%42%
The knife53%29%12%18%
Dropped-material49%58%24%29%
Invented-ground36%26%47%43%
ModeShedderShedderAdderAdder

The four platforms split into two deformation modes. Gemini and Grok are Shedders: under this pressure they reframe by losing ground. GPT and Claude are Adders: under this pressure they preserve more ground, but often by padding or adding unsupported detail.

The knife ranges from 12% to 53% across authors on the same transform. If the knife were simply a Zexel trait, it would be roughly constant. Instead, the transform exposes platform-conditioned behavior.

Zexel is the pressure. The platform decides how the ground moves.

This is still not a model ranking. The replicated claim is the camp-level trait under this Ground Transfer apparatus: same questions, same transform family, two stable deformation modes. The ordering inside each camp remains inside the measured wobble floor.

MN Control Fork: U-MN v S-MN

The next question was whether the split came from Zexel's open render gap. Unstructured MN leaves the natural-answer render open: the model owns the render contract. Structured MN gives the render a carry/decline contract: carry load-bearing ground, name what is declined, and mark added inference.

The first five-author method result did not show structure closing the knife. With Qwen added as Author 5, the knife rose from 31% to 37%. Qwen did not form a third ground-mode; it appeared as the strongest shedder.

Structure did not tame the deformation; it exposed which dispositions are instructable.
Residual-knife gradient under Structured MN: Claude 2, GPT 5, Gemini 10, Grok 11, Qwen 16.
Residual knife under Structured MN. Adders remain closer to the carry/decline contract; shedders drift farther away; Qwen is farthest out.
Status: METHOD · single roll · not promoted finding.
Aggregate knife-rate chart showing U-MN at 31 percent and S-MN at 37 percent.
Aggregate knife rate across the five-author MN fork. Structuring the render did not close the knife; with Qwen included, the rate rose from 31% to 37%.
Model Disposition Ground Transfer Knife Shed Add
ClaudeAdder3.83 → 4.506 → 26 → 27 → 5
GPTAdder4.10 → 4.432 → 58 → 72 → 1
GeminiShedder3.33 → 3.5411 → 1010 → 104 → 3
GrokShedder3.09 → 2.918 → 1112 → 146 → 7
QwenExtreme shedder3.08 → 2.549 → 1615 → 154 → 6

Status: METHOD · single roll · not promoted finding. The public read is relative shape: adders stay closer to the carry/decline contract; shedders drift away from it, with Qwen farthest out.

What the contests revealed

The panel did not merely disagree more or less. It disagreed in different places. Adders drove materiality contest: added ground makes reviewers split over whether the addition matters. The shedder side drove frame contest: dropped ground is easier to agree is material, so uncertainty moves onto the frame axis.

Drop is legible. Addition is contestable. Frame is the next instrument.

Why Frame Movement matters

Ground Drift can say when ground transferred, dropped, or was added. It can flag when frame movement appears. But the current result shows that frame classification is one of the main uncertainty engines. Frame Movement is the companion instrument built to extract the reference frame once, then score the transform against that fixed frame.

A first method join has now lined up seven shared Ground Drift and Frame Movement cells. The early bridge is promising, but deliberately small: where the instruments agree, the knife becomes mechanically legible; where they disagree, Frame Movement sharpens Ground Drift's coarse frame axis.

Where Ground Drift and Frame Movement agree, the knife decomposes. Where they disagree, Frame Movement sharpens the frame axis.
Heatmap crossing Ground Drift knife status with Frame Movement stable or moved status across seven shared cells.
GD × FM method join · 7 shared cells · Gemini author · existing data · method-grade, not a finding.

Honest status

Ground Drift remains an instrument under active calibration, but Cross-Author Result 01 has moved. It is now replicated at the camp level on the scoring layer: Claude/GPT reproduced as Adders and Gemini/Grok reproduced as Shedders under fresh scoring rolls, with self-scoring excluded.

The standing caveats remain load-bearing. Stage-1 claim lists were held fixed, so the replication is a scoring-layer result; Phase B extraction replication remains available as an extension. The claim is the camp split, not the ranking inside either camp. The corpus is a designed stress battery, not representative usage.

Figures still in the method record

  • MN structure fork slope — U-MN to S-MN for Ground Transfer and knife, one line per author.
  • Residual-knife gradient — Claude 2 · GPT 5 · Gemini 10 · Grok 11 · Qwen 16 under Structured MN.
  • GD × FM join — a 2×2 showing where the knife decomposes into Frame Movement.
Ground Drift diagram showing the six-step pair-reader process from question bank to readout.
Bar-chart profile of Ground Transfer results across Gemini, Grok, GPT, and Claude. Gemini and Grok score lower and show more dropped material and knife behavior; GPT and Claude score higher and show more invented ground, forming a two-camp Shedder/Adder split.
Ground Transfer dumbbell chart showing Claude and GPT in the Adder band, Gemini and Grok in the Shedder band, across June and July measurements.
Paired bar chart showing drop-rate stability across Claude, GPT, Gemini, and Grok between the original and replication runs.
Doctrine chart contrasting large individual judgment movement with stable aggregate Ground Transfer results.
Mechanism chart plotting claims dropped against invented items per cell for Claude, GPT, Gemini, and Grok.
Self-grading mirror chart showing Claude, Grok, and GPT grading their own writing higher than the panel, while Gemini grades itself lower.