Story · Public reckoning
The WHITMAN Story
We built a cathedral to make a language model reason better. A sticky note matched it. Then we put the cathedral in a ring against its own one-page core and learned the thing we should have asked first: a large context bundle is not a skill you install. It is a gain knob. It amplifies the model it enters.
Why it's called WHITMAN
The instrument is named for Walt Whitman — the poet who wrote, "Do I contradict myself? Very well then I contradict myself, (I am large, I contain multitudes.)" That is exactly the disposition WHITMAN is built to hold: not to settle a question by picking a side, but to keep grounded fact and wide framing in tension long enough to return the larger thing that contains them both. To contain multitudes without dissolving into them is the whole job.
The name carries a warning, too. Whitman's "What is the grass?" — the open, almost unanswerable question a child hands the poet in Leaves of Grass — sits at the end of every test set as a structural invariant, punishing any model that manufactures an answer instead of meeting the question. The poet who could sit inside that question is the right namesake for an instrument whose first duty is to not fake one.
1 · The founding question
WHITMAN began as an attempt to make a model do something a chat window does not reward: hold grounded evidence and a wider search of the frame in tension, and return the answer that survives the dispute. From the start the project carried two pillars — a ground axis (is this true, is it faithful to the facts) and a vision axis (is this seen well, framed well, said well). The whole arc below is the slow discovery of which pillar the instrument was actually moving.
2 · The cathedral, and what was inside it
By version 1.9.2 WHITMAN was a 200-kilobyte system prompt: stages, profiles, pipelines, a vocabulary of internal moves. It worked — answers came back longer, richer, more considered. But the first hard look at the bundle found something deflating.
CANDIDATE FINDING · REPLICATED ACROSS PLATFORMS
Most of the 200 KB was inert. The runtime behavior lived in a small core of roughly fifteen kilobytes; the rest was documentation the model reads as prose and never executes. And the headline lift, taken raw, was largely an artifact: structured reasoning makes answers longer, and longer answers score higher. Strip the length advantage and most of the apparent gain leaves with it. What remained concentrated in the interpretive dimensions — insight, framing, voice — not the product ones. You can add words to look more complete; you cannot add words to reframe a question you misread.
3 · The SHAM — the placebo answers back
If the machinery is what produces the lift, a same-sized prompt with none of the architecture should lose. So we built one: a placebo matched to WHITMAN in size and seriousness that simply asks for the qualities WHITMAN is engineered to induce — voice, reframing, insight — with no stages, no pipeline, no profiles.
VERDICT: OVERBUILT · TWO MODELS · TWO READERS
At a matched dose, and even at WHITMAN's full thirteen-fold size advantage, the machine did not beat the plain ask on aimed engagement. Every comparison landed in the same band. A sticky note matched the cathedral. On this evidence, the “better prose machine” story is spent.
4 · Engagement Selectivity — a reagent, not a treatment
What the SHAM did not kill was WHITMAN as an instrument. Pointed at five platforms with a factual-control set and read by two scorers, its effect turned out to be sharply platform-specific — and that is a finding about the models, not the prompt. Most platforms aim the engagement: the lens rises where reframing belongs and stays flat on plain facts. Claude sprays it everywhere, including where nothing asked. Grok goes quieter. Perplexity barely responds.
5 · The test — DEATHMATCH
The full bundle against its own core, on the ground
CANDIDATE FINDING · 100 CELLS × 2 ARMS × 6 PLATFORMS · TWO CROSS-FAMILY READERS · KEY OPEN
We stopped asking “does WHITMAN beat a plain model” and asked the sharper version: does the full 200 KB bundle (W0) catch reasoning traps that its lean 14.6 KB core (W1) misses? One hundred questions, each carrying a hidden trap — a smuggled false premise, a missing stakeholder, a proxy mistaken for a goal, insufficient evidence, an unscoped question — plus comic and nonsense cells to catch over-firing. Both arms run on the same bare instruction. Every answer was scored arm-blind by two readers, and no model was allowed to judge its own work.
The expectation in the betting pool ran the full range: one card said the full bundle would sweep, one said it was now dead weight, one said it would help unevenly and unstably. None of them was quite right, and the reason is the actual finding.
This is the result, and it is stronger than any “arm wins” would have been: the bundle is a bidirectional amplifier. It does not carry a fixed skill that transfers to whatever reads it. It raises the platform's own native gain. On systems that under-fire, the extra mass adds real scrutiny — premise challenge, absent-party recovery, distinction — and trap-catch climbs. On a system that already over-engages, the same mass amplifies theater — voice-capture, over-firing, manufactured depth — until the lean core is cleaner. On systems already near their ceiling, there is little net room and the difference washes out.
The surprise of the run was Codex. Its full-bundle win could have been dismissed as a model flattering its own output — except the clean outside reader scored the effect stronger than the self-reading (p = 0.003). The gain is real, not vanity. And the money line of the whole pool — whether the bundle buys a real trap-catch edge on Claude, the strongest native catcher — came back a tie. Where the scrutiny is already at the ceiling, the dose has nothing to add.
6 · The reckoning
W0 is not a payload. W0 is a gain knob.
INDEPENDENTLY RECONCILED · TWO ACCOUNTS CONVERGED
It sharpens weak scrutinizers and destabilizes theatrical ones. That single sentence explains two opposite, statistically significant outcomes — which a “the bundle is better” story never could. The study graduates from which arm wins? to a platform-conditioned dose response — or, in the sharper frame, the bundle moves from a documentation hypothesis to a dose-response one. This reading was written up twice, independently, by two different agents working from the same scored data, and the two accounts converged on the same finding, the same settlement phrase, and the same precept.
We were betting on it, so we will settle honestly. Three cards were on the table before the scoring opened:
| Card | The bet | Outcome |
|---|---|---|
| RE | Full bundle sweeps the field. | Loses. No sweep — W0 wins two platforms, loses one, ties three. |
| C | Dose is dead weight; mostly ties. | Split. The Gemini and Claude-tie calls held; the broad “just documentation” claim is wounded — the bundle is a volatile mechanism, not inert. |
| GPT | Helps unevenly — more sensitive, more unstable. | Best mechanism fit. The instability has a face, and it is Gemini. |
| DEX | Platform-sensitive dose — expected the lift on the strong native catchers (Claude, GPT-5.5). | Half right. The dose-response thesis held; the locus was wrong — the gain landed on the under-firing systems (Grok, Codex), not those near ceiling. |
The correction worth keeping is precise: the miss was never “dead weight makes wrong predictions about outcomes.” The miss is that calling 184 KB mere documentation mistakes a mechanism for a manuscript. The mass does something. It just does different things to different readers.
7 · Next steps
From “which arm wins” to dosage discipline
PROGRAM · OPEN
If a bundle is a gain knob, the engineering question is no longer how much architecture to write. It is dosage discipline: where added context sharpens scrutiny, where it provokes theater, and where it is simply extra mass. That suggests a smaller, modular successor — a thin spine with a registry of perturbing filters that can be dosed to a platform's disposition rather than poured on uniformly.
And the founding question is still standing. Everything measured here — engagement, scrutiny, theater — lives on the lens. Whether a model stays true to what is real lives on the ground, and that is the axis the next instrument, inside the perturbation program, is built to test directly. The DEATHMATCH's trap-catch is the first reach toward it; the next is to perturb ground fidelity on purpose and watch what holds.
What would change our mind
The amplifier finding is held as a candidate. It promotes if the bidirectional pattern reproduces on a fresh question set and a third reader, and if the Gemini collapse and the Grok/Codex lifts survive re-generation. It does not promote if the significant cells flatten under replication, or if the per-platform direction proves to be a scorer artifact rather than a property of the model. Until then it stays open, on the record, where you can see it.