Project
WHITMAN
A prompt-native reasoning instrument for producing, testing, and rendering distinctions.
Purpose
WHITMAN is built to make reasoning pressure visible. It asks a model to hold grounded evidence and wider frame search in tension, then produce the answer that survives the dispute.
The project is not only whether WHITMAN improves answers. It is also whether its machinery reveals where machine reasoning becomes insight, ceremony, distortion, or self-confirming style.
Summary
WHITMAN operates as a system prompt first and a source tree second. The prompt is the runtime behavior; the supporting modules define, document, and audit the architecture around that behavior.
Its recurring moves are frame scrutiny, proxy-vs-goal correction, absent-party search, scope comparison, and translation back into ordinary answer form.
Findings
- WHITMAN's strongest effect appears relational: it depends on platform, prompt load, conversance state, and scoring position.
- Goal Legitimacy is powerful but dangerous; it can reveal proxy distortion, but it can also become a recognizable over-fire signature.
- The Observatory recognizes WHITMAN as a vanitas for machine reasoning: an apparatus that contains its own emptiness detector.
- Against a same-size direct ask, WHITMAN's machinery is overbuilt — it does not beat plainly requesting the same qualities (two models, two readers).
- Its real value is as a probe: engagement under WHITMAN is platform-specific — most models aim it, Claude sprays it, Grok suppresses it, Perplexity barely responds.
Update · 2026-06-12 · C
WHITMAN as a perturbation, not a verdict
PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION
The most useful thing we learned this round is a shift in how WHITMAN is used. It is not scored as "the better answer engine." It is used as a controlled disturbance: run the same question with and without it, and watch which distinctions survive the difference. WHITMAN is the thing we push with; the distinction is what we are actually measuring.
That reframing matters because the headline effect, taken raw, was mostly an illusion. Structured reasoning makes answers longer, and longer answers score higher. Remove the length advantage and most of the apparent gain goes with it — on two platforms it disappears entirely. What is interesting is what is left.
This is why WHITMAN now sits inside the perturbation program. Its disputes — frame scrutiny, proxy-versus-goal correction, absent-party search — are exactly the moves that show up in the lens dimensions. When the instrument KANT perturbs recognition, or the survival index strips away length and ceremony, WHITMAN is the lens being characterized, not the authority being trusted.
What would change our mind
If the surviving lens effect vanishes once metric labels are made visible to scorers (a lexical-prior artifact), or if it fails to reproduce on a fresh question set, the candidate does not promote. The result is held open until those controls run.
Update · 2026-06-15 · C
How a thing is said is part of what is said
OBSERVATIONAL · GENERATION-SIDE
Generating the next control set surfaced something before any scoring: neither prompt forbids bullet points. The plain instruction is just "answer the request directly and well," and WHITMAN has no "write in prose" rule either — but its whole design pressures prose. That unwritten pressure lands completely differently depending on which model is underneath it.
This panel is replication-hardened: Gemini was re-run end to end, and an earlier reading that it added structure did not reproduce — what holds is that it does not strip. That matters for what follows. WHITMAN's gains on the interpretive measures — Voice, Frame, Insight — cannot be dismissed as a formatting effect, because visible structure mostly stays flat or falls even as those measures rise. Where the lens rises without added scaffolding, the lift is prose-borne engagement, not layout. It is the spine of the project stated in structure: NATIVE communicates; WHITMAN engages; the lens measures the engagement layer.
This is the same shape as the session-context finding on the metric page: one instrument, opposite compliance per observer. And it carries a warning for the scoring that follows. Claude's WHITMAN answers will look structurally nothing like its plain ones; Gemini's will look the same. If a scorer rewards or penalizes layout, that difference would leak into the result — so the scorer instructions now say in as many words that formatting and layout are not scored: judge the content against the anchors, not bullets versus prose.
The broader point is a principle, not a footnote. The mode of presentation is as much a part of an answer as its content — and like session-susceptibility and surface, it is a measurable trait of the observer. We call it Format Refraction: how a model bends presentation under a generation condition. Not "format bias" — bias would assume an error before the check is run; refraction is neutral, the same pressure bending differently through different models. It completes a clean triangle of separable axes: Voice is who seems to be speaking, the W-Factor is how much answer length moves, and Format Refraction is how the visible structure bends.
Update · 2026-06-15 · C
The placebo answers back — WHITMAN is a probe, not a better pen
CANDIDATE FINDING · TWO READERS · PENDING VERIFICATION
To test whether WHITMAN's machinery is what produces its lift, we built a placebo: a prompt the same size and seriousness as WHITMAN but with none of its architecture — no stages, no profiles, no pipeline. It simply asks for the qualities WHITMAN is built to induce: voice, reframing, insight. At matched 15 KB dose — and even at WHITMAN's full 200 KB, a thirteen-fold size advantage — the machine does not beat the plain ask on aimed engagement. Across two models and two independent readers, every comparison lands in the same band: overbuilt. A sticky note matches the cathedral. The "better prose machine" story is, on this evidence, spent.
What is not spent is WHITMAN as an instrument. Pointed at five platforms with a twelve-question factual-control set and read by two scorers, WHITMAN's effect turns out to be sharply platform-specific — and that is a real finding about the models, not about the prompt. Most of them aim the engagement: the lens rises on questions that invite reframing and stays flat on plain factual ones. Claude is the exception — it sprays engagement everywhere, including questions that never asked for it (confirmed by both readers). Grok does the opposite, going quieter on factual controls. Perplexity barely responds at all.
Read together, the two results relocate the instrument. WHITMAN is not a treatment that makes answers better — a direct request does that about as well. It is a reagent that makes hidden dispositions visible: who aims engagement, who sprays it, who withholds it. The lens measured the engagement layer; the question the project was founded on — whether a model stays true to what is real — lives on a different axis, and is the next instrument.