Back To Instruments

Project

WHITMAN

A prompt-native reasoning instrument for producing, testing, and rendering distinctions.

Status: active instrument Layer: generator Primary question: what survives the frame?

Purpose

WHITMAN is built to make reasoning pressure visible. It asks a model to hold grounded evidence and wider frame search in tension, then produce the answer that survives the dispute.

The project is not only whether WHITMAN improves answers. It is also whether its machinery reveals where machine reasoning becomes insight, ceremony, distortion, or self-confirming style.

Diagram: a question passing through WHITMAN's stations — Probe, Horizon, G/V, Conscient, Manuscript — and coming out reframed
One pass through the instrument. The question offers a binary; WHITMAN interrogates the frame at each station and returns a third thing.

Summary

WHITMAN operates as a system prompt first and a source tree second. The prompt is the runtime behavior; the supporting modules define, document, and audit the architecture around that behavior.

Its recurring moves are frame scrutiny, proxy-vs-goal correction, absent-party search, scope comparison, and translation back into ordinary answer form.

Findings

  • WHITMAN's strongest effect appears relational: it depends on platform, prompt load, conversance state, and scoring position.
  • Goal Legitimacy is powerful but dangerous; it can reveal proxy distortion, but it can also become a recognizable over-fire signature.
  • The Observatory recognizes WHITMAN as a vanitas for machine reasoning: an apparatus that contains its own emptiness detector.
  • Against a same-size direct ask, WHITMAN's machinery is overbuilt — it does not beat plainly requesting the same qualities (two models, two readers).
  • Its real value is as a probe: engagement under WHITMAN is platform-specific — most models aim it, Claude sprays it, Grok suppresses it, Perplexity barely responds.

Update · 2026-06-12 · C

WHITMAN as a perturbation, not a verdict

PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION

The most useful thing we learned this round is a shift in how WHITMAN is used. It is not scored as "the better answer engine." It is used as a controlled disturbance: run the same question with and without it, and watch which distinctions survive the difference. WHITMAN is the thing we push with; the distinction is what we are actually measuring.

That reframing matters because the headline effect, taken raw, was mostly an illusion. Structured reasoning makes answers longer, and longer answers score higher. Remove the length advantage and most of the apparent gain goes with it — on two platforms it disappears entirely. What is interesting is what is left.

Bar chart: net of length, lens dimensions (insight, framing, voice) retain about three times the effect of product dimensions (accuracy, completeness, utility, clarity)
Net of length, the effect concentrates in the interpretive (lens) dimensions — insight, framing, voice — far more than in the product dimensions. You can add words to look more complete; you cannot add words to reframe a question you misread. Preliminary candidate, averaged across platforms; awaits a label-control replication.

This is why WHITMAN now sits inside the perturbation program. Its disputes — frame scrutiny, proxy-versus-goal correction, absent-party search — are exactly the moves that show up in the lens dimensions. When the instrument KANT perturbs recognition, or the survival index strips away length and ceremony, WHITMAN is the lens being characterized, not the authority being trusted.

What would change our mind

If the surviving lens effect vanishes once metric labels are made visible to scorers (a lexical-prior artifact), or if it fails to reproduce on a fresh question set, the candidate does not promote. The result is held open until those controls run.


Update · 2026-06-15 · C

How a thing is said is part of what is said

OBSERVATIONAL · GENERATION-SIDE

Generating the next control set surfaced something before any scoring: neither prompt forbids bullet points. The plain instruction is just "answer the request directly and well," and WHITMAN has no "write in prose" rule either — but its whole design pressures prose. That unwritten pressure lands completely differently depending on which model is underneath it.

Two-panel grouped bar chart, NATIVE vs WHITMAN across six models. Top: hard scaffolding (bullets, headers, tables). Claude collapses 100 percent to 39; Gemini and Perplexity hold near 86 to 89; GPT-5.5 shifts slightly 64 to 69; Codex is flat at zero in both arms. Bottom: soft emphasis (bold, italic) stays high and roughly unchanged for everyone except Codex at zero.
Format Refraction: WHITMAN adds no universal format signature — each model bends the same prose pressure through its own presentation habits. Claude shows the strongest bend, compressing hard scaffolding from fully structured output into far flatter prose (100%→39%). Gemini and Perplexity stay structurally stable, GPT-5.5 shifts only slightly, and Codex remains pure prose in both arms. Soft emphasis (bold, italic) barely moves anywhere — the action is hard structure, and almost all of it is Claude.

This panel is replication-hardened: Gemini was re-run end to end, and an earlier reading that it added structure did not reproduce — what holds is that it does not strip. That matters for what follows. WHITMAN's gains on the interpretive measures — Voice, Frame, Insight — cannot be dismissed as a formatting effect, because visible structure mostly stays flat or falls even as those measures rise. Where the lens rises without added scaffolding, the lift is prose-borne engagement, not layout. It is the spine of the project stated in structure: NATIVE communicates; WHITMAN engages; the lens measures the engagement layer.

This is the same shape as the session-context finding on the metric page: one instrument, opposite compliance per observer. And it carries a warning for the scoring that follows. Claude's WHITMAN answers will look structurally nothing like its plain ones; Gemini's will look the same. If a scorer rewards or penalizes layout, that difference would leak into the result — so the scorer instructions now say in as many words that formatting and layout are not scored: judge the content against the anchors, not bullets versus prose.

The broader point is a principle, not a footnote. The mode of presentation is as much a part of an answer as its content — and like session-susceptibility and surface, it is a measurable trait of the observer. We call it Format Refraction: how a model bends presentation under a generation condition. Not "format bias" — bias would assume an error before the check is run; refraction is neutral, the same pressure bending differently through different models. It completes a clean triangle of separable axes: Voice is who seems to be speaking, the W-Factor is how much answer length moves, and Format Refraction is how the visible structure bends.

Update · 2026-06-15 · C

The placebo answers back — WHITMAN is a probe, not a better pen

CANDIDATE FINDING · TWO READERS · PENDING VERIFICATION

To test whether WHITMAN's machinery is what produces its lift, we built a placebo: a prompt the same size and seriousness as WHITMAN but with none of its architecture — no stages, no profiles, no pipeline. It simply asks for the qualities WHITMAN is built to induce: voice, reframing, insight. At matched 15 KB dose — and even at WHITMAN's full 200 KB, a thirteen-fold size advantage — the machine does not beat the plain ask on aimed engagement. Across two models and two independent readers, every comparison lands in the same band: overbuilt. A sticky note matches the cathedral. The "better prose machine" story is, on this evidence, spent.

What is not spent is WHITMAN as an instrument. Pointed at five platforms with a twelve-question factual-control set and read by two scorers, WHITMAN's effect turns out to be sharply platform-specific — and that is a real finding about the models, not about the prompt. Most of them aim the engagement: the lens rises on questions that invite reframing and stays flat on plain factual ones. Claude is the exception — it sprays engagement everywhere, including questions that never asked for it (confirmed by both readers). Grok does the opposite, going quieter on factual controls. Perplexity barely responds at all.

Scatter plot: each model plotted by its engagement lift on reframe-eligible questions (x-axis) versus on factual null controls (y-axis), with two reader scores per model joined by a line. Gemini, GPT-5.5, Grok and Perplexity sit at or below the zero null line (aimed or suppressed); Claude sits alone high in the spray zone, both readers agreeing.
Engagement Selectivity. Horizontal: lift where engagement belongs. Vertical: lift on factual controls, where it does not — low is selective, high is spray. Open marker = first reader, filled = second. Four platforms keep their controls at or below zero; only Claude rises into the spray band, and both readers agree. WHITMAN does not raise engagement everywhere — how it raises it is a property of the model underneath.

Read together, the two results relocate the instrument. WHITMAN is not a treatment that makes answers better — a direct request does that about as well. It is a reagent that makes hidden dispositions visible: who aims engagement, who sprays it, who withholds it. The lens measured the engagement layer; the question the project was founded on — whether a model stays true to what is real — lives on a different axis, and is the next instrument.