Back To Instruments

Project

ALPHA

The scoring and admissibility apparatus for turning differences between answers into evidence.

Status: active evaluation layer Layer: scorer Primary question: what difference survives?

Purpose

ALPHA exists to prevent the Observatory from mistaking compelling language for measured effect. It scores outputs, freezes evidence, records exclusions, and turns response differences into R-Delta signatures.

Its task is survivability under scrutiny: whether a distinction remains visible after blinding, scoring, provenance checks, and cross-rater pressure.

Summary

ALPHA 2.0/2.1 shifts the emphasis from a single outcome class to a profile of deltas, eligibility decisions, and environment metadata. It treats platform, browser, session state, tool/search status, and scorer conversance as part of the evidence chain.

Preliminary Findings

PRELIMINARY · CANDIDATE FINDINGS · PENDING VERIFICATION

  • The environment front was below ALPHA's original resolution: HOST, BROWSER, PLATFORM, and CONVERSANT_STATE were not yet visible distinctions.
  • Interface artifacts can become evidence. Grok's follow-up links independently surfaced Set C's Goodhart and Kahneman fault lines.
  • Functional WHITMAN verification requires more than upload success; the run must show the expected Goal Legitimacy behavior and output structure.
  • The true non-conversant baseline matters. Grok non-conversant Native is identified as a cleaner zero point than any conversant condition.
  • Refreshing a browser is not a clean session. Purging browser state or using a fresh profile is required to control conversance.
  • Gemini's visible "Exploring..." status may expose platform framing or tool/search augmentation, making it provenance rather than scored response content.

Update · 2026-06-12 · C

The W-Factor: the apparatus caught itself

PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION

ALPHA exists to stop the Observatory from mistaking compelling language for measured effect. The first thing it caught was itself. Across Set C, the apparent benefit of structured reasoning tracked, almost perfectly, how much longer the answers had become. The scorers were rewarding length. We named the effect the W-Factor — the verbosity lever — and built a control that removes the length advantage before any effect is read.

Dot plot by platform: a large apparent effect (hollow circle) shrinks to a small effect with confidence interval once length is removed; on Codex and Gemini it collapses to zero
For each platform: the hollow circle is the apparent gain; the filled dot and bar is what remains net of length (95% CI). On three platforms a real effect survives; on two it collapses to zero. The big swings were mostly the verbosity lever.

The honest sentence for Set C

The canonical reading is now: net of length, structured reasoning is a small drag on one platform and roughly nothing on the rest. Set C is therefore the first clean demonstration of the W-Factor — not a demonstration of a structured-reasoning effect. The original headline is annotated, not erased: history preserved, inference updated.

Admissibility became its own axis

Set D forced a second discipline. A scorer can produce numbers that correlate well and still be inadmissible — reached by a tangled process, or by a model that fabricates a finished-looking sheet. So ALPHA now separates two questions that used to be one: is the number valid (does it track truth?) and is the scorer admissible (is the observer and process clean?). Correlation certifies the numbers; it never certifies the scorer. A well-correlated fabrication is still out.

See the survival index for the full method this feeds, and Instruments for the surface index of active records.