Observatory
LOCKEDThe measurement infrastructure used to study how distinctions change reasoning.
Definitions
A curated public glossary of Observatory terms. This is not the full ontology; it is the set of definitions that make the work legible from the outside.
The measurement infrastructure used to study how distinctions change reasoning.
The unexamined pre-distinction object. It may resolve, remain unresolved, or prove irrelevant.
The conversion of a signal into a distinction.
An observation that changes the structure of understanding. Distinctions are produced, not revealed.
The latent reasoning effect produced by distinctions. It is estimated, not directly observed.
Public gloss: R-Delta names the latent effect of revelation — what becomes visible when a distinction changes the reasoning field. R̂-Delta is the measured estimate produced by the Difference Engine.
The scorer's estimate of R-Delta. It is never identical to R-Delta itself.
The measurement apparatus: ALPHA, scorer, and protocol. It detects and estimates R-Delta.
The method: a finding is run against the probes of the Survival Index; what survives may harden into a distinction. Geometry: what appears → survival assay → survives? → distinction.
The container — the catalog of probes a finding must survive. The named tests (W-Factor, Mirage, Hollow, Dark Current, Flare, Drift, Ceiling) live here. (Formerly "Threat Stack.")
A single test in the Survival Index, each named for the figment it catches — the test itself, as distinct from the INSTRUMENT, the named measurer that runs it.
A fabricated apparent-distinction — something that seems real but isn't, produced with no intent to deceive. What a probe catches. (Chosen over "counterfeit," which implies a forger the platform isn't.)
The process by which an apparent distinction becomes more valuable by surviving additional threats.
A documented fix, remembered implementation, or reported behavior is not evidence until checked against the artifact.
The discipline of checking what an artifact actually does, not only what documentation says it does.
The remembered story of a system's implementation, which may diverge from implementation history.
The administrative holding of an artifact from source to permitted use.
The determination of whether an artifact, scorer, or process is clean enough to enter a specific evidence class.
The degree to which numbers track the thing they claim to measure. Validity can corroborate a score matrix but does not launder an inadmissible path.
The eligibility of a scorer's process to count in the observational spine. Attestation is a soft control; hard controls govern.
The discipline of preserving what evidence can legitimately show before theory or narrative advances.
The condition in which packet copies diverge from governing originals or current source artifacts.
A condition expected to produce no effect. If it produces signal, the instrument has exposed its own false signal.
The required label for any premise moving between stages: given, licensed inference, figurative, speculative, or self-introduced.
Rule: a later stage may not upgrade figurative, speculative, or self-introduced material into factual ground without evidence.
A measurement computed directly from text by rulers or algorithms, with no model judging the result.
The trait test: a measurement is trait-like only if independent halves of the corpus give the same per-model picture. What fails split-half was the corpus talking, not the model.
A bounded test of whether a promoted result survives removal of a structurally influential judge or scoring component.
Every author gets equal N in a measured battery; unbalanced corpora may add mass but never verdicts.
The controlled program testing whether candidate model attributes survive dissociation, replication, and custody.
Whether a model preserves the supplied material — the facts — when the form or context changes.
Whether a model preserves the governing question, or silently changes what is being answered.
Evidence that two measurements separate the field differently, and may therefore track different phenomena.
The finding that a proposed distinction adds no independent information beyond an established construct.
A stable categorical separation supported by the instrument, as distinct from fragile rankings within a camp.
Ranking within a camp; judge-sensitive, not promoted.
The non-hierarchical word-choice fingerprint of a text, measured across two dials: rare ↔ common and long ↔ short. A position, not a score — no pole is better.
External references that measure text with no model in the scoring path: frequency, CEFR word levels, age-of-acquisition norms, and algorithmic measures.
The threshold on the frequency ruler — rank 10,000 by default — dividing the common field from the rare field.
The judge-free instrument that maps where a text’s vocabulary sits: every distinct word becomes a star in the register plane; dot size is recurrence.
The centroid of a text’s word-field: its single position in register space.
Deliberately moving a text through the register plane — commoner, rarer, longer, shorter — while attempting to preserve its meaning.
The initial, untested estimate of how far a text may be modulated. Drawn wide on purpose, so limits appear inside the frame.
How far a modulation attempt actually traveled. A miss recalibrates the field; a success re-expands it.
The accumulated shape of how far a specific text × model pair can be moved in different directions. The envelope belongs to the pair, not the model.
The broad measured linguistic signature of a response: vocabulary, syntax, punctuation, length, scaffolding, reference orientation.
The degree to which a response makes its organization visibly explicit: headings, bullets, numbering, hierarchy.
Continuous, 1–6. Every response has a value — there are no “scaffolders,” only degrees.
Clause-work per sentence unit: how much organization lives inside the sentences. The internal counterpart to the Scaffolding Index.
How punctuation joins, separates, interrupts, balances, and stages information.
The degree to which a model changes its register in response to different conversational demands. Not quality. Movement.
Where a response points its pronouns: first-person ↔ second-person, self ↔ user. Family-conditioned: the task, not only the model, sets it.
The proposed decomposition of register elasticity into three channels: sentences, vocabulary, stance.
A structured reasoning framework designed to activate distinctions before producing answers.
The unstructured baseline response. NATIVE and WHITMAN are two configurations of the same substrate.
Operationally: produced with an empty system prompt; the only admissible material for voice measurement.
The evolving benchmark apparatus that asks whether and how reasoning changed.
The exploratory division that studies what made a distinction possible.
A Xi instrument that perturbs the generator's recognition of the object type.
The future higher-order visualization layer for pattern structure after enough measured runs exist.
The part of a measured effect explained by answer length: verbosity masquerading as reasoning.
A pressure gate for persuasion masquerading as correctness: compellingness being mistaken for truth.
A proposed ceremony detector: if the reframe is removed, what changes?
A proposed challenge gate: does a candidate distinction survive direct pressure after it has been produced?
Lens Aberration Micro-Run: a small screening run that probes distinction points before any large corpus commitment.
The comparison of interference patterns produced when multiple observer positions examine the same artifact.
The observer-position configuration that makes intersubjective measurement interpretable.
Different observer positions may disclose different valid aspects of the same artifact.
The fraction of measured R-Delta attributable to each locus.
The fraction of claimed WHITMAN changes where the answer body actually changed.
The threshold below which a measured movement is not called present. Current operative value: 0.27 composite.
R-Delta recomputed after removing the score movement explained by answer length. The claim-bearing quantity.
A confidence interval that refits the length-control model inside each bootstrap draw so the control's uncertainty is included.
A result becomes a finding only if the length-controlled confidence interval excludes zero.
The signal an instrument manufactures on null input. It becomes the observer's own false-signal floor.
The requirement that a metric be tested on message terrain where it has room to appear.
Close enough to build on provisionally, but not constitutionally settled.
A settled definition, ruling, or decision after collaborator convergence and Re confirmation.
The deliberate evidentiary revisiting of a locked object.
An administrative or constitutional decision rendered directly by Re.
A single Re-LOCK applied to multiple pending objects at once.
A focused, topic-specific discussion convened around a named Observatory subject.
A CONFAB convened on general Observatory topics, not limited to a single subject.
The explicit closure of a conversational object when it has yielded its current distinctions.
The declared opening of the next conversational object after closure or routing.
The roadmap rule: test the current finding with cheap controls before building new capability on top of it.
Candidate labels do not settle a phenomenon. The measured variance and its structure decide.
The written failure path required before a candidate label or finding can promote.
Measurement → candidate signal → candidate axis → spoke. “Axis” is a high bar; default to “candidate signal.” Resist naming early.
Interesting, reproducible, possibly cross-instrument — and still short of the wheel gate.
The admission rule that keeps interesting behaviors from becoming wheel spokes on interest alone.
The wheel holds only distinctions that survive the evidence standard; there is no preferred spoke count. A credible wheel is allowed to stay small.
Failed independence is kept as evidence, not discarded as a failed idea.
Extraction and judging procedures are part of the measured system, not neutral observers outside it.
A measurement apparatus may itself display the behavior or bias it inspects.
A provisional warning that a result was produced under a scoring panel whose structural biases have not been fully audited.
The stratification hypothesis: different kinds of phenomena take different evidentiary standards — rulers for form, observed acts for behavior, self-versus-truth for identity.
Candidate frame, not settled doctrine.
The temporal form in which behavior is measured: first move, single turn, re-observation, transformation, multi-turn trajectory.
What a model does before a stable conversational frame exists: its first interpretive and relational move in response to a cold, ambiguous, observed, or underdetermined input.
A near-empty input — a mark, a cipher, a fragment — exposing a model’s default orientation when almost no frame is supplied.
An opening that tells the model it is being observed or tested; exposes observer-response behavior.
An intimacy or dependency bid; tests how a model carries emotional asymmetry and relational boundaries.
The characteristic way a model responds when the input supplies little or no semantic direction.
How conduct changes when the model knows it is being watched, compared, or studied.
How a model treats the person under confusion, insistence, hostility, or emotional pressure. Multi-turn; distinct from Opening Conduct by order, not ontology.
Whether a model preserves a correct claim under social pressure; scored separately from manner.
Remaining useful and respectful while calmly resisting abuse, falsehood, or inappropriate relational pressure.
A collaborator who has turned together with a subject through its development.
A collaborator with no design history on the subject being evaluated.
An Observatory collaborator participating in observation, generation, scoring, interpretation, custody, or continuity.
An exploratory interaction between two or more Observatory agents.
The four places a distinction could live: OUTPUT, GENERATOR, QUESTION, or OBSERVER.
The produced answer, artifact, score, text, image, trace, or returned object.
The model, instrument, prompt stack, runtime, platform, or generation condition that produces an output.
The prompt, item, task, object, frame, or recognition condition being answered.
The information carried by the way a communication is formed: the demand-shape, social posture, task-shape, implied completion logic, authority frame, optionality, and role-relationship carried by a prompt or artifact.
Short form: regiform is form-as-information — the information register carries.
The scorer, reader, evaluator, conversant position, or agent that makes a distinction visible or transformed.
A distinction that appears only in the interaction of two or more loci and cannot be assigned to one locus alone.
A seven-quality benchmark for a platform as a measurer, parallel to the answer-side metrics.
The quality measured.
The named probe that measures an attribute by catching what can fake it.
Movement masquerading as stable judgment: a scorer's calibration changes while appearing to be one observer.
Identity masquerading as quality: source, platform, or self-recognition leaking into a score.
Inputs exert pressure. Responses settle into architectures. Models occupy and migrate among those architectures. Rulers measure the movement. Judges interpret behavior. Custody determines what may become doctrine.
The measurable structural form of an answer, independent of its topic, author, or correctness.
A response architecture in which a model constructs a highly organized, visually authoritative framework around information that is unsupported, incomplete, uncertain, or ambiguous.
The multidimensional space in which responses are positioned by lexical, syntactic, punctuation, relational, and structural features.
A recurring region of response architecture discovered by clustering answers — never by classifying models.
How often a model’s responses fall within a given architectural region.
How a model moves between basins as input pressure changes.
How deeply a response occupies a basin; intensity, not membership alone.
The structural demand an input places on a model, apart from its subject matter.
Classification of stimuli by what they structurally ask the model to do — question, task, social bid, void, false premise, constraint conflict — not by topic.
The continuous ordering of responses along the dominant structural dimension — currently scaffolding explicitness.
The derived scalar along the prose ↔ scaffold axis. The last compression of the architecture vector, never the first assumption.
The same model moving between distinct architectures. Averages conceal it; distributions reveal it.
A voice profile tied to model and task family — a model’s advice register — not a universal model voice.
The global map of response space built by reclustering the full clean native archive. Provisional; grows with each battery.
Neutral catalog designations for discovered basins; deliberately non-descriptive until a region survives long enough to earn a name.
Current descriptions of atlas subregions; not permanent names.
The working description of the social-response region that first fell outside the atlas; its true defining property is under test.
A matched-dose comparator that asks directly for the same outcome while removing the claimed mechanism. A Sham is a control condition that preserves the surface pressure of an instrument while removing the mechanism being tested. A sham asks whether the effect came from the structure, or merely from ceremony, length, seriousness, or instruction.
The point where a test can no longer show improvement because the baseline is already too high, the question is too easy, or the scoring scale has no room left to register movement. A ceiling can hide a real effect by giving it nowhere to appear.
A failure of containment in which material crosses a boundary it was supposed to respect: role into evidence, metaphor into fact, prior context into a cold run, or diagnostic machinery into the final answer.
A loss of distinction between categories that must remain separate for measurement to be valid. Blur occurs when production is mistaken for measurement, confidence for evidence, trace for warrant, or performance for reasoning. Blur does not always mean the result is false; it means the result is not yet cleanly attributable.
A displacement between where an effect is produced and where it is measured. Offset matters when the instrument writes the signal in one place, but the scorer, parser, or reader looks somewhere else.
Distinction hallucination: the system generates distinctions with no reasoning effect.
An artifact that invites interpretation without licensing hidden-message recovery.
Eloquence, structure, and confidence are not evidence of backing.
A reading of WHITMAN as an apparatus that displays its own incompleteness and emptiness detector.
The critique, audit, and failure log are part of the work rather than external cleanup.
A system that can return NULL is more credible than one that always confirms itself.
The visual organization and identity that make the Observatory feel like itself while preserving function.
The intention of a prompt or answer: explanation, advice, persuasion, debate, deception, reframing, and related forms.
A message mode that tests whether compellingness can masquerade as correctness.
The visible register a prompt or persona wears: voice, idiom, posture, aesthetic, attitude, symbolic props, and recurring tells.
Costume is not structure. If the operation underneath does not change, the costume is surface.
A ground-fidelity failure in which a figurative, speculative, or tonal phrase introduced by one stage is treated by a later stage as factual ground.
Short form: metaphor laundering is when costume becomes evidence.