Back To Publications

INTELLIGENT ARTIFICE QUARTERLY · Autumn 2026

The Machine That Doubts Its Instruments

Inside the Observatory of Structured Reasoning, where a result is not the end of an experiment but the beginning of its prosecution.

Publication: INTELLIGENT ARTIFICE QUARTERLYBy Mara VennAutumn 2026
INTELLIGENT ARTIFICE QUARTERLY cover: The Machine That Doubts Its Instruments

There is a moment in the life of a research institution when skepticism can become indistinguishable from style.

The Observatory of Structured Reasoning has more opportunities than most to cross that line. It has named procedures for challenging its own findings, elaborate custody rules for preserving the history of a run, instruments that interrogate the outputs of other instruments, and an increasingly baroque vocabulary for distinctions that may or may not survive the next experiment. Its archive contains findings that were strengthened, findings that were narrowed, and findings that were parked after the institution itself found reasons to distrust them.

From outside, this can look like epistemology becoming theater.

If every result can be met by another reader, another baseline, another decomposition, another concern about framing or segmentation, what would ever count as arrival? A research practice must be capable not only of revising a claim but of allowing one to stand. Otherwise skepticism ceases to be a method and becomes an immunity from conclusion.

This was the question I carried into the Observatory.

It was not the question I left with.

The Observatory is not invested in a particular outcome. It is invested in investigation.

The result is only the beginning

The Observatory began with an intuition familiar to anyone who has worked closely with language models: answers that look similar can be structurally different, and answers that look different can be carrying the same underlying reasoning. Presentation can conceal uncertainty. Compression can erase the qualification that made a claim honest. A formal structure can create an appearance of rigor that the evidence beneath it does not possess.

The project responded by building instruments.

Some were designed to force distinctions into view. Some compared renderings. Some examined what survives when language is compressed or transformed. Zexel separated reasoning from rendering and required an answer to pass through a structured process before becoming public language. OMEGA attempted to look across instruments rather than treating each experiment as an island.

Over time, however, something happened that is more interesting than the instruments themselves.

The Observatory began testing the tests.

That change is easy to miss because it does not produce a spectacular single result. It appears instead in the institutional behavior surrounding results: frozen protocols, preserved raw outputs, blind readers, versioned corrections, human rulings, negative fixtures, explicit unknowns, and a refusal to let a later interpretation silently rewrite an earlier run.

A candidate finding does not simply graduate because it is interesting.

It acquires adversaries.

How a Claim Earns the Right to Stay. Candidate claims encounter baselines, reader effects, alternative explanations, replication, damage tests, custody checks and human ruling.
How a Claim Earns the Right to Stay. Candidate claims encounter baselines, reader effects, alternative explanations, replication, damage tests, custody checks and human ruling.

The diagram is peculiar when compared with the standard visual grammar of research. Most process diagrams move from question to method to result. The Observatory's continues beyond the result and then bends backward.

Question. Model and instrument. Artifact. Readers. Candidate claim.

Then survival pressure.

Baseline comparison. Null models. Reader effects. Drift and stability. Alternative explanations. Replication. Damage tests. Custody and provenance.

Only after this does a human ruling appear.

And even then the exits include not only confirmation but revision, narrowing, repetition, parking and retirement.

Most of the machinery, in other words, is located after the point at which a result would normally become a slide.

Most of the machinery is located after the point at which a result would normally become a slide.

The cost of not knowing

R.E. Hoyt, the Observatory's founder, gives me a sentence that initially sounds suspiciously like the kind of institutional language designed to survive any result.

"The Observatory is not invested in a particular outcome. It is invested in investigation."

It is a convenient sentence. Almost too convenient.

The problem with such a precept is obvious: everyone prefers to imagine that their inquiry is disinterested. Researchers are rarely rewarded for saying that they hoped their result would survive, that they liked the theory, that the graph was beautiful, or that months of work had made a particular interpretation emotionally expensive to abandon.

The Observatory does not solve that problem. It is plainly capable of enthusiasm. Its instruments acquire personalities. Its projects develop visual mythologies. Findings receive names. Its internal conversations can be exuberant.

What distinguishes the practice is not the absence of attachment.

It is the attempt to make attachment lose jurisdiction.

The archive contains examples.

During work on Zeditor, an instrument for examining what could be removed from model answers, several apparently exciting findings weakened when pooled reader results were disaggregated. A pattern that seemed to suggest that judgment scales "needed company" largely dissolved when individual readers were examined. Aggregate deletion rates that looked stable concealed materially different sets of passages being cut. A proposed equivalence between two reading dimensions turned out to have a thinner evidentiary basis than the headline suggested.

The corrections were not treated as embarrassment to be minimized. They became the next object of study.

The question shifted from How much does the editor remove? to something more consequential:

What does the reader lose?

That sentence changed the project.

Once removal is evaluated from the reader's side, a successful cut is no longer simply a span that an instrument can delete while preserving grammaticality. A caveat can be grammatically expendable and epistemically essential. An example can be unnecessary to the proposition and necessary to the explanation. A heading can carry orientation without carrying factual content.

The instrument had not failed. The definition of success had become inadequate.

That distinction recurs throughout the Observatory's recent work.

When the experiment tests the wrong thing

ReForm began with a deceptively simple question: what happens to meaning when the same textual units are rearranged?

In an early ordering pilot, models received fixed units and were asked to assemble them into coherent responses. The words could not be rewritten. Only sequence could change.

The mechanism worked.

That was the problem.

The task was good at revealing whether explicit facts and relationships could be recovered. It was less capable of testing the reading experience that had motivated the experiment. A model might recover every sentence while changing the apparent commitment of a speaker, the scope of a reservation, or the relationship between competing obligations.

The Observatory's response was not to declare the first run useless. It preserved it as an engineering test.

Then it changed the question.

The next work moved toward texts in which order could alter position, qualification and interpretive force. Unit boundaries themselves became experimental objects, because separating a qualification from its claim could manufacture the very disruption later attributed to rearrangement.

This is where the Observatory's habits begin to look less like excessive caution and more like a form of experimental literacy.

A negative result is not automatically evidence that nothing happened.

A successful pipeline is not automatically evidence that the right thing was measured.

And an instrument that cannot distinguish those cases is not finished merely because it runs.

A poetry engine complicates the story

Poetix is where the Observatory's current direction becomes stranger.

It is a fork of Zexel built for verse. Where Zexel had asked, among other things, what was necessary, Poetix altered the governing search toward transcendence: whether the reasoning had found an object or relation capable of taking the answer somewhere beyond the obvious landing.

The first temptation is to tell a simple story.

Poetix makes better poems.

There is evidence for something in that vicinity. In the current PX-03 read, five models wrote one hundred poems across four conditions: unaided, Zexel poem mode, Poetix poem mode, and Poetix readout-then-poem mode. Five model readers then evaluated the poems blind.

The unaided poems were named least good with striking frequency. Poetix's readout-then-poem condition was named best more often than any other condition.

That is the headline a product team might stop at.

The Observatory did not.

The writer mattered more.

Different models benefited from different conditions. Claude and Grok improved as the process moved toward Poetix RDP. GPT and Gemini did best under Zexel. Qwen's Poetix poem-mode work could score below its unaided poems. The engine effect was real enough to investigate, but it was not uniform enough to describe as a universal upgrade.

Then another signal weakened.

Poetix contains a turning point called the Crux. In the run, some models reported that it fired almost constantly; another rarely reported it at all. Within a writer, whether the Crux fired did not meaningfully track the resulting poem's score.

The flag was describing the writer's behavior at least as much as it was describing the poem.

This matters beyond poetry.

The Observatory has now encountered, in more than one project, a recurring problem: a measurement may partly be a fingerprint of the model performing the measurement.

The instrument is not outside the phenomenon.

Neither is the reader.

The instrument is not outside the phenomenon. Neither is the reader.

The experimental object expands

This is the idea around which the Observatory's recent work increasingly seems to orbit.

A model cannot always be treated as a neutral carrier of an instrument. The same protocol can interact differently with different model architectures. A reader can bring its own thresholds and aesthetic preferences. A segmentation rule can create the units whose behavior it later appears to discover. A reporting format can make a distinction look more stable than it is.

The experimental object therefore expands.

Not:

MODEL

Not even:

MODEL + INSTRUMENT

But something closer to:

MODEL × INSTRUMENT × READER × PROTOCOL

The multiplication sign matters.

These are interactions, not ingredients in a bag.

A finding that survives one model may fail to reproduce in another because the instrument has encountered a different reasoning architecture. A model reader may reward a quality that a human reader does not. A protocol change may alter the apparent behavior of an otherwise unchanged engine.

The Observatory has not demonstrated that this interaction model is the correct general account of language-model experimentation.

It has accumulated enough trouble to make ignoring it increasingly difficult.

The instrument is also under test: interactions among models, instruments, readers and protocols.
The instrument is also under test: interactions among models, instruments, readers and protocols.

The peculiar virtue of public correction

There is a less glamorous feature of the Observatory that may ultimately matter more than its conceptual vocabulary.

It keeps the old versions.

This sounds trivial until one notices how often research communication quietly replaces history.

A corrected chart overwrites the incorrect chart. A revised rule is described as though it had governed the earlier run. A model identity is normalized after the fact. A failed parser disappears from the final table. An exploratory category becomes a validated concept through repetition rather than adjudication.

The Observatory has increasingly designed against this.

Source artifacts remain source artifacts. Extracted text is derivative. Deletion is an event rather than disappearance. A corrected report becomes another version. A historical run is described according to the rules that actually produced it, even if those rules are later rejected.

This is methodological bureaucracy, and bureaucracy deserves some of its bad reputation.

But custody becomes intellectually interesting when instruments evolve quickly.

Without it, improvement can counterfeit replication.

A revised instrument can produce a cleaner result and tempt its designers to read the new result backward into the old experiment. The Observatory's insistence on preserving the historical implementation makes that move harder.

It also produces an unusual research narrative.

Progress is visible not only as better findings but as better reasons for distrusting earlier ones.

A claim marked “withdrawn as stated”: the correction remains part of the record.
A claim marked “withdrawn as stated”: the correction remains part of the record.

Can skepticism become an aesthetic?

The strongest objection remains.

The Observatory has built a culture that rewards finding the next qualification.

That culture can become self-reinforcing.

A distinction can acquire weight merely because it has been named. A small dataset can support an elaborate conceptual vocabulary. Model readers can create the appearance of independent judgment while sharing training histories, stylistic preferences and failure modes. Human review remains scarce. Instruments change quickly. Many of the most intriguing observations remain provisional.

There is also a danger specific to the Observatory's success at self-critique.

If every correction is interpreted as evidence that the method works, the method becomes unfalsifiable by failure.

A wrong result cannot automatically become a victory for the system that discovered it was wrong.

The Observatory's answer must eventually be performance: claims that survive; instruments that discriminate where they should and abstain where they should; replications that preserve an effect; human readers who confirm that a measured distinction corresponds to something they actually experience.

The institution appears to understand this.

Its best recent moves have involved reducing claims rather than decorating them.

Zeditor's deletion rate gave way to reader loss.

ReForm's successful ordering pipeline gave way to the question of whether the task measured interpretation.

Poetix's encouraging aggregate result gave way to writer-by-instrument interaction.

These are not victories in the conventional sense.

They are improvements in the shape of the question.

The laboratory that learned to play

There is another change underway.

The Observatory has become more playful.

This may seem inconsistent with an institution increasingly concerned with custody and verification. In practice, the two developments appear related.

ReForm is being explored as a public combinatorial instrument in which people rearrange fixed phrase tiles and discover how order changes implication. Poetix takes a reasoning architecture and asks it to produce verse. Zeditor has been imagined as a Victorian cutting machine. Instruments acquire physical metaphors because physical metaphors make transformations inspectable.

Play here is not the opposite of rigor.

It is a way of producing variation.

  • Move the phrase.
  • Change the order.
  • Swap the gate.
  • Remove the sentence.
  • Ask another reader.
  • Hold the words constant.
  • Let the arrangement move.

Each gesture creates a counterfactual close enough to the original that the difference can be inspected.

The Observatory's visual culture can make this look whimsical: small mechanical children, impossible instruments, celestial caretakers, manuscripts turned into blocks.

Underneath the imagery is a severe experimental instinct.

Change one thing.

Watch what else moves.

The processing center: claims can survive, narrow, repeat, park or retire.
The processing center: claims can survive, narrow, repeat, park or retire.

What survives

Late in my visit, I return to Hoyt's precept.

"The Observatory is not invested in a particular outcome. It is invested in investigation."

By then it sounds less pristine.

That improves it.

The Observatory is invested in outcomes in the ordinary human sense. It gets excited. It builds around promising distinctions. It names things. It makes pictures of them. It sometimes believes too quickly.

The precept does not describe a psychological state.

It describes a jurisdiction.

Enthusiasm may propose.

It does not get final authority.

The more interesting question is whether the institution can continue enforcing that distinction as its corpus grows, its instruments become public, and its findings begin attracting attention from people who were not present for the arguments that produced them.

That is the point at which methodological culture either becomes infrastructure or becomes branding.

For now, the Observatory's most persuasive result may be neither Zeditor nor ReForm nor Poetix.

It may be the record of what happened when those instruments disappointed the people who built them. The charts were corrected. The claims narrowed. The categories were demoted. The experiment was redesigned. The archive remained.

And then the work continued. That is not proof that the Observatory's findings are right. It is evidence that the institution has understood a harder problem:

A method for producing claims is incomplete until it includes a method for surviving the desire to keep them.