Obversary Studios

Independent research organization studying AI behavior as something that can be observed directly.

Brian Moran Obversary Studios LLC May 2026
Anscombe's quartet: four datasets with identical summary statistics — computed live from the plotted points — and four entirely different shapes. Data from Anscombe, F. J. (1973), "Graphs in Statistical Analysis," The American Statistician 27(1), 17–21.

About Obversary Studios

Obversary Studios is an independent research organization studying AI behavior as something that can be observed directly.

Interpretability looks inside a model — weights, activations, circuits. Evaluation looks at what comes out — scores, benchmarks, pass rates. What a system actually did, step by step and on what basis, sits between those two and is barely instrumented. We measure it by its extremes, when something breaks, or by its aggregates, when a number moves. Rarely as a thing in itself.

That layer is the work here. Founded 2023, and organized as a studio rather than a lab because the output is meant to be seen, not only read.

The commitment is straightforward: build AI systems whose memory, mistakes, and judgment stay observable, so behavior can be examined rather than inferred from an outcome.

Those three are not a list. They are the structure of the problem.

Memory

What a system carries forward, and what it quietly drops. Behavior over time is unreadable without a record of what the system was holding when it acted.

Flow of experiments through source, prompt strategy, and model family — what is carried forward between stages, drawn as width. Demonstration chart with illustrative data, shown as an example of the instrument.

Judgment

What a system chose, and what it chose among. A decision is only legible next to the alternatives that were available and rejected. Most evaluation records the path taken and discards the neighborhood it was taken from.

What one coding agent actually chose, next to what it chose among: 4,605 consecutive action pairs from 60 real sessions of a single model (Fable 5) in a single harness, drawn as flows from each action to the one that followed it. Aggregated from Glint-Research/Fable-5-traces (AGPL-3.0) — observed behavior of one system, not a capability claim.

Mistakes

Where the other two become visible. Failure is not the subject. It is the condition under which the subject stops hiding.

Row-normalized confusion matrices, baseline (left) and delta (right) — where errors pool is visible as structure, not a scalar. Demonstration chart with illustrative data, shown as an example of the instrument.

Everything published here is an instrument for one of these: benchmarks and harnesses built to expose structure rather than produce a number, tracing for multi-step agents, memory substrates that can be inspected, and prompt-injection research that treats adversarial input as a mappable surface rather than a catalogue of exploits.

The first result

The program is behavior. The first claim staked against it is narrow: failure has geometry. Errors are not a list of incidents but a field with shape — they pool in particular places, they have preconditions, and whether those preconditions are observable before the break is a question with an answer rather than a rhetorical flourish.

That question is being tested against public step-level annotation data. If the answer turns out to be no, that will be published too.

A field with shape: a point cloud whose structure — arms, density, a center — is visible at a glance and invisible to any aggregate. Demonstration chart with generated data, shown as an example of the instrument.

Why it’s a studio

A lab produces findings. A studio produces things made to be looked at.

The work here is rendered, not only written. Behavior gets drawn — as fields, traces, distributions, surfaces — because structure is invisible to a summary statistic and obvious to the eye. Anscombe demonstrated this in 1973 with four datasets sharing identical means, variances, correlations, and regression lines, and four entirely different shapes. Tufte made the same argument about the Challenger O-rings, where the data existed and the plot did not. The score is the summary statistic. The render is the structure.

Reading

Documentation

Article index and published notes.

Enter the documentation

Projects overview

Map of current public work.

Open projects overview

Map

How repos, articles, and practice connect.

Open map