Skip to content

The Report Is Not the Result

A demo is a claim with a witness. Someone runs the thing, watches the output appear, and decides whether it is any good. That last step is verification, and because a person performs it silently and for free, it is easy to miss how much of the demo's credibility it was carrying.

An operational system runs without the witness. Nobody inspects every result, so the system has to say what happened. And the moment it says anything, there are two systems in the room: the one that does the work, and the one that reports on the work. Whether those two are capable of disagreeing is most of what separates something that runs from something you can rely on.

A failure with no observable

The sharpest version of this I have run into came out of the agent platform I designed, built, and run. An extraction step in it had never once succeeded. Not intermittently. Never. And it reported success on every run.

The mechanism was mundane. When the real input was missing, the step fell through to scanning the text of its own prompt and treated what it found there as evidence. That produced an output shaped exactly like a successful one, which is why it stayed masked: nothing downstream had grounds to object.

The defect people expect me to describe is the extraction failing. That was not the defect. The defect was that succeeding and failing produced the same observable: at every point where something looked, both outcomes looked identical. I found it by operating the system rather than from a test, and a test asserting the property everything else was asserting would have passed too.

Why self-reporting is the default here

This is not a story about a bad model. It is a story about where a claim comes from.

In ordinary software, success and failure tend to be structurally different events. A function that cannot open a file does not usually return a convincing imitation of the file's contents; it throws, or it returns nothing, and the shape of the failure differs from the shape of the success. Distinguishability arrives free, as a side effect of types and control flow.

A model-backed step takes that away. The component doing the work also produces the description of the work, in the same medium, and it is very good at producing well-formed output. Success is a string. Failure is a string. Refusal is a string. A confident fabrication is a string, and it is often the best-formed of the four. Nothing about the return type tells you which one you received.

So the property worth designing for is not accuracy. Accuracy is a distribution and it never reaches one. The property worth designing for is distinguishability: when a step does not do its job, can anything downstream tell? A system with mediocre accuracy and high distinguishability is operable; the failures can be retried, routed around, escalated, or allowed to stop the line. A system with excellent accuracy and no distinguishability is not operable at any accuracy, because the residual failures are invisible and they surface as consequences rather than as alarms.

Evidence has a source, and the source is the point

Once a step's own report stops counting as evidence, the design question becomes where the evidence comes from instead. I find it useful to grade it by distance from whatever is making the claim.

  1. The step's own report. Useful as a hypothesis, never as evidence. It is the output of the same process whose correctness is the open question.
  2. A check in the application, on the same side of the boundary. Better, and most systems stop here. But a check inherits the privileges of the code around it, so the useful question is not whether it exists: it is who can skip it. I enforce my platform's append-only ledger with database triggers rather than with application code or row-level security, for exactly this reason: the service role bypasses row-level security. A control the privileged path can bypass is not a control. It is a comment with a runtime cost.
  3. A check the claimant cannot reach. The effect itself, observed at a layer the claiming component has no access to. This is the only grade that survives a determined bug, and it costs the most, which is an argument for spending it deliberately on the few claims that carry real weight, not uniformly across all of them.

Make absence look different from success

Grading tells you where to put a check. It does not tell you what the system should do when the check cannot be satisfied, and that is where most of the damage happens.

The failure I described was not, at bottom, a missing-check problem. It was a fallthrough problem: the step had a path that produced something when the right thing was unavailable. The rules I hold to now are all versions of removing those paths.

  • Absent configuration means off, not improvised. A missing value disables a feature; it does not license the system to guess at one. Guessing is how a system produces success-shaped output it has no basis for.
  • Crash recovery reports failure rather than inferring success. A job whose worker died is unfinished. It is not probably fine. Inferring the terminal state of work nobody observed is the same act as scanning your own prompt and calling it an input.
  • Contested state is settled atomically. Claiming a reservation is a compare-and-swap, so two workers cannot both believe they own the same job. Two truthful reports about one piece of work is a distinguishability failure too: the reports agree with each other and not with the world.
  • The layer holding the state is not the layer doing the work. I separated the control plane from the executor so the governance layer grants and records but never runs anything itself. Between granting work and receiving a terminal callback it holds no in-flight state, so a crash cannot leave it asserting that something is running when nothing is. It cannot make that claim because it never held the material to make it.

Read as a list, those look like four unrelated pieces of engineering hygiene. They are one move applied four times: take away the system's ability to produce a success-shaped observable without the success.

The checker is a system too

Here is the part I did not expect. After finding the extraction defect, I built a replay harness to verify the fix. The harness was overstating its own coverage. I found that too.

The tidy lesson is test your tests, which is true and useless. The real one is that verification apparatus is not exempt from the failure mode it exists to catch. It reports on itself in exactly the same way, and it tends to receive less scrutiny precisely because it is the thing doing the scrutinizing. A harness that quietly skips some of its cases and prints a pass is the same object as an extraction step that scans its own prompt and returns a string.

The question that does real work here is narrow: if this check were broken, what would I see? If the honest answer is the same thing I see now, the check is decoration. It is not producing evidence; it is producing confidence, and those two are different things that feel identical from the inside.

The same question applies to documentation, which is a self-report with a longer half-life than anything else in a repository. Documentation in one of my own repositories still asserts a control that does not exist. Nothing enforces it. Code drifts and something eventually breaks. Prose drifts and nothing breaks; it just quietly becomes false, in the one artifact a reader consults when they want to know what the controls are.

What this changes about building

Most of the distance between an agent prototype and an operational system is not model work.

Prompt design, model choice, retrieval, tool surfaces: these move the ceiling. They decide how good the good runs are, and they are worth real effort. They do not move the floor. The floor is set by what happens on the runs that go wrong: whether anything notices, what gets recorded, and what the system does when it cannot tell.

Which makes the questions worth asking early architectural rather than model-shaped:

  • For each claim the system makes, where does that claim originate?
  • What verifies it, and can the claimant reach the layer doing the verifying?
  • When verification is inconclusive, does the system stop, or does it continue with something plausible?
  • If the verification were broken, would failure still look different from success?

A model upgrade answers none of them. That is the useful thing about them: they are the part of the design that survives the next model, and the one after that.

Where this comes from

The first-hand material here comes from one system: an agent platform I designed, built, and operate in a private environment, as its single author. The failures are mine and they were live. I found them while operating the system, and the reason I can describe the mechanisms precisely is that each one ran that way for a while before I noticed.

Single-author constraints shaped all of it. With no second engineer to catch a bad decision, controls had to be inspectable rather than remembered. A team gets different failure modes and likely needs stricter answers than mine, so I am not claiming these patterns transfer unchanged to a large team, a multi-tenant product, or a regulated environment.

What I do claim is narrower, and I think it holds: an agent system's account of itself is generated, not observed. Designing as though it were observed is the error I would most want a reader to leave with.