Skip to content
Work
Built Independently

AI Operations Platform

An agent platform that cannot quietly lie about its own state.

Context

  • Roledesigned, built, operated; single author
  • PeriodApril 2026 – present
  • Statusrunning in a private environment I operate

Since April 2026, I have designed, built, and run an agent platform in a private environment I operate. Its purpose was to run my own work: intake, lead response, outbound, and the dispatch and governance of coding work itself.

It is independently built and operated in a private environment, separate from my work at Mapped.

The relevant part is not that it exists. It is that it was operated — long enough to accumulate an incident history, an append-only decision record, and a set of failures I found in my own system rather than being told about.

Problem

Agent systems fail in ways ordinary software does not.

  • Silent successthey report success while producing nothing.
  • Plausible fallthroughthey substitute something plausible when the real input is missing.
  • Governance driftdocumentation describes controls that no longer exist, or never did.

So the engineering problem was never "make an agent work." It was: make an agent system that cannot quietly lie about its own state.

Constraints

  • Single author.No second engineer to catch a bad decision — controls had to be inspectable rather than remembered.
  • Live while changing.The system ran my own intake and outbound, so it had to keep working while being modified.
  • Nothing autonomous outward.Anything with an effect outside the system required a human decision, by design.
  • Unbypassable controls.A control the privileged path can bypass is not a control — which ruled out enforcing audit in application code.

Decisions

Each is recorded in the decision log with the alternative it rejected.

  1. Separate the control plane from the executor.The governance layer grants and records; it never runs work itself. Between grant and terminal callback it holds no in-flight state, so a crash cannot leave it asserting a build is running when nothing is.
  2. Make failure the default.Absent configuration means a feature is off, not that the system breaks. Crash recovery reports failure rather than inferring success. Claiming a reservation is an atomic compare-and-swap, so two workers cannot both believe they own the same job.
  3. Make the ledger unfalsifiable at the database layer.Append-only tables are enforced by database triggers rather than by application code or row-level security — because the service role bypasses row-level security, and a control the privileged path can bypass is not a control.

Architecture

  • Layer 1 — control plane. Grants, records, never executes. Command router, approval and consent ledger, atomic reservation grant, callback ingestion. Four named agent surfaces — intake, lead response, outbound, SDR — each with run and health, behind a fleet-level health aggregate.
  • Human gate. Agent prepares, a person decides, the action is logged append-only, enforced by database trigger.
  • Layer 2 — typed execution runtime. Does the work, disposable clones: pull, claim, clone, launch, verify, call back. Module boundaries: reservations, ledger, outbox, queue, workers, operational projection.
  • Layer 3 — LLM workflow layer. Model-backed agents, outside the application: SDR, outreach, outbound execution, lead scoring, business intelligence, contact enrichment.

The application repository has no LLM SDK dependency. It orchestrates, governs, and dispatches — to workflow agents and to coding-agent bridges. Calling the whole platform "an LLM application" would be an overclaim, and anyone who opens the repository sees it immediately.

  • Layer 1 — control plane

    Grants, records, never executes.

    • command router
    • approval and consent ledger
    • atomic reservation grant
    • callback ingestion

    Four named agent surfaces

    • intake
    • lead response
    • outbound
    • SDR

    each with run and health, behind a fleet-level health aggregate

  • Human gate

    Agent prepares, a person decides, the action is logged append-only, enforced by database trigger.

  • Layer 2 — typed execution runtime

    Does the work, disposable clones

    pull, claim, clone, launch, verify, call back

    Module boundaries

    • reservations
    • ledger
    • outbox
    • queue
    • workers
    • operational projection
  • Layer 3 — LLM workflow layer

    Model-backed agents, outside the application

    • SDR
    • outreach
    • outbound execution
    • lead scoring
    • business intelligence
    • contact enrichment
Three layers and the human gate. The LLM workflow layer sits outside the application.

Implementation

  • Invariants, not counts.Contract tests enforce architectural invariants: a single navigation registry, one approval destination, redirect grammar, ledger hygiene, decision history, route authority.
  • Pure modules under test.Business logic extracted out of routes into pure modules, held under an automated unit-test suite, in two tiers — control plane and scheduling domain.
  • Append-only decision log.Each record names the implementation choice, the alternative rejected, and the code the decision governs. The direct answer to "did you decide this, or did the model?"
  • Change control on live automation.Every live workflow change is backed up before and verified after — paired snapshots, every time.
  • Release assurance across the project's repositories.A centralized CI and release-assurance layer: reusable, SHA-pinned GitHub Actions workflows driven by a per-repository release declaration; post-deployment verification of the serving revision — deployed-SHA and health checks against the live hosts; an append-only release-evidence store; and a policy evaluator that judges each revision from stored evidence and detects drift between the CI-approved revision and the one actually serving.

What operating it taught me

The defect I found in my own system. An extraction path had never once succeeded — and it was masked, because the system fell through to scanning its own prompt text and treated that as evidence. So it reported success on every run. I found it while operating the system, not from a test.

The follow-on is the part I actually care about: the replay harness I built to verify the fix was itself overstating its own coverage, and I found that too.

  • Documentation drifts fastest.Code that drifts eventually breaks; prose that drifts just quietly becomes false, in the one artifact a reader consults to learn what the controls are. So documentation gets read against the running system rather than trusted — that is where the drift shows up.
  • Instrument before you build.No baseline was captured before anything was built, so no before-state exists to compare against.
  • Counts drift; structure doesn't.Commits, files, and test counts moved measurably inside a single week — so they are described, not printed. API routes, agent surfaces, and the migration series did not move.

Close

The repositories are private. A walkthrough is how they get shown.