AI Operations Platform
An agent platform that cannot quietly lie about its own state.
Context
- Roledesigned, built, operated; single author
- PeriodApril 2026 – present
- Statusrunning in a private environment I operate
Since April 2026, I have designed, built, and run an agent platform in a private environment I operate. Its purpose was to run my own work: intake, lead response, outbound, and the dispatch and governance of coding work itself.
It is independently built and operated in a private environment, separate from my work at Mapped.
The relevant part is not that it exists. It is that it was operated — long enough to accumulate an incident history, an append-only decision record, and a set of failures I found in my own system rather than being told about.
Problem
Agent systems fail in ways ordinary software does not.
- Silent successthey report success while producing nothing.
- Plausible fallthroughthey substitute something plausible when the real input is missing.
- Governance driftdocumentation describes controls that no longer exist, or never did.
So the engineering problem was never "make an agent work." It was: make an agent system that cannot quietly lie about its own state.
Constraints
- Single author.No second engineer to catch a bad decision — controls had to be inspectable rather than remembered.
- Live while changing.The system ran my own intake and outbound, so it had to keep working while being modified.
- Nothing autonomous outward.Anything with an effect outside the system required a human decision, by design.
- Unbypassable controls.A control the privileged path can bypass is not a control — which ruled out enforcing audit in application code.
Decisions
Each is recorded in the decision log with the alternative it rejected.
- Separate the control plane from the executor.The governance layer grants and records; it never runs work itself. Between grant and terminal callback it holds no in-flight state, so a crash cannot leave it asserting a build is running when nothing is.
- Make failure the default.Absent configuration means a feature is off, not that the system breaks. Crash recovery reports failure rather than inferring success. Claiming a reservation is an atomic compare-and-swap, so two workers cannot both believe they own the same job.
- Make the ledger unfalsifiable at the database layer.Append-only tables are enforced by database triggers rather than by application code or row-level security — because the service role bypasses row-level security, and a control the privileged path can bypass is not a control.
Architecture
- Layer 1 — control plane. Grants, records, never executes. Command router, approval and consent ledger, atomic reservation grant, callback ingestion. Four named agent surfaces — intake, lead response, outbound, SDR — each with run and health, behind a fleet-level health aggregate.
- Human gate. Agent prepares, a person decides, the action is logged append-only, enforced by database trigger.
- Layer 2 — typed execution runtime. Does the work, disposable clones: pull, claim, clone, launch, verify, call back. Module boundaries: reservations, ledger, outbox, queue, workers, operational projection.
- Layer 3 — LLM workflow layer. Model-backed agents, outside the application: SDR, outreach, outbound execution, lead scoring, business intelligence, contact enrichment.
The application repository has no LLM SDK dependency. It orchestrates, governs, and dispatches — to workflow agents and to coding-agent bridges. Calling the whole platform "an LLM application" would be an overclaim, and anyone who opens the repository sees it immediately.
Layer 1 — control plane
Grants, records, never executes.
- command router
- approval and consent ledger
- atomic reservation grant
- callback ingestion
Four named agent surfaces
- intake
- lead response
- outbound
- SDR
each with run and health, behind a fleet-level health aggregate
Human gate
Agent prepares, a person decides, the action is logged append-only, enforced by database trigger.
Layer 2 — typed execution runtime
Does the work, disposable clones
pull, claim, clone, launch, verify, call back
Module boundaries
- reservations
- ledger
- outbox
- queue
- workers
- operational projection
Layer 3 — LLM workflow layer
Model-backed agents, outside the application
- SDR
- outreach
- outbound execution
- lead scoring
- business intelligence
- contact enrichment
Implementation
- Invariants, not counts.Contract tests enforce architectural invariants: a single navigation registry, one approval destination, redirect grammar, ledger hygiene, decision history, route authority.
- Pure modules under test.Business logic extracted out of routes into pure modules, held under an automated unit-test suite, in two tiers — control plane and scheduling domain.
- Append-only decision log.Each record names the implementation choice, the alternative rejected, and the code the decision governs. The direct answer to "did you decide this, or did the model?"
- Change control on live automation.Every live workflow change is backed up before and verified after — paired snapshots, every time.
- Release assurance across the project's repositories.A centralized CI and release-assurance layer: reusable, SHA-pinned GitHub Actions workflows driven by a per-repository release declaration; post-deployment verification of the serving revision — deployed-SHA and health checks against the live hosts; an append-only release-evidence store; and a policy evaluator that judges each revision from stored evidence and detects drift between the CI-approved revision and the one actually serving.
What operating it taught me
The defect I found in my own system. An extraction path had never once succeeded — and it was masked, because the system fell through to scanning its own prompt text and treated that as evidence. So it reported success on every run. I found it while operating the system, not from a test.
The follow-on is the part I actually care about: the replay harness I built to verify the fix was itself overstating its own coverage, and I found that too.
- Documentation drifts fastest.Code that drifts eventually breaks; prose that drifts just quietly becomes false, in the one artifact a reader consults to learn what the controls are. So documentation gets read against the running system rather than trusted — that is where the drift shows up.
- Instrument before you build.No baseline was captured before anything was built, so no before-state exists to compare against.
- Counts drift; structure doesn't.Commits, files, and test counts moved measurably inside a single week — so they are described, not printed. API routes, agent surfaces, and the migration series did not move.
Close
The repositories are private. A walkthrough is how they get shown.