adventureonthewave

production-ai / observability

Agent Observability

Classical monitoring watches services and requests. Agents add a new unit that nobody is watching: the action — what the agent decided, did, spent, and changed.

Agent observability is the instrumentation that makes agent behavior visible, explainable, and billable to a cause.

The four questions

When something involving agents goes wrong in production, someone must be able to answer four questions quickly, from data, not from memory:

  1. What did it do? — the trace of the run: inputs, decisions, tool calls, outputs, changes made.
  2. What did it cost? — tokens, calls, and spend per run, per agent, per day — before the invoice lands.
  3. Why did it do that? — the context and prompts that drove the behavior, retained long enough to diagnose.
  4. Who stopped it? — evidence that guardrails (budgets, retries caps, permission walls) actually fired.

What breaks without it

What good looks like

In practice

On my own rig, every one of the 967 gate reviews from the audited September 2026 week logged its verdict, tokens, and duration to an append-only log — because a review process you cannot audit is a review process you cannot trust. The same principle scales from a coding rig to a production platform: instrument first, then automate.

Figures below come from one audited production week (September 21–28, 2026) across my own GitHub account — merged-PR counts, gate logs, and token accounting, published weekly on the Turbo Rig stats page. They are measurements of my own workflow, not industry averages.

Talk it through with someone who runs this stack on his own systems every day.