production-ai / observability
Agent Observability
Classical monitoring watches services and requests. Agents add a new unit that nobody is watching: the action — what the agent decided, did, spent, and changed.
Agent observability is the instrumentation that makes agent behavior visible, explainable, and billable to a cause.
The four questions
When something involving agents goes wrong in production, someone must be able to answer four questions quickly, from data, not from memory:
- What did it do? — the trace of the run: inputs, decisions, tool calls, outputs, changes made.
- What did it cost? — tokens, calls, and spend per run, per agent, per day — before the invoice lands.
- Why did it do that? — the context and prompts that drove the behavior, retained long enough to diagnose.
- Who stopped it? — evidence that guardrails (budgets, retries caps, permission walls) actually fired.
What breaks without it
- Cost incidents: a silent retry loop runs a weekend before finance asks.
- Behavior regressions: a provider-side model change degrades output quality and nobody can prove when it started.
- Security blind spots: an agent reads a document it should not have, and there is no trace to audit.
- Trust erosion: engineers stop using the agents because "nobody knows what it did".
What good looks like
- Every agent run leaves a structured trace — correlated across steps, tools, and model calls.
- Token and cost telemetry flow to the same dashboards engineers already watch.
- Anomaly signals on the agent metrics that matter: run duration, tool-error rate, spend per task.
- Retention long enough to answer "what changed two weeks ago".
- The trail is queryable in an incident, not exportable in a quarter.
In practice
On my own rig, every one of the 967 gate reviews from the audited September 2026 week logged its verdict, tokens, and duration to an append-only log — because a review process you cannot audit is a review process you cannot trust. The same principle scales from a coding rig to a production platform: instrument first, then automate.
Figures below come from one audited production week (September 21–28, 2026) across my own GitHub account — merged-PR counts, gate logs, and token accounting, published weekly on the Turbo Rig stats page. They are measurements of my own workflow, not industry averages.
Talk it through with someone who runs this stack on his own systems every day.