production-ai / agentops
AgentOps
AgentOps is the operations discipline for systems whose components are AI agents: keeping agents reliable, observable, governed, and affordable once they run in production.
It is what DevOps was to servers and MLOps was to models — applied to the new moving part: the agent that acts.
The short definition
An agent is software that decides and acts with partial autonomy: it plans, calls tools, spends money, and changes systems. Operations was built for software whose behavior is specified. AgentOps is operations rebuilt for software whose behavior is emergent — you govern outcomes and boundaries instead of instructions.
Practically, AgentOps covers four things: observability (traces of what agents did), governance (permissions, approval boundaries, human merge authority), cost management (token budgets and spend telemetry), and incident response (what to do when an agent misbehaves at scale).
AgentOps vs MLOps
MLOps manages the lifecycle of models: training data, versioning, evaluation, deployment, drift monitoring. Its unit of work is the model artifact. AgentOps assumes the model is a commodity you call, and manages the agent: its tools, memory, permissions, and traces. If your problem is "the model quality regressed", that is MLOps. If your problem is "the agent looped all night and rewrote the wrong file", that is AgentOps.
AgentOps vs DevOps
DevOps gave us the culture and tooling for reliable delivery of deterministic software: CI/CD, infrastructure as code, SLOs. All of that still applies — agents ship through pipelines too. AgentOps adds the parts DevOps never needed: reviewing code written by a non-human author, auditing autonomous actions after the fact, and budgeting a compute cost that varies with the agent's judgment, not the traffic.
AgentOps vs AIOps
AIOps is using AI to operate infrastructure — anomaly detection, event correlation, noise reduction on alerts. AgentOps is the inverse problem: operating the AI itself. The two compose: an AgentOps platform may well use AIOps techniques on agent telemetry. But "AI watching the servers" and "us watching the AI" are different jobs with different failure modes.
The practices, concretely
- Every agent run leaves a trace: inputs, decisions, tool calls, spend.
- Agents run under least-privilege credentials, scoped per tool.
- Budgets and retry caps are enforced by the platform, not by prompts.
- AI-written code passes a review gate — cross-model, then human merge.
- Agent SLOs exist: success rate, cost per task, intervention rate.
- Incident runbooks cover agent-specific failures: loops, context poisoning, silent regressions.
Where the evidence lives
This site practices a slice of AgentOps in public: my own coding rig publishes its gate logs and weekly metrics (967 gate reviews and 6.57B tokens in the audited week of September 21–28, 2026). The Turbo Rig case study walks through the architecture; Ancuria shows the same discipline applied to a production SaaS.
Talk it through with someone who runs this stack on his own systems every day.