production-ai / readiness-audit
AI Production Readiness Audit
Before you give agents production access, find out what they will break.
A readiness audit is a structured review of everything an AI system touches before it goes live — so the expensive lessons happen in a report instead of in your incident channel.
Why before
Every failure mode an agent can hit in production is cheaper to find on paper. Permissions that are too broad, secrets that travel in plaintext, a pipeline with no rollback, a model dependency nobody owns — each of these is a paragraph in an audit and an outage in production.
Teams usually request this audit at one of two moments: right before a pilot goes live, or right after leadership asks why the pilot is "taking so long". Both are good moments. The first is cheaper.
What gets reviewed
- Architecture — what the agents are, what they connect to, where the trust boundaries are.
- Agent permissions — what each agent may read, write, call, and spend. Least privilege, actually checked.
- Tool access — every tool, MCP server, and API an agent can reach, and what happens when each misbehaves.
- Secrets — where keys live, how they travel, what an agent (or a prompt injection) could exfiltrate.
- CI/CD — how AI-written code reaches production, and whether that path has a gate, a reviewer, and a rollback.
- Model dependencies — what breaks when a provider ships a change, deprecates a model, or rate-limits you.
- Context and memory — what the agents remember, where it is stored, and who else can read it.
- Observability — whether anyone can answer "what did the agents do last night" without archaeology.
- Rollback — the honest answer to "an agent broke it, now what".
- Human approval — which actions require a human, and whether that boundary is enforced by the system or by convention.
- Cost controls — budgets, ceilings, and alerts before the invoice, not after.
- Failure modes — retry storms, silent loops, context poisoning, runaway spend, and what contains each one.
What comes out
- Architecture findings — where the system is thin, ranked by blast radius.
- A risk matrix — likelihood × damage, in terms your CTO can act on.
- A production readiness score — defensible, criteria-backed, re-runnable after remediation.
- A prioritized remediation plan — what to fix first and why.
- A 30/60/90-day roadmap — remediation sequenced into engineering-sized steps.
How deep it goes
The audit draws on patterns proven at production scale on my own systems: the same review-gate, human-merge-authority, and observability rules that govern the Turbo Rig coding rig and the Ancuria SaaS. Where a finding needs evidence, you get the mechanism, not a best-practice citation.
The engagement version
Deliverables, timeline, and how to book: Agentic Readiness Audit — scope & deliverables.
Talk it through with someone who runs this stack on his own systems every day.