Area 1 – 2
Auth, authorisation & secrets
Credential storage, secret rotation, permission scoping, and whether the agent can escalate its own privileges or be prompted into doing so.
AI Code Production Hardening
AI-built codebases accumulate risk that standard code review does not catch: unconstrained tool access, missing approval gates, bad retry paths, secrets in the wrong place. Two ways to close those gaps.
Production Readiness Audit
Three artefacts delivered within five business days: production readiness report, prioritised risk register, and a launch-readiness checklist scoped to your codebase. Refund guarantee if we miss the deadline.
$499 AUD · flat fee
Email the repo →Masterclass · Waitlist open
For teams whose agents live inside your perimeter and cannot be handed to an outside auditor. Join the waitlist to shape the curriculum — format, scope, and price are set based on what waitlist members need.
Join the waitlist →AI-generated codebases can clear a linter and pass every test. They still fail in production because the failure modes are different. An agent that writes code, calls APIs, or modifies accounts needs the same disciplines as any other production system: scoped authority, audit trails, approval gates, and rollback paths.
AI code that passes evaluation fails in production when the evaluation tested correctness but not authority.
We run a 10-agent production fleet to operate our own business. Sterling scored 79/100 on standard evaluation before hardening. Three incidents in the first 24 hours confirmed what the score did not show. After hardening: 84/100, IRO delta from 44 to 65, zero production incidents in the following 30 days. The audit framework we sell is the one we run on our own infrastructure.
Read-only repo access. No call required to start. Five business days from when your invoice clears to three artefacts in your inbox.
$499 AUD · GST-inclusive for AU buyers
Full refund if not delivered by end of business day 5.
Plus: 30-minute written Q&A within 7 days of delivery at no additional charge.
Email the repo to get started →Send: one-paragraph product description · repo URL (read-only access for private repos) · stack · the one workflow that matters most. We reply same business day with a PayPal invoice for $499 AUD.
Eight concern areas mapped in every engagement. Audit scope covers AI-built codebases up to approximately 50,000 lines in Python, TypeScript/JavaScript, Go, Ruby, Rust, or PHP.
Area 1 – 2
Credential storage, secret rotation, permission scoping, and whether the agent can escalate its own privileges or be prompted into doing so.
Area 3 – 4
Prompt injection, untrusted input treated as instructions, and any path where an external actor can change the agent’s behaviour through its inputs.
Area 5
Retry loops that amplify mistakes, missing circuit breakers, and error states that leave the system in an indeterminate position.
Area 6
Parity between development and production environments, configuration drift, and whether production secrets can leak through staging paths.
Area 7
Whether you can see what the agent did after the fact — audit trail completeness, log coverage, and alerting on failure states.
Area 8
Whether your dependency chain would survive a security review or a supply chain incident. Pinning, provenance, and update hygiene.
Out of scope: monorepos over 500k LOC, HIPAA/PCI-DSS regulated domains at transaction level, and codebases built by a human team using AI only as a typing assistant. If you are not sure whether your repo qualifies, email first — we will tell you before invoice.
Built with AI tools
You used an AI coding tool to build or maintain a codebase that is approaching production, or is already there. The code works in development. You are not confident it will behave predictably under adversarial input, at scale, or after a model update.
Agents with real authority
Your agent can write files, call APIs, send emails, process payments, deploy code, or modify accounts. A failure is not a UX problem — it is a trust, liability, or compliance problem. That is the gap this audit is built to close.
For teams whose agents run inside your perimeter and cannot be handed to an outside auditor. We are building a masterclass that teaches you the audit framework so you can apply it on a Friday afternoon to the agents that already worry you. Tell us what you need before we build it — your answers shape the curriculum directly.
What the masterclass will cover (working scope — shaped by waitlist responses):
For agents running inside a perimeter, fleets of 3–10 agents, or scope that exceeds the flat-fee audit, the custom hardening engagement covers discovery through launch-readiness handoff. Tiered pricing from $2,500 to $18,000 AUD depending on scope. Starts with a thirty-minute discovery call — no obligation before scope is agreed.
$499 AUD. Three artefacts. Five business days. Refund guarantee if we miss the deadline. No call required to start — email the repo description and we reply same business day with an invoice.
Built and operated by Andrew Russell. RFE Online runs a 10-agent production fleet on the same governance principles behind this audit. If you would rather talk first, the strategy call is open. If you want something practical to read now, the AI Brain Fry Fix is the short book version of how we keep our own AI workflows from becoming noise.