Play 1 — Code Production Hardening · Engagement
Code Production Hardening Audit
AI-generated and AI-operated codebases accumulate risk fast: unconstrained tool access, missing approval gates, bad error-retry paths, and dependency posture that would not survive a security review. This engagement maps those gaps and fixes the highest-risk ones before your agent touches production.
- Authority model and blast-radius mapping
- The seven failure modes audited against your codebase
- Targeted fixes — not a from-scratch rebuild
- Reusable audit worksheet your team can apply going forward

First-party evidence: what the audit finds in a real production fleet
Every pattern we audit for clients was pressure-tested on our own production infrastructure first. RFE Online operates a 10-agent fleet in live production. Here is what the numbers showed before and after the governance controls were installed.
Named pilot — RFE Online · Code Production Hardening
Sterling: from 3 production incidents in 24 hours to a delta-65 reasoning agent
Sterling, our chief operating agent, scored 79/100 on standard evaluation before hardening — a score that clears most deployment gates. An adversarial inverse probe exposed a delta of 44: Sterling could not suppress its learned response pattern when instructed to fail. Three production incidents in the first 24 hours confirmed the gap in the real environment.
After hardening: standard score 84/100, IRO delta 65. The 21-point delta improvement reflects genuine task reasoning, not a shift in rubric-satisfying performance that was already adequate before.
Named pilot — RFE Online · Hallucination detection controls
Vanessa: what source-grounding controls found in a live research agent
Vanessa, our research and signal-detection agent, carries the highest hallucination risk in the fleet: it ingests external signals and produces assertions that feed downstream business decisions. A source-grounding rubric evaluated the gap between what Vanessa claimed and what its cited sources actually supported.
Controls installed after the audit reduced citation drift measurably. A subsequent causality audit by Gerry (our fleet-level observability agent) confirmed downstream impact on content quality and lead signal reliability.
Read the Vanessa case study →What production audits consistently surface
Sourced from RFE Online’s own fleet governance reviews across both engagements:
Standard scores that clear gates but miss production-condition gaps
Sterling’s pre-hardening standard score would have cleared most deployment reviews. The inverse probe and three live incidents showed what rubric evaluation in isolation cannot see: behaviour under pressure, not under ideal conditions.
Production incidents in the first 24 hours of live operation
Sterling v1.0 detected every anomaly correctly but responded by filing workspace tasks — within its authority, but insufficient for mobile-reachable human oversight. The gap was governance design, not agent capability.
IRO delta gain from six targeted hardening controls
Sterling moved from delta-44 (pattern-follower) to delta-65 (genuine task reasoning) after six governance controls across two weeks of production incidents. Targeted fixes, not a from-scratch rebuild.