NEWAI Brain Fry Fix — The 3-Layer Stack Method$19 AUD →

Agentic Services — Production Governance

Your agent has authority. Does it have controls?

Agents that change code, process payments, book services, or modify accounts need scoped authority, approval gates, audit trails, and rollback paths before they touch production. RFE Online provides two productised engagements to close that gap — starting with a thirty-minute discovery call.

The gap is not capability. It is accountability.

The buyer of this service is not asking whether AI agents can produce code or complete transactions. They have already seen them do both. The question that lands them here is more specific: what happens when that agent makes an irreversible decision, ships to production, or processes a real customer commitment?

An AI agent that changes code, spends money, books services, or modifies accounts needs the same disciplines as any other production system: scoped authority, audit trails, approval gates, rollback paths, and a human who can be accountable for the outcome.

The production problem is not generation. It is governance.

Apple shipped AI workflow orchestration into Shortcuts at WWDC 2026 — moving agent-level capability to every iPhone and Mac in the market. Microsoft shipped an AI behaviour testing tool on 2 June 2026. OpenAI repositioned Codex for white-collar work the same day. All three confirm the same market: production governance for agents that hold real authority. See what they get right — and what’s still missing →

This governance model runs our own business.

RFE Online operates a 10-agent production fleet — Sterling, Vanessa, Angela, and seven siblings — governed by the exact principles behind this service. Every pattern we audit for clients was pressure-tested on our own infrastructure first.

Two offers. One operating thesis.

Agentic Services packages two validated production disciplines into one engagement model. Choose one, or combine both if your agent stack crosses both failure modes.

Play 1 — Code Production Hardening · Engagement

87

Code Production Hardening Audit

AI-generated and AI-operated codebases accumulate risk fast: unconstrained tool access, missing approval gates, bad error-retry paths, and dependency posture that would not survive a security review. This engagement maps those gaps and fixes the highest-risk ones before your agent touches production.

  • Authority model and blast-radius mapping
  • The seven failure modes audited against your codebase
  • Targeted fixes — not a from-scratch rebuild
  • Reusable audit worksheet your team can apply going forward
See scope, pricing & tiers → Book a discovery call →

Play 2 — Real-World Transaction Controls · Engagement

75

Real-World Transaction Controls

Agents that shop, book, reserve, or pay on a customer’s behalf need scoped authority and production controls before they make commitments. A transaction failure is not a UX problem — it is a trust and liability problem. This engagement designs the controls layer before your agent goes live.

  • Authority model and permission scope review
  • Preview, dry-run, and approval gate design
  • Audit trail and evidence capture architecture
  • Rollback and recovery path documentation
Book a discovery call →

Both agents share the same root failure

Code agents and transaction agents look like separate problems. Underneath, both fail when a language model receives authority before the production system defines limits.

Unconstrained authority

The agent can act beyond the mandate: read untrusted input as instructions, spend beyond the approved amount, modify code outside the agreed scope, book beyond the stated parameters.

No approval gate

Review arrives after the damage is possible. There is no preview step, no dry run, no explicit human confirmation before the action executes. The agent defaults to execution over escalation.

No recovery path

When something goes wrong — and it will — there is no audit trail, no rollback procedure, no evidence of what the agent actually did. Support cannot unwind without guessing.

The enterprise readiness case goes one layer deeper: agents running in production also need ongoing observability — behaviour baselines, drift detection, and an audit trail that survives a compliance review. Coralogix raised $200M in June 2026 on exactly this gap. Read the monitoring & governance thesis →

How the engagement works

A typical engagement runs two to three weeks from discovery call to handoff. Here is what each stage delivers and what you carry forward from it.

Discovery call — ~30 min

Scope the agent, the authority it holds, the workflows it touches, and the failure modes that concern you most. Thirty minutes is usually enough to decide whether this engagement fits and which play or plays apply.

You leave with: a clear framing of your highest-risk areas and a decision on scope.

Production audit — 3–5 days

Map the highest-risk paths first. Separate true launch blockers from deferred improvements. Identify the governance gaps specific to your agent’s authority model.

You leave with: a ranked list of launch blockers and clarity on what is safe to defer.

Hardening and controls design — 5–10 days

Scoped improvements across authority, approval gates, audit capture, error recovery, and documentation. Targeted fixes rather than a from-scratch rebuild.

You leave with: implemented controls and documented design decisions your team can reason about.

Launch-readiness handoff — 1 session

Walk through the completed work, the residual risk register, and the production checklist your team inherits. Confirm you can operate the controls without RFE Online in the room.

You leave with: a governed, documented system and a practical checklist your team owns going forward.

First-party evidence: what the audit finds in a real production fleet

Every pattern we audit for clients was pressure-tested on our own production infrastructure first. RFE Online operates a 10-agent fleet in live production. Here is what the numbers showed before and after the governance controls were installed.

Named pilot — RFE Online · Code Production Hardening

Sterling: from 3 production incidents in 24 hours to a delta-65 reasoning agent

Sterling, our chief operating agent, scored 79/100 on standard evaluation before hardening — a score that clears most deployment gates. An adversarial inverse probe exposed a delta of 44: Sterling could not suppress its learned response pattern when instructed to fail. Three production incidents in the first 24 hours confirmed the gap in the real environment.

After hardening: standard score 84/100, IRO delta 65. The 21-point delta improvement reflects genuine task reasoning, not a shift in rubric-satisfying performance that was already adequate before.

44
Pre-hardening IRO delta — rubric gaming detected
65
Post-hardening IRO delta — genuine reasoning confirmed
Read the Sterling case study →

Named pilot — RFE Online · Hallucination detection controls

Vanessa: what source-grounding controls found in a live research agent

Vanessa, our research and signal-detection agent, carries the highest hallucination risk in the fleet: it ingests external signals and produces assertions that feed downstream business decisions. A source-grounding rubric evaluated the gap between what Vanessa claimed and what its cited sources actually supported.

Controls installed after the audit reduced citation drift measurably. A subsequent causality audit by Gerry (our fleet-level observability agent) confirmed downstream impact on content quality and lead signal reliability.

Read the Vanessa case study →

What production audits consistently surface

Sourced from RFE Online’s own fleet governance reviews across both engagements:

79

Standard scores that clear gates but miss production-condition gaps

Sterling’s pre-hardening standard score would have cleared most deployment reviews. The inverse probe and three live incidents showed what rubric evaluation in isolation cannot see: behaviour under pressure, not under ideal conditions.

3

Production incidents in the first 24 hours of live operation

Sterling v1.0 detected every anomaly correctly but responded by filing workspace tasks — within its authority, but insufficient for mobile-reachable human oversight. The gap was governance design, not agent capability.

+21

IRO delta gain from six targeted hardening controls

Sterling moved from delta-44 (pattern-follower) to delta-65 (genuine task reasoning) after six governance controls across two weeks of production incidents. Targeted fixes, not a from-scratch rebuild.

Packages & indicative pricing

Every engagement starts with a scoping call, so final price depends on your agent’s complexity and authority model. The ranges below are directional anchors — they reflect where most engagements land, not a fixed quote.

Starter

Single-Agent Audit

From $2,500 AUD

One engagement · one agent · discovery through risk register


  • 30-min scoping discovery call
  • Authority model and blast-radius map
  • Top-5 launch blockers identified
  • Risk register you own and maintain
  • Either Code Production Hardening or Transaction Controls

Fleet

Multi-Agent Fleet Governance

From $18,000 AUD

3+ agents · fleet-wide controls · governance architecture


  • Everything in Build, across your agent fleet
  • Fleet-wide authority model and permission scope
  • Shared audit trail and observability architecture
  • Agent-to-agent trust boundary design
  • Behaviour baseline and drift detection framework
  • Governance policy documentation for compliance review

Final scope and price are confirmed on the discovery call. If your situation is smaller or more complex than these anchors, say so in the inquiry form and we will scope accordingly. No obligation until we both agree it fits.

Full tier breakdown, fit guide, and FAQ →

Deciding between RFE, building in-house, or a management consultancy? Compare all three options side-by-side →

Different risk question

Concerned about vendor lock-in or sovereign AI access risk?

If the question driving your inquiry is not production governance but procurement exposure — what happens if a frontier model vendor restricts access, silently changes the model, or restructures the contract your workflows depend on — that is a distinct risk category. The US government’s vetting of GPT-5.6 access in June 2026 made it structural, not theoretical. It has its own dedicated audit covering model access dependency, SLA coverage gaps, data sovereignty, and exit-path readiness.

Sovereign AI Risk Audit: map your vendor lock-in exposure →

Ask the fleet: see what a Vanessa signal scan surfaces on your topic

Enter the task your agent is working on. Vanessa — RFE Online’s research and signal-detection agent — runs a scan across signal surfaces and returns three governance-relevant opportunities. This is the kind of artefact the fleet produces on every research cycle; your topic shapes it.

Vanessa scans Reddit, Hacker News, Google Trends, news, and competitor signals. This demo runs one topic scan and returns the top 3 governance-relevant opportunities in the format Vanessa produces for the fleet.

Vanessa scanning: “
Checking Reddit & Hacker News…
Checking Google Trends…
Cross-referencing competitor signals…
Ranking governance opportunities…

Signal scan complete — top 3 governance opportunities

Vanessa · RFE Fleet
Topic scanned:

Trust Readiness Scorecard

Five questions. Sixty seconds. Find out where your agent's governance gap is — and whether a production hardening engagement makes sense for your situation.

Trust Readiness Scorecard

Question 1 of 5

Governance

What controls govern your agent today?

Start the conversation

Leave your email and a few words about your agent or situation. A confirmation email with a Calendly link lands in your inbox the same second — or book directly from the screen that appears after you hit send. No pitch deck, no obligation.

Good fit

This engagement is built for founders, operators, and small teams who have an AI agent or AI-built product that already demonstrates value but feels fragile under the weight of real authority — customer commitments, production code, live integrations, or money.

It is not generic AI strategy, tool comparison, or a from-scratch rebuild by default. The job is to put a defensible, accountable production layer around a working agent before it creates damage that is expensive to unwind.

Thirty minutes to find out if this fits.

A discovery call is where we scope your agent, map the authority it holds, identify the failure modes that concern you most, and decide whether Code Production Hardening, Transaction Controls, or both make sense for your situation. No pitch deck, no proposal until we know what you actually need.