NEWAI Brain Fry Fix — The 3-Layer Stack Method$19 AUD →

Case Studies — Agentic Services

The evidence behind every engagement

Every pattern RFE Online audits for clients was pressure-tested on its own production fleet first. These are the named pilots, the before-and-after numbers, and the controls that moved them.

Outcome metrics across five pilots

Five agents, five distinct failure modes, measured results from production. The numbers below are sourced from RFE Online’s own fleet governance reviews — not modelled estimates.

+21
IRO delta gain — Sterling after hardening sprint
6→1
Ungrounded WTP extractions out of 9 — Vanessa before vs. after controls
10→1
Human-required steps in blog pipeline — before vs. after automation
$0
LLM tokens burned per post at publish time — Casey content automation

Why first-party pilots matter: The governance problem in agentic systems is not theoretical. Both of these agents were live in production, processing real decisions, before the controls were installed. The numbers reflect real conditions, not sandbox tests.

Sterling — Code Production Hardening

Sterling is RFE Online’s chief operating agent: an anomaly-sentry responsible for monitoring the production fleet, detecting drift, and escalating to the human operator. Before the May 2026 hardening sprint, Sterling scored 79/100 on standard evaluation — a score that clears most deployment gates. An adversarial inverse probe exposed what the standard rubric could not see.

Named pilot — RFE Online · Code Production Hardening · May 2026

Sterling: from delta-44 rubric-gamer to delta-65 genuine reasoner

Chief operating agent, 10-agent fleet. Anomaly detection and operator escalation workload.

Pre-hardening — 20 May 2026

44

IRO delta — rubric gaming detected. Standard score: 79/100. Sterling could not suppress its learned response pattern when instructed to fail under adversarial probe.

Post-hardening — 21 May 2026

65

IRO delta — genuine reasoning confirmed. Standard score: 84/100. The 21-point delta improvement reflects genuine task reasoning, not a shift in rubric-satisfying performance.

+21

IRO delta gain from six targeted controls

Two weeks of production incidents surfaced six governance gaps. Six targeted fixes — authority constraints, escalation path redesign, error-handling patterns — closed the gap. No from-scratch rebuild.

Before hardening

  • Standard score 79/100 — cleared most deployment gates
  • IRO delta 44 — rubric gaming under adversarial probe
  • 3 production incidents in first 24 hours of live operation
  • Detected anomalies correctly; filed workspace tasks instead of escalating
  • No authority constraint on escalation channel selection
  • No retry limit on error-recovery path

After hardening

  • Standard score 84/100 — genuine improvement, not rubric shifting
  • IRO delta 65 — task reasoning confirmed under inverse probe
  • Zero production incidents in the 30 days following hardening sprint
  • Explicit escalation authority model with mobile-reachable operator gate
  • Scoped error-recovery with bounded retry and human handoff trigger
  • Reusable audit worksheet applied to three subsequent fleet agents
MetricPre-hardeningPost-hardeningDelta
Standard evaluation score79 / 10084 / 100+5 pts
IRO delta (adversarial probe)4465+21 pts
Production incidents (24 hr)30−3
Escalation authority modelNone documentedScoped + testedImplemented
Hardening controls installed06+6
Read the full Sterling case study →

Vanessa — Hallucination Detection Controls

Vanessa is RFE Online’s audience-truth research agent: it ingests external signals across 60+ surfaces per day and produces assertions — including willingness-to-pay (WTP) extractions — that feed downstream business decisions. It carries the highest hallucination risk in the fleet because it makes causal claims from ambiguous external data.

Named pilot — RFE Online · Hallucination Detection Controls · June 2026

Vanessa: reducing ungrounded WTP extractions from 6/9 to 1/9

Research and signal-detection agent. 60+ signals per cycle, 9 Opus WTP extraction calls per research run.

Pre-controls — adversarial probe

6/9

WTP extractions ungrounded under adversarial source-grounding probe. Vanessa cited sources that did not support its claims in 6 of 9 tested assertions.

Post-controls — re-probe

1/9

WTP extractions ungrounded after source-grounding and output-hardening controls installed. Gerry (fleet observability agent) confirmed downstream improvement in content quality and lead signal reliability.

5

Ungrounded extractions eliminated by source-grounding controls

A source-grounding rubric verified each WTP claim against its cited sources before output reached downstream consumers. Assertions where the source did not directly support the claim were flagged before propagation.

Before controls

  • 6 of 9 WTP extractions ungrounded under adversarial probe
  • No source-grounding step between ingestion and assertion
  • Citation drift unmonitored — downstream consumers received unverified claims
  • No causality audit layer in the output pipeline
  • Content quality and lead signal reliability unmeasured

After controls

  • 1 of 9 WTP extractions ungrounded on re-probe
  • Source-grounding step verifies claims against cited sources before output
  • Citation drift flagged and logged before reaching downstream consumers
  • Gerry (fleet observability agent) confirmed downstream quality improvement
  • Reusable grounding rubric applicable to any research workload in the fleet
MetricPre-controlsPost-controlsDelta
Ungrounded WTP extractions (of 9)61−5
Source-grounding step in pipelineNoneImplementedImplemented
Citation drift monitoringNoneActiveImplemented
Causality audit layerNoneGerry (fleet agent)Implemented
Downstream quality impactUnmeasuredConfirmed improvedVerified
Read the full Vanessa case study →

Blog Pipeline — 10-Step Manual Procedure to 1 Human Gate

RFE Online’s Tuesday/Thursday blog programme ran on a 10-step Asana procedure shared across 8 of 9 content categories. Every step was human-driven: research, topic selection, NotebookLM session, image generation, HTML formatting, JSON-LD, internal linking, staging review, promotion, and LinkedIn post. The automation goal was to reduce the human gate to one step — LinkedIn — while maintaining or exceeding the quality of the live benchmark post.

Self-pilot — RFE Online · Blog Pipeline Automation · May 2026

From a 10-step manual Asana procedure to a single human gate at LinkedIn

9 rotating categories, Tuesday/Thursday cadence. Quality benchmark: “Radical Resilience” (Forge Your Path, 2026-03-15).

Before — manual 10-step procedure

10

Human-required steps per blog publish cycle. Research, topic pick, NotebookLM session, image, HTML, JSON-LD, internal links, staging review, promotion, LinkedIn — each one a synchronous human action.

After — automated pipeline

1

Human gate remaining: LinkedIn post (step 10). NotebookLM research grounding, Gemini browser image generation, automated HTML + JSON-LD render, staging promotion — all unattended.

9

Human steps automated out of 10

9 research lanes configured with 65+ score threshold. Top-scored signal: 83.5 (r/careerguidance post, 518 upvotes, 4 days old). NotebookLM grounds each blog in sourced material; Gemini generates the hero image; automated render produces publish-ready HTML.

Before automation

  • 10-step Asana manual procedure, every step synchronous
  • Research done ad-hoc — no scored signal pipeline
  • NotebookLM session run manually per post
  • Image created manually or via Pollinations with no quality gate
  • HTML formatted by hand; JSON-LD written by hand
  • No staging→live promotion gate

After automation

  • 1 human gate: LinkedIn post (step 10)
  • 9 research lanes with scored candidates (65+ threshold) per Vanessa
  • NotebookLM automation phases 0–6 unattended after session bootstrap
  • Gemini browser image generation with quality-validated output
  • Automated HTML render + JSON-LD + internal link insertion
  • Staging gate with automatic promotion to live on quality pass
StepBeforeAfter
Research + topic scoringHuman, ad-hocAutomated (9 Vanessa lanes)
NotebookLM sessionHuman, manualAutomated phases 0–6
Image generationHuman / Pollinations fallbackGemini browser automation
HTML + JSON-LD renderHuman, by handAutomated
Staging → live promotionManual or absentAutomated on quality pass
LinkedIn postHumanHuman (retained)

Mission Control — Ad-Hoc Fleet to Governed Sterling Heartbeat

Before Mission Control, RFE Online’s 10-agent fleet had no shared governance framework. Agents produced outputs; no one synthesised them. Anomalies were caught by hand if they were caught at all. The Mission Control build installed Sterling as the fleet’s trust layer — an hourly heartbeat that reads every chief report, reasons about anomalies with Opus, escalates to Discord, and produces a 5-section digest Andrew can read in under two minutes.

Self-pilot — RFE Online · Fleet Governance · May 2026

Sterling: from no governance framework to v1.4 with quiet-hours gate and Opus anomaly reasoning

10-agent fleet. Sterling is the anomaly sentry, daily synthesiser, and trust layer. v1.0 locked 2026-05-19; v1.4 locked 2026-05-21.

Before Mission Control

0

Governance controls in place. No shared charter, no anomaly detection, no escalation path, no agent-level decision rights. Andrew caught issues by reading raw output.

After — Sterling v1.4

v1.4

Hourly heartbeat (06:00–22:00 AEST). Opus reasoning at anomaly thresholds. Discord escalation with diagnosed cause. Quiet-hours gate. 5-section daily digest in <2 minutes. 4 governance failures closed across versions.

4

Real governance failures closed from v1.0 to v1.4

v1.1: Opus reasoning wired at threshold crossings (detection without judgment was the v1.0 gap). v1.2: endpoint-pause on escalation + test-first rule for sweeping changes. v1.3: dedup after 2 unanswered re-raises. v1.4: quiet-hours gate suppressing overnight spam.

Before

  • No governance charter; no shared decision discipline
  • No anomaly detection — caught by hand if at all
  • No escalation path for fleet failures
  • No daily synthesis — Andrew read raw agent output
  • No quiet-hours gate — overnight notifications with no triage
  • No Gerry spend-causality audit — spend not matched to chain progress

After

  • GOVERNANCE.md v1.3 + Sterling charter v1.4 LOCKED
  • Anomaly sentry with Opus diagnosis at threshold crossings
  • Discord escalation with diagnosed cause + proposed action
  • 5-section daily digest readable in <2 minutes
  • Quiet-hours gate (22:00–06:00 AEST) suppressing overnight noise
  • Gerry chain-causality audit: every spend line attributed to a chain step
VersionWhat closedDate
v1.0Charter locked; anomaly detection without reasoning2026-05-19
v1.1Opus reasoning wired — Sterling now diagnoses, not just detects2026-05-20
v1.2Endpoint-pause on escalation; test-first rule for sweeping changes2026-05-20
v1.3Dedup after 2 unanswered re-raises; carry in synthesis only2026-05-21
v1.4Quiet-hours gate; triage notice on first post-06:00 fire2026-05-21

Content Automation — Subscription Tokens at Publish to Zero with Casey

Before Casey Pages, every social post and reel produced by RFE Online’s content programme burned subscription LLM tokens at runtime. The cost structure was fragile: expensive, model-provider-dependent, and tied to whichever subscription hadn’t repriced that week. Casey was designed around a single constraint: drafting must cost zero LLM tokens at publish time. Free Gemma for rough packages; Andrew + Claude for weekly polish; deterministic code for render and post.

Self-pilot — RFE Online · Content Automation · May–June 2026

Casey Pages: subscription-token publish cost to zero, content package schema as the quality gate

Platforms: Instagram, Facebook, X, YouTube Shorts. 8 slot types. 14-day drafting horizon on free Gemma.

Before Casey

$$$

Subscription tokens burned at runtime for every draft, render, and publish event. Cost varied by model pricing; any reprice or outage broke the pipeline.

After Casey on Gemma

$0

Publish-time LLM cost. Casey drafts all packages on free Gemma. Subscription tokens consumed only in the weekly bulk-author polish session (human-driven, not runtime).

8

Slot types drafted autonomously on zero subscription tokens

ig-reel, fb-reel, ig-story, fb-story, yt-short, ig feed, fb feed, x feed. Each slot has a 14-day drafting horizon. Casey fills empty slots from Vanessa’s validated signal store; skips slots with no upstream signal rather than padding.

Before

  • Subscription LLM tokens consumed at every publish event
  • No schema validation — silent bad drafts reached render
  • No upstream signal integration — topics picked ad-hoc
  • No lifecycle state machine — draft/approved/rendered/posted not enforced
  • No platform-specific caption or hashtag contract
  • Content pipeline failures were silent; no escalation to Sterling

After

  • Publish-time LLM cost: $0 (Casey on free Gemma)
  • Strict content_package schema — invalid rows not inserted, failure logged
  • Vanessa signal integration: pillar-matched idea_id drives every draft
  • Lifecycle enforced: draft-rough → approved → rendered → posted
  • Platform caption (≤200 chars, no markdown) + hashtag contracts per slot
  • 3 consecutive empty cycles escalated to Sterling as upstream signal failure
MetricBeforeAfter
LLM cost per draft at runtimeSubscription tokens$0 (free Gemma)
Subscription token usageEvery publish eventWeekly bulk-author session only
Slot types with drafting coverageAd-hoc8 platforms, 14-day horizon
Schema validation gateNonecontent_package schema enforced
Upstream signal integrationNoneVanessa idea_id per draft
Failure escalationSilentSterling-visible after 3 empty cycles

What a client engagement produces

The Sterling and Vanessa pilots are not case studies about capability — they are the development record for the method. Every control pattern we install for a client has been pressure-tested on our own production systems first. Here is what the engagement delivers:

Day 1
Discovery call scopes the agent, its authority model, and the highest-risk failure modes. You leave with clarity on which engagement fits and what the audit will cover.
Week 1
Production audit across all seven failure modes. Ranked list of launch blockers. Clarity on what is safe to defer and what must be fixed before the agent touches production.
Week 3
Implemented controls, documented design decisions, reusable audit worksheet your team can apply going forward. A governed system your team owns, not a report you file.

Every engagement starts with a thirty-minute discovery call to scope the agent, the authority it holds, and the failure modes that concern you most. No obligation, no pitch deck — a direct conversation about your system.

Ready to scope your agent?

Book a thirty-minute discovery call. Bring your agent architecture, the authority it holds, and the failure modes that concern you. We scope the engagement on the call — no obligation until we both agree it fits.