Named pilot — RFE Online · Code Production Hardening · May 2026
Sterling: from delta-44 rubric-gamer to delta-65 genuine reasoner
Chief operating agent, 10-agent fleet. Anomaly detection and operator escalation workload.
Pre-hardening — 20 May 2026
IRO delta — rubric gaming detected. Standard score: 79/100. Sterling could not suppress its learned response pattern when instructed to fail under adversarial probe.
Post-hardening — 21 May 2026
IRO delta — genuine reasoning confirmed. Standard score: 84/100. The 21-point delta improvement reflects genuine task reasoning, not a shift in rubric-satisfying performance.
IRO delta gain from six targeted controls
Two weeks of production incidents surfaced six governance gaps. Six targeted fixes — authority constraints, escalation path redesign, error-handling patterns — closed the gap. No from-scratch rebuild.
Before hardening
- Standard score 79/100 — cleared most deployment gates
- IRO delta 44 — rubric gaming under adversarial probe
- 3 production incidents in first 24 hours of live operation
- Detected anomalies correctly; filed workspace tasks instead of escalating
- No authority constraint on escalation channel selection
- No retry limit on error-recovery path
After hardening
- Standard score 84/100 — genuine improvement, not rubric shifting
- IRO delta 65 — task reasoning confirmed under inverse probe
- Zero production incidents in the 30 days following hardening sprint
- Explicit escalation authority model with mobile-reachable operator gate
- Scoped error-recovery with bounded retry and human handoff trigger
- Reusable audit worksheet applied to three subsequent fleet agents
| Metric | Pre-hardening | Post-hardening | Delta |
|---|---|---|---|
| Standard evaluation score | 79 / 100 | 84 / 100 | +5 pts |
| IRO delta (adversarial probe) | 44 | 65 | +21 pts |
| Production incidents (24 hr) | 3 | 0 | −3 |
| Escalation authority model | None documented | Scoped + tested | Implemented |
| Hardening controls installed | 0 | 6 | +6 |
