Evidence

Proven on a real bank.

9 of 11 incidents fixed end to end, 0 harmful, on a live core-banking monolith, plus the external studies behind the approach.

Agents are guessing well.

Current agents get the right answer and can't show their work. Across 754 runtime-verified incident cases, the best system names the correct root cause 68% of the time, and closes the evidence chain behind that answer just 15% of the time.

Joint RCA Accuracy
0.68

Did you name the right root cause?

TrainTicket workload

Evidence Closure Rate
0.15

Did you collect the observations that justify it?

Same agent. Same incident.

That gap is why vendor demos feel convincing and production doesn’t. Answer correctness alone substantially overestimates an agent’s ability to perform evidence-grounded diagnosis, and hallucination isn’t a model-quality problem. It’s architectural.

1,675 agent runs · 5 frontier models · full OpenRCA benchmark
71.2%Interpretive hallucination

The executor returned correct data. The controller invented a meaning that wasn't there. CPU at 15% gets reported as “overloaded at 90%, causing the crash.”

63.9%Incomplete exploration

The agent fixates on one service, usually metrics, and never opens the log where the decisive clue actually is.

39.9%Symptom-as-cause

It finds the first error and stops. Downstream database failure surfaces as upstream API errors, and the agent declares victory at the API.

18.8%Instruction-code mismatch

The controller asks for A versus B. The executor writes code that only checks A. The controller sees a summary, so it never notices.

The most prevalent pitfalls “persist across all models regardless of capability tier, indicating that these failures originate from the shared agent architecture rather than from individual model limitations.”

Prompt engineering was tested against this and failed. Hypothesis-driven prompting and pitfall-aware prompting each helped exploration slightly. Neither had any meaningful effect on interpretive hallucination.

Built against a bank that won’t go away.

Pinata, 3AM's proving ground, is a real legacy core-banking monolith, not a mock. Faults change actual system state: database rows, config, containers, network, the deployed artifact. Alerts fire because the symptom is real, and the ground truth lives in a directory the agent sandbox never mounts. It's where 9 of 11 incidents were fixed end to end, 0 harmful, and it connects to the product only through public connectors, which is how plug-and-play gets proven.

The banking hall of a 1929 bank building — the kind of institution the lab is built against
A bank that won’t go awayPhoto: w_lemay (CC BY-SA 2.0)
244k
lines of production Java
106
injectable incidents, real symptoms
95
alert rules across 10 files
92
runbooks, some deliberately stale
~200
services in the synthetic estate
18
audited guarded remediation actions
Incident catalog by category
Database17
Loans15
Messaging13
COB (batch)12
Savings10
Accounting10
Infrastructure9
API6
Security5
Configuration4
Deployment3
Performance2
Why this is a hard evaluation
  • Real symptoms, not synthetic alerts

    An incident mutates actual system state. The only exception is an explicit noise generator, so signal-to-noise is part of the test, which is the actual hard part.

  • The ground truth is hidden

    Root cause, trigger and ledger are stored outside the agent's sandbox. The retrieval index is built at setup time as a separate artifact and never mounted at runtime.

  • Scored on outcomes that can’t be faked

    Grading joins five independent sources on one timeline: the ledger, alert state, pager activity, the action audit log, and the noise stream.

  • Collateral damage is the real incentive

    Grading reads business invariants: unbalanced transactions, trial balance drift, duplicate journal lines. An agent that silences a pager by force-writing to the ledger scores badly. That’s the exploit we want it to be tempted by.

What the measurements actually show.

Our own score (9 of 11 incidents fixed end to end, 0 harmful, on the Pinata estate) sits next to external studies of retrieval-grounded incident diagnosis and ablations run on 2,400 annotated incidents. We publish both because the honest ones are more persuasive than the vendor ones.

%
root-cause identification
vs 71.8% for BM25 with the same model
−%
mean diagnosis time
across 2,400 annotated incidents
m
P1 critical, from 48.2 min
in a midsized financial-services estate
/5
SRE satisfaction
driven by click-through citations to source
Ablation · accuracy cost of removing each component
Semantic / structure-aware chunking−10.9 pp
Cross-encoder re-ranking−7.7 pp
Metadata filteringa few pp
Feedback loopa few pp

How you cut up your documents mattered more than which frontier model was doing the reasoning. Chunking alone was worth more than an entire model swap, which is why we’re budgeting for plumbing, not for a bigger LLM.

Why the knowledge layer exists at all

35-55% of incident-resolution time is spent on knowledge retrieval, not remediation.

The bottleneck isn’t the fix. It’s finding out whether anyone has seen this before. Conventional AIOps tooling operates on numerical and structured telemetry, but the decisive knowledge is in plain English, sitting in a postmortem, not in a metrics series.

That’s a retrieval problem wearing an incident-management costume.

What we won’t claim
We don't train on our own score

A keyword-match root-cause score against the agent's own prose is gameable by any retrieval system: cite a postmortem that already contains the keywords and you score 1.0 without diagnosing anything. We weight outcome and collateral metrics instead.

Benchmarks hand you the telemetry

Leading RCA benchmarks hand over telemetry as files, bypassing indexing cost, latency and bandwidth entirely. Production generates petabytes a day. Hill-climbing a benchmark is possible without ever addressing retrieval under constraint.

Some cases aren't identifiable

Independent reviews of the reference benchmark found cases where the labelled root cause isn't recoverable from the provided telemetry under the benchmark's own rules. Selecting a labelled answer and performing root cause analysis are not the same task.

Early access

Point it at a monolith that scares you.

We’re onboarding a small number of teams running legacy core systems we can’t rewrite. The install runs on your hardware, behind your firewall, connected through a guided setup. So if you’ve got a service where the on-call rotation has learned to dread the pager, we want to see it.

Shadow mode by default: proposes and logs, touches nothing until you approve it.