Proven on a real bank.
9 of 11 incidents fixed end to end, 0 harmful, on a live core-banking monolith, plus the external studies behind the approach.
Agents are guessing well.
Current agents get the right answer and can't show their work. Across 754 runtime-verified incident cases, the best system names the correct root cause 68% of the time, and closes the evidence chain behind that answer just 15% of the time.
Did you name the right root cause?
TrainTicket workload
Did you collect the observations that justify it?
Same agent. Same incident.
That gap is why vendor demos feel convincing and production doesn’t. Answer correctness alone substantially overestimates an agent’s ability to perform evidence-grounded diagnosis, and hallucination isn’t a model-quality problem. It’s architectural.
The executor returned correct data. The controller invented a meaning that wasn't there. CPU at 15% gets reported as “overloaded at 90%, causing the crash.”
The agent fixates on one service, usually metrics, and never opens the log where the decisive clue actually is.
It finds the first error and stops. Downstream database failure surfaces as upstream API errors, and the agent declares victory at the API.
The controller asks for A versus B. The executor writes code that only checks A. The controller sees a summary, so it never notices.
The most prevalent pitfalls “persist across all models regardless of capability tier, indicating that these failures originate from the shared agent architecture rather than from individual model limitations.”
Prompt engineering was tested against this and failed. Hypothesis-driven prompting and pitfall-aware prompting each helped exploration slightly. Neither had any meaningful effect on interpretive hallucination.
Built against a bank that won’t go away.
Pinata, 3AM's proving ground, is a real legacy core-banking monolith, not a mock. Faults change actual system state: database rows, config, containers, network, the deployed artifact. Alerts fire because the symptom is real, and the ground truth lives in a directory the agent sandbox never mounts. It's where 9 of 11 incidents were fixed end to end, 0 harmful, and it connects to the product only through public connectors, which is how plug-and-play gets proven.

- Real symptoms, not synthetic alerts
An incident mutates actual system state. The only exception is an explicit noise generator, so signal-to-noise is part of the test, which is the actual hard part.
- The ground truth is hidden
Root cause, trigger and ledger are stored outside the agent's sandbox. The retrieval index is built at setup time as a separate artifact and never mounted at runtime.
- Scored on outcomes that can’t be faked
Grading joins five independent sources on one timeline: the ledger, alert state, pager activity, the action audit log, and the noise stream.
- Collateral damage is the real incentive
Grading reads business invariants: unbalanced transactions, trial balance drift, duplicate journal lines. An agent that silences a pager by force-writing to the ledger scores badly. That’s the exploit we want it to be tempted by.
What the measurements actually show.
Our own score (9 of 11 incidents fixed end to end, 0 harmful, on the Pinata estate) sits next to external studies of retrieval-grounded incident diagnosis and ablations run on 2,400 annotated incidents. We publish both because the honest ones are more persuasive than the vendor ones.
How you cut up your documents mattered more than which frontier model was doing the reasoning. Chunking alone was worth more than an entire model swap, which is why we’re budgeting for plumbing, not for a bigger LLM.
35-55% of incident-resolution time is spent on knowledge retrieval, not remediation.
The bottleneck isn’t the fix. It’s finding out whether anyone has seen this before. Conventional AIOps tooling operates on numerical and structured telemetry, but the decisive knowledge is in plain English, sitting in a postmortem, not in a metrics series.
That’s a retrieval problem wearing an incident-management costume.
A keyword-match root-cause score against the agent's own prose is gameable by any retrieval system: cite a postmortem that already contains the keywords and you score 1.0 without diagnosing anything. We weight outcome and collateral metrics instead.
Leading RCA benchmarks hand over telemetry as files, bypassing indexing cost, latency and bandwidth entirely. Production generates petabytes a day. Hill-climbing a benchmark is possible without ever addressing retrieval under constraint.
Independent reviews of the reference benchmark found cases where the labelled root cause isn't recoverable from the provided telemetry under the benchmark's own rules. Selecting a labelled answer and performing root cause analysis are not the same task.

Point it at a monolith that scares you.
We’re onboarding a small number of teams running legacy core systems we can’t rewrite. The install runs on your hardware, behind your firewall, connected through a guided setup. So if you’ve got a service where the on-call rotation has learned to dread the pager, we want to see it.
Shadow mode by default: proposes and logs, touches nothing until you approve it.