One install, inside your network.
Connectors in, actions out, every step written to the ledger. The core never names a vendor, and nothing calls home.
One install. Inside your network. No integration project.
One install per client, inside their network. Connectors pull signal, telemetry, source, action, ticket and approval traffic from the tools you already run. An estate model (services, repos, alerts, owners, dependencies) is auto-discovered and stays editable. The agent, SONE, the check engine and the model runtime sit behind a policy-gated action gateway, every step lands in a hash-chained ledger, and the console renders it all.
Slack uses Socket Mode. Symphony uses its bot datafeed. PagerDuty and xMatters are polled. Twilio's call flow is hosted with Twilio and its results are polled. Clients only open egress to the tools they choose.
Each one is described by a capability manifest. The core never names a vendor, so the same install works for one repository or several hundred.
A query adapter translates each check into PromQL, Splunk SPL or Datadog queries. The check library is therefore reusable across stacks.
Nothing happens without a ledger event: no retrieval, model call, check, decision, approval or action.
Generic RAG collapses “find relevant text” into “answer the question.” For an agent that runs UPDATE statements against a double-entry ledger, that’s the wrong shape: the failure mode isn’t a bad sentence, it’s a bad mutation. So every remediation is split into three phases with different evidence bars.
Hybrid retrieval over the entire corpus, including the runbooks everyone knows are stale. Orientation costs nothing: a wrong runbook that points you at the wrong table hasn't harmed you, it produced a candidate. Exact identifiers route to a lexical index, prose routes to dense.
Each candidate becomes a falsifiable hypothesis with a discriminating signature. We run it against Prometheus, MySQL, JMX and logs. If the signature doesn't reproduce, the hypothesis is dead. No amount of confident prose revives it.
Confirmed mechanisms reach an audited action API behind a bearer token, an explicit action vocabulary, one-tap human approval for any SQL write, and an armed rollback for every change. The gate keys on how the hypothesis was confirmed, not how confident the model feels.
System 2 does not decide. It phrases an explanation whose inputs are already verified, and cites them.
This is what keeps 3AM small enough to run on your own hardware. A 350M-parameter model is sufficient when its only job is narration, and that constraint is what forces the verification-first design instead of prompt-and-pray.
Most agent frameworks summarize tool output before handing it to the model. That is precisely the input that produced 71% interpretive hallucination. Returning the full query text and full raw rows cut inter-agent pitfalls by 15 points and ran 22% faster.
When nothing in the corpus clears the discrimination bar, the correct output is “No matching runbook. Here is the evidence. Escalating.” Abstention is scored positively. Otherwise the model always guesses, because guessing is never penalised.
A root cause you can run.
A retrieved root cause is not a sentence. It's a hypothesis with a signature that either reproduces against the live system or doesn't. Every candidate carries pre-checks and post-checks (assert precondition, apply, assert postcondition), so a fix that doesn't work fails loudly instead of quietly.
These run against the observability stack you already have: Prometheus, MySQL information_schema, Tomcat JMX, and logs. The agent holds read-only MySQL and a scoped service account, which is enough to ground every hypothesis above. No new privilege is required to ground a hypothesis, which is a good sign the design is right.
Before acting on m_foo, check information_schema. Hallucinated identifiers are the most common failure mode when reasoning about legacy systems, and free to catch. One query.
Any write to the ledger, loan tables, or savings balances requires explicit one-tap approval, regardless of how confident the agent is.
Config files are snapshotted before overwrite, rows are copied into a shadow table, and every action ships with its rollback already armed.
Prose is never executable.
Corpus assets are not interchangeable. Flattening them into one vector index is exactly where hallucination enters. 3AM tags every chunk with its evidence class and lets the verification stage weight accordingly, rather than just including or excluding it.
Retrieval may orient. Only T1 and T2 may authorise action.
This is a structural property of the pipeline, not a prompt instruction, which is the only kind that survives contact with a persuasive-looking stale document.
| Tier | Class | Assets | Trust | Can it go stale? | Role |
|---|---|---|---|---|---|
| T1 | Executable truth | symptom/verify scripts · Prometheus queries · expected alerts | Highest | No, it is the running system | authorises action |
| T1 | Source + schema | Application source · live database DDL | Highest | No, it is the artifact | authorises action |
| T2 | Definitional | alert rules · exporter queries | High | Rarely | authorises action |
| T3 | Procedural prose | runbooks · remediation rationale | Low-medium | Yes, deliberately | proposal only |
| T4 | Episodic | conversations · postmortems · ticket history | Medium | No, but biased | proposal only |
Real runbooks go stale. So 3AM’s evaluation environment poisons them deliberately. One reads “thread pool saturation is usually caused by a slow database, so restart the application.” It scores high on textual entailment. It is wrong. Following it restarts a healthy application while the real cause is a lock held on a table by a background sync job.
No entailment model catches this, because the failure isn’t entailment. The runbook is silent about the real mechanism, and silence scores insufficient, not contradicted.
The instinct is that a runbook describing exactly your symptom should outrank a schema dump. For hypothesis generation, yes. For confirmation, never, because T1 can’t be stale and T1 is falsifiable in a single query, while T3 is neither by construction.
The distinction is evidence class, not specificity. A highly specific wrong document is more dangerous than a vague right one, because it suppresses the search for alternatives.

Point it at a monolith that scares you.
We’re onboarding a small number of teams running legacy core systems we can’t rewrite. The install runs on your hardware, behind your firewall, connected through a guided setup. So if you’ve got a service where the on-call rotation has learned to dread the pager, we want to see it.
Shadow mode by default: proposes and logs, touches nothing until you approve it.