On-prem AI SRE, no integration project

The AI SRE that fixes prod, not just describes it.

It watches your legacy monolith, proves each hypothesis with executable checks against the live system, and only then touches anything.

incoming
incident detected
payments-coredetected

FineractApiHighErrorRate

5xx 4.1% over 5m · p99 3.2s

iam_syncdetected

LockWaitAboveBudget

p99 2.4s on m_appuser

iam_syncgrounded

H3 lock tables held

confirmed against live state

payments-corefix applied

Guard patch on 083

applied · rollback armed

payments-coreverified

Error rate 4.1% → 0.0%

trial balance intact

auth-edgedetected

TokenRefreshStorm

3.1x baseline · 1.9k in flight

auth-edgeverified

Token storm cleared

0 in flight · no session drops

nightly-rollupdetected

WorkerOOMKilled

exit 137 · retry 2/3

nightly-rollupfix applied

Rollup completed

0 rows lost · 4m 51s

hypotheses grounded: 6collateral damage: none

Five minutes, page to fix.

A real incident from the Pinata estate, replayed from the decision ledger: the page lands, three hypotheses are grounded against live state, one survives, and the fix is applied with rollback armed. Every line below is a ledger event.

incident/pager-core-banking/083
resolved
03:14:02DETECTFineractApiHighErrorRate
03:14:41ACCEPTpager-core-banking acknowledged
03:15:19RETRIEVE42 candidates, stale runbooks included
03:16:03GROUNDH1 metadata lock on m_appuser: FALSIFIED
03:16:44GROUNDH2 conn leak: FALSIFIED
03:17:58GROUNDH3 iam_sync holds LOCK TABLES: CONFIRMED
03:18:31ACTblocked 083, applied guard patch, rollback armed
03:19:12VERIFYerror rate 4.1% to 0.0%, trial balance intact
time-to-mitigate: 5m 10scollateral damage: none

One lab’s proof, turned into a plug-and-play install.

Today everything is wired to one estate: one stack's URLs, container names, tables and alert names. That was right for proving the approach: 9 of 11 incidents fixed end to end, 0 harmful. But a client has its own repos, monitoring, pager, chat and runtime, sometimes hundreds of each. 3am.si turns the proof into a product: a client installs it on-prem and connects to any legacy estate through a guided setup, with no integration project.

3am.si
The company and the platform
3am Core
The on-prem install: the agent, action gateway, approval router, ledger, connectors and console
SONE
The knowledge layer: hybrid retrieval, System 1 verification, trust tiers. One index per repo, federated search across them
Pinata
The proving ground. Runs the incident labs and the evaluation harness, and connects to 3am Core only through public connectors
/11
incidents fixed end to end
harmful, across all of them
→
repositories per install, same product
Decisions taken
Deployment

On-prem only. Nothing leaves the client's network apart from the SaaS tools they choose to connect.

IP protection

Compiled binaries in signed images; an offline signed licence; our knowledge sealed in encrypted packs; a EULA.

Models

The client's GPU, auto-detected. A CPU fallback profile so a pilot runs on any VM.

Autonomy at switch-on

Shadow mode first (proposes, logs, touches nothing), then L1 one-tap approval, per service.

Packaging

Helm chart first, plus a single-host Compose bundle for VM-only estates, built from the same images.

v1 integration scope

Source repos, monitoring and alert sources, action targets, ticketing and CMDB.

Approval channels

Slack, Symphony, PagerDuty, xMatters/Everbridge, Twilio voice, plus e-mail as a fallback.

First connectors

Prometheus/Grafana, Splunk + ServiceNow, Datadog + Jira.

Audit

Every action, every reason and every model call logged in a tamper-evident form.

Early access

Point it at a monolith that scares you.

We’re onboarding a small number of teams running legacy core systems we can’t rewrite. The install runs on your hardware, behind your firewall, connected through a guided setup. So if you’ve got a service where the on-call rotation has learned to dread the pager, we want to see it.

Shadow mode by default: proposes and logs, touches nothing until you approve it.