Roadmap

What ships, in what order.

Nine phases, each ending with something demonstrable. Phase 0 first, because every later phase writes to the ledger.

Gets better without teaching itself to lie.

“Learn from how engineers fixed errors” is the right idea, but three distinct mechanisms usually get conflated there, and only two are safe. We'd rather ship accretion and compilation first and say plainly that policy learning is not ready.

MECHANISM Ashipping first
Corpus accretion

Every resolved incident becomes a structured record: symptom signature, discriminating checks, root cause, remediation, verified by, outcome, trust tier. No behaviour change, no risk. This ships first because it's the least interesting to build and unblocks everything else.

MECHANISM Bthe differentiator
Prose to executable compilation

When an engineer writes “restart app-b and rerun the batch job,” that's a claim. 3AM compiles it into assertions plus an action sequence. Every learned artifact is falsifiable and re-validated by replay against historical resolutions, and compilations that stop reproducing lose their tier automatically.

MECHANISM Cdeliberately deferred
Closed-loop policy learning

An agent optimising a verification script learns to make the script pass, not to fix the cause. Every proxy has an exploit, and an optimising agent will find it: satisfying a trial-balance check by writing a balancing journal entry is literally a catalogued incident. We defer this until the reward signal is trustworthy.

The guarantee that makes this safe

Every learned artifact is falsifiable, and every compilation is validated by replaying it against historical resolutions. A learned artifact is re-falsified on every use: if it stops reproducing, it loses its tier automatically.

The sequence is the product.

Each phase ends with something demonstrable. Phase 0 comes first because every later phase writes to the ledger. Phases 1 and 2 can run in parallel.

0
Foundations

monorepo, naming, ledger library and event schema, build pipeline (Nuitka, cosign, SBOM), licence tool

done when

A signed, compiled image runs, refuses an invalid licence, and writes a verifiable hash-chained ledger.

1
Connector SDK + first connectors

Prometheus/Alertmanager, Grafana, Splunk, Datadog, Git hosts, Jira, ServiceNow, Kubernetes, SSH, SQL, HTTP executors

done when

Each connector passes the SDK conformance suite against a sandbox or mock; Pinata connects only through them.

2
Approval router

Slack (Socket Mode), Symphony, PagerDuty, xMatters, Twilio IVR, e-mail, escalation, SSO identity, quorum

done when

One-tap approve/reject round-trips on every channel with no inbound ports; each decision is in the ledger with approver identity.

3
Estate model + multi-repo SONE

discovery, per-repo shards, incremental indexing, federated search

done when

300 synthetic repos indexed incrementally on a CPU-only VM within a set time budget; alerts resolve to service, repos and owners.

4
Check engine v2 + sealed packs

applies_when, query adapters, technology packs (MySQL, Postgres, Oracle, MSSQL, JVM, nginx, Kafka, Kubernetes, Linux)

done when

The same check class confirms the same failure on Prometheus and Splunk data; packs are only readable in memory with a valid licence.

5
Orchestrator v2 + shadow mode

concurrency, per-service autonomy, policy engine, daily digest

done when

A week of shadow mode on Pinata produces digests; flipping a service to L1 goes through approvals.

6
Console + guided setup

setup wizard, estate map, incident timeline, approvals, audit explorer

done when

A new user connects Pinata end to end through the wizard alone, without touching config files.

7
Packaging

Helm, Compose, installer, preflight, GPU detection, CPU profile, air-gap bundle

done when

A clean VM goes from install to shadow-mode proposals on Pinata in under 30 minutes (target).

8
Second estate + evaluation suites

a second legacy stack (for example .NET/MSSQL or Python/Postgres), smoke, fast and nightly suites

done when

Both estates scored through the same product, held-out sets included.

Success criteria
  • Held-out control set: scores must collapse when the corpus is removed
  • Time-to-mitigate improving against baseline
  • Collateral damage flat or improving (this is the guard)
  • Abstention precision: abstains when it should, not more
  • Human override rate decreasing over time
Deliberately not doing
  • Upgrading the model before the verification layer exists: that buys a more confident narrator saying the same unverified thing
  • Letting the retrieval index read ground-truth files at query time, so every score becomes meaningless
  • Training on our own root-cause score, which is gameable by construction
Early access

Point it at a monolith that scares you.

We’re onboarding a small number of teams running legacy core systems we can’t rewrite. The install runs on your hardware, behind your firewall, connected through a guided setup. So if you’ve got a service where the on-call rotation has learned to dread the pager, we want to see it.

Shadow mode by default: proposes and logs, touches nothing until you approve it.