Platform/Verified Remediation
Verified Remediation

Prove an action is safe, then pace it within limits

Before a drain, reboot or rollout touches production, Raincurve checks it against live state and in-flight work — then schedules it so the fix does not cause the next incident.

98.4%Of unsafe maintenance actions approved by standard runbook checksVerification study
0Unsafe approvals with exact graph verificationVerification study
0.19 sTo verify a 14,680-device network on one CPU coreBlock-cut tree
43%Less expected damage with churn-capped rollout pacingAlmgren-Chriss study
Verify

Runbooks check the action. Raincurve checks the network.

An action that is safe in isolation becomes dangerous next to maintenance elsewhere, or after a device has already failed. Runbook checks ignore both — and even state-aware reachability checks miss the loss of backup paths.

Raincurve enforces a do-no-harm rule: after the action, every server keeps at least as much resilience as it had, up to its required level. Using biconnected components, the check covers the whole estate in one linear-time pass and explains any violation in plain terms.

Pace

How fast should the fix roll out?

Go too fast and reconvergence storms cause their own outage; go too slow and devices stay degraded. Raincurve adapts the Almgren-Chriss optimal-execution model from trading to remediation, with a hard cap on changes per minute where the control plane stops coping.

When a fix might itself be bad, the schedule starts slowly to learn, then accelerates to the churn limit once enough devices succeed.

Verification resultblocked
action
Drain agg-4 for firmware upgrade
requested by change CHG-2291
state
agg-3 already degraded · 1 drain in flight
live
check
p3-srv-1-2 would drop from 2 paths to 1
required: 2
result
Not safe now
retry after agg-3 restored
pacing
Otherwise: ≤ 8 changes/min, canary first
churn cap

Guardrails that travel with every action

Works for any proposer

Humans, scripts and models all pass through the same check before execution.

Explanations, not just verdicts

Every block names the endpoint, the resilience before and after, and the requirement it violates.

Sandboxed execution

Approved actions run in isolated microVMs with snapshots for fast, clean rollback.

Progressive automation

Start with recommendations, move to approvals, then to automatic execution inside agreed limits.

Change-system aware

In-flight maintenance and freeze windows are part of the state being verified.

Full audit trail

Proposal, verification, approval, execution and outcome are recorded together.

Make infrastructure intelligence operational.

Start with a conversation about your environment.