Platform/Failure & Impact Prediction
Failure & Impact Prediction

Know the blast radius before customers feel it

Raincurve scores emerging conditions, estimates what they will take down and ranks them by operational impact — so the right problem gets attention first.

The problem

Every dashboard is red. Which one matters?

Degradations rarely announce themselves as outages. A rising error rate on an optic, a slow drift in a queue, a GPU throwing correctable errors — most are harmless, a few are the start of an incident. Teams triage by severity labels that know nothing about what sits downstream.

Raincurve ranks conditions by consequence: which services, tenants and redundancy margins depend on the thing that is degrading, and how likely it is to fail.

From signal to impact

01

Score

Every signal gets a continuous risk score from the Curve-1 pre-filter, so rare, dangerous patterns stand out from routine noise.

02

Relate

Trajectory encoding compares the current sequence of events to known failure trajectories across layers.

03

Project

The dependency graph projects the condition forward: which paths, services and tenants lose redundancy or capacity.

04

Prioritize

Conditions are ranked by likelihood × impact, with the next responsible action attached.

Resilience, measured

Redundancy is a number, not a diagram.

For every endpoint Raincurve computes resilience — the number of device-disjoint paths to the core. A condition that takes a server from two paths to one is flagged before it becomes an outage, even when everything is still reachable.

Because the computation is a single linear-time graph pass, it stays current on estates with tens of thousands of devices.

Impact forecastranked
0.91
agg-3 optical Rx degrading
36 servers lose redundancy · 2 tenants
0.64
gpu-p3 correctable ECC rising
4 inference endpoints at risk
0.22
core-1 CPU spike
base-rate event · no downstream impact
0.08
acc-9-2 port flap
single-homed test host
0.897AUC-ROC predicting failures across 16,000+ CI/CD trajectoriesCOSMED deployment
~0.88AUC on held-out data in multi-language evaluationCurve-1 evaluation
95Confirmed failure trajectories across 13 categories used in trainingCurve-1 training
~2 msPre-filter scoring per event, keeping continuous operation affordableCurve-1 pipeline

Make infrastructure intelligence operational.

Start with a conversation about your environment.