Root-cause reasoning across carrier-grade networks
Turn alarm floods from transport, IP/MPLS and access layers into incidents with a named origin — and roll out fixes at a pace the control plane can absorb.
A fiber cut is one event and ten thousand alarms.
Carrier networks are layered by design: optical transport underneath IP/MPLS, underneath access and services. When something fails low in the stack, every layer above it reports its own version of the problem, often from different management systems with different clocks.
NOC teams correlate by time window and adjacency rules that were tuned years ago. When background noise rises — a storm, a maintenance night, a vendor bug — those rules either merge unrelated events or split one incident into dozens of tickets.
Raincurve models the network the way it actually fails: faults propagate along topology, with delays, and trigger further alarms. It learns those cascades from your own alarm history, without labels.
Where time is lost today
Rules collapse under noise
Time-window correlation that works at quiet hours fails exactly when alarm volume spikes.
Duplicate tickets
One incident fans out into tickets per layer and per region, each triaged independently.
Clock skew between systems
Transport and IP management systems disagree on timestamps, scrambling the order of events.
Remediation storms
Pushing a fix to hundreds of devices at once triggers reconvergence and a second incident.
Concurrent maintenance
Planned work in one region silently removes the protection path another region relies on.
Tribal knowledge
Which alarm usually causes which lives in senior engineers' heads, not in the tooling.
Cause, not coincidence.
Raincurve fits a topology-aware Hawkes process to your alarm stream. Each alarm is explained either as background or as triggered by an earlier alarm on a nearby device, with learned strengths between alarm types and two decay timescales for propagation and polling.
In our study, performance held as noise rose: at 15 alarms per minute of noise the model kept an F1 of 0.73 while a tuned rule engine fell to 0.37, and at 30 per minute the rules collapsed to 0.03. Clock skew up to 60 seconds had little effect.
2/mintopology-window rules 0.75
6/minrules 0.67
15/minrules 0.37
30/minrules 0.03
From alarm to safe rollout
Ingest
Alarms and telemetry from transport, IP and access management systems, normalized and aligned in time.
Group
Topology-aware causal grouping produces one incident per fault, with a parent chain back to the origin.
Verify
Proposed actions are checked against protection paths and in-flight maintenance before approval.
Pace
Fleet-wide fixes are scheduled with Almgren-Chriss pacing under a hard churn cap, starting slowly while the fix is unproven.
Platform capabilities used here
Every solution runs on the same Raincurve platform.
How engagements start
Carrier engagements typically start by replaying a quarter of historical alarms through Raincurve in the Tier 1 assessment.
Infrastructure Assessment
A scoped read of your topology, telemetry and recent incidents. We map cross-layer dependencies, replay past incidents through Curve-1 and show where time to resolution is lost.
Production Pilot
Raincurve runs alongside your existing tools on a defined slice of production — a hall, a region, a cluster — with success criteria agreed up front and measured weekly.
Enterprise Reliability Platform
Estate-wide deployment in your environment: continuous cross-layer reasoning, verified remediation workflows, and integration with your NOC, ticketing and change processes.
Technical documentation
Deployment, integrations and concepts for this environment.
Deployment guide
Private, hybrid and cloud topologies, sizing and security boundaries.
Open docs →DocumentationIntegrations
Telemetry, inventory, orchestration and ticketing sources Raincurve reads and writes.
Open docs →DocumentationPlatform concepts
Topology graph, contracts, hypotheses, incidents and verification.
Open docs →