Solutions/Carriers & Telecom
Carriers & Telecom

Root-cause reasoning across carrier-grade networks

Turn alarm floods from transport, IP/MPLS and access layers into incidents with a named origin — and roll out fixes at a pace the control plane can absorb.

Built forTransportIP / MPLSAccessNOC operations
The operating reality

A fiber cut is one event and ten thousand alarms.

Carrier networks are layered by design: optical transport underneath IP/MPLS, underneath access and services. When something fails low in the stack, every layer above it reports its own version of the problem, often from different management systems with different clocks.

NOC teams correlate by time window and adjacency rules that were tuned years ago. When background noise rises — a storm, a maintenance night, a vendor bug — those rules either merge unrelated events or split one incident into dozens of tickets.

Raincurve models the network the way it actually fails: faults propagate along topology, with delays, and trigger further alarms. It learns those cascades from your own alarm history, without labels.

Where time is lost today

Rules collapse under noise

Time-window correlation that works at quiet hours fails exactly when alarm volume spikes.

Duplicate tickets

One incident fans out into tickets per layer and per region, each triaged independently.

Clock skew between systems

Transport and IP management systems disagree on timestamps, scrambling the order of events.

Remediation storms

Pushing a fix to hundreds of devices at once triggers reconvergence and a second incident.

Concurrent maintenance

Planned work in one region silently removes the protection path another region relies on.

Tribal knowledge

Which alarm usually causes which lives in senior engineers' heads, not in the tooling.

How grouping works

Cause, not coincidence.

Raincurve fits a topology-aware Hawkes process to your alarm stream. Each alarm is explained either as background or as triggered by an earlier alarm on a nearby device, with learned strengths between alarm types and two decay timescales for propagation and polling.

In our study, performance held as noise rose: at 15 alarms per minute of noise the model kept an F1 of 0.73 while a tuned rule engine fell to 0.37, and at 30 per minute the rules collapsed to 0.03. Clock skew up to 60 seconds had little effect.

Grouping under noisepairwise F1
2/min
Hawkes + topology 0.92
topology-window rules 0.75
6/min
Hawkes + topology 0.85
rules 0.67
15/min
Hawkes + topology 0.73
rules 0.37
30/min
Hawkes + topology 0.60
rules 0.03

From alarm to safe rollout

01

Ingest

Alarms and telemetry from transport, IP and access management systems, normalized and aligned in time.

02

Group

Topology-aware causal grouping produces one incident per fault, with a parent chain back to the origin.

03

Verify

Proposed actions are checked against protection paths and in-flight maintenance before approval.

04

Pace

Fleet-wide fixes are scheduled with Almgren-Chriss pacing under a hard churn cap, starting slowly while the fix is unproven.

0.85Pairwise F1 at default noise vs 0.67 for topology-window rulesHawkes study, simulated
1.6 sTo fit 24 hours of alarms (13,638 events) on a laptopHawkes study
43%Less expected damage with churn-capped rollout pacingAlmgren-Chriss study
8/minIllustrative churn cap: the point where reconvergence storms beginAlmgren-Chriss study

How engagements start

Carrier engagements typically start by replaying a quarter of historical alarms through Raincurve in the Tier 1 assessment.

Tier 1

Infrastructure Assessment

A scoped read of your topology, telemetry and recent incidents. We map cross-layer dependencies, replay past incidents through Curve-1 and show where time to resolution is lost.

Tier 2

Production Pilot

Raincurve runs alongside your existing tools on a defined slice of production — a hall, a region, a cluster — with success criteria agreed up front and measured weekly.

Tier 3

Enterprise Reliability Platform

Estate-wide deployment in your environment: continuous cross-layer reasoning, verified remediation workflows, and integration with your NOC, ticketing and change processes.

Questions

Make infrastructure intelligence operational.

Start with a conversation about your environment.