How Telecom Service Impact Analysis Works: From Resource Alarms to Customer Impact

In the field of Telecom Assurance, delivering reliable telecommunications services to customers requires restoring faults in the network as fast as possible, especially if they are customer service impacting. It is preferable to be able to identify a probable fault before it occurs, thus preventing service impact on customers. 

In addition to the unplanned outage scenario described above, maintenance of services and resources as planned outages (hardware and software upgrades, data fixes) also requires identifying impacted customer services so customers can be informed of planned service disruptions, especially for enterprise service deliveries. 

1 Introduction: Why Service Impact Matters

So the question to ask in such cases are

  • A customer experience or service is degraded – what underlying resource caused it? This is a top-down diagnosis.
  • A resource has failed – which services and customers are now affected? This is a bottom-up containment. 

These questions are relevant because in telecommunication environments,

  • delivered services are built on layered technical components
  • one low-level fault can result in many alarms from network elements and network element management systems
  • operators need to understand both the root cause of the fault and the service impact to business/customer. 
  • operators need to understand how to bound and govern the impact of deliberately taking down a service or resource out of service for maintenance, upgrade or data fix. 

Service Impact Analysis is the ability for a carrier to keep its operational-level reality aligned with its commercial commitments in real time. It provides value by translating network faults into a breached SLA that, in turn, may risk customer churn or regulatory compliance. 

For planned work, SIA helps schedule maintenance and upgrades around business exposure instead of causing it. 

2 The Core Problem: Why Impact Analysis is Hard

2.1 Telecom systems are layered

A failure at a low layer is visible at that layer but matters at the higher layer. A customer buys a Product via a Subscription to a Plan (Consumer space). The Product is realised by one or more Services (Customer-Facing Services at the Business Layer, Resource-Facing Services at the Technical Layer). The Services run on physical, logical and virtual resources. e.g., a router, VNF, a fibre path, a database). 

2.2 The dependencies are directional, and not every dependency is equal

Product, Service and Resource dependencies can be hard, as in load-bearing, or soft as in alternate paths, diverse paths, HA clusters. 

Such dependencies are therefore best represented as a graph, not a flat list, requiring distinguishing dependency types (hard/impacting vs soft/informational).

Critically, whether a soft dependency is currently load-bearing is a fact about the live network, not a fact about the design. The same edge is informational while the primary path is carrying traffic and impacting the moment it fails. Section 5.4 returns to where that live state has to come from, and what goes wrong when it is stale.

2.3 One fault creates many symptoms

An extreme example of this is a fibre cut that results in an alarm storm. Alarms are generated on the physical link, the logical links over it, the network services and the customer-facing services in near real time. These alarms have to be correlated back to the originating fault as Root Cause Analysis. So RCA is the process of collapsing this fan-out back to one or a handful of originating faults. 

Impact Analysis, on the other hand, is the opposite process. A known fault is propagated forward to enumerate the affected services and customers. 

RCA and SIA then require walking the same dependency graph in opposite directions. 

2.4 Different audiences need different expressions of impact

Engineers need to know which resource failed in the NOC.

Operations managers in the SOC need to know which services are degraded and how severely. 

Customer-facing teams need to know which customers and SLAs are at risk and by how much (financially, contractually). 

2.5 Impact analysis runs in two temporal directions, not only in two structural directions

As mentioned in sections 2.3 and 2.4, for unplanned outages, impact analysis walks the dependency graph either bottom-up (a fault occurred – what does it affect) or top-down (a service is degraded – what caused it). 

The second orthogonal axis is time. 

Reactive impact analysis answers the question, “What is affected by the fault that has already happened?” 

Proactive impact analysis answers “what would be affected if this resource or service was deliberately taken out of service?” Typically asked before every planned maintenance window, hardware/software upgrade, or data-correction activity. 

Proactive impact analysis also answers “what would likely be affected based on system behaviour known and given events that have occurred”. 

Thus, Proactive Impact analysis comes in two distinct forms.

All three use the identical graph and identical information distinctions for impact analysis identified in section 2.2; the difference is whether the triggering event is observed reactively (an alarm is generated), proactively (effect of failure = a change is proposed in the future), or inferred as likely. 

SIA should support all three from one shared graph.

3 The Conceptual Backbone: Resource -> Service -> Customer

As mentioned in section 2.1 Telecom services are layered with Resources at the bottom layer allocated and consumed by Services at the middle layer that know how to deliver Products that the Customer ordered at the Customer Layer.

What that means a single fault (at the resource layer) get translated three times before it can be acted on correctly

  • once into “what broke”
  • once into “so what’s degraded”
  • and once into “now who’s affected and by how much”

Service Impact Analysis should make these tree translations fast, accurate and automatic.

4 How TM Forum SID Models Service Impact

The following is based on TMForum’s SID Model.

4.1 SID does not define one single “SIA ABE”

TM Forum expresses SIA through a pattern across the domains identified in section 3 rather than via a single dedicated entity. 

4.2 Resource layer : Resource Alarm

The Resource layer includes resources like a router, a VNF, a fibre path, and a database. Any degradation or failure of a component at this layer is detected as an alarm, acting like a technical signal that something is not running as normal in the component. 

TM Forum’s SID information model defines a ResourceAlarm with a severity, a probable cause, raise/clear timestamps, and two attributes that exist purely to initiate the translation process: serviceAffecting (does this alarm actually touch a service, or is it noise?) and potentialRootCauseIndication (is this alarm a candidate root cause, or a downstream symptom of something else?). The moment an alarm fires, it’s already being asked to answer questions that belong one layer up.

On the process side, eTOM’s Resource Trouble Management (1.5.8) governs what happens next: survey and analyse the raw alarm stream, filter and correlate it, localise the root cause, and — critically — decide whether this resource-level event needs to be escalated upward.

Correlation – the translation step from noise to signal

Any failure at the resource layer within an NE, for example, rarely produces a single alarm. In many instances, a flood of near-duplicate alarms is produced that end up on the alarm management system screens. De-duplication and correlation turn that flood into a signal, grouping alarms that share one underlying cause and surfacing a root-cause alarm that explains what happened. In a sense, the raw operation telemetry is becoming impact-aware. 

4.3 Service Layer: Service Problem

So what comes out of correlation is not a resource alarm anymore; it’s a service problem. SID models this as ServiceProblem, and it explicitly quantifies the blast radius with an affectedServiceNumber attribute: not just “something is wrong,” but “this many service instances are touched.”

eTOM’s Service Problem Management (1.4.6) describes the process for this tier, and it is similar to the workflow of the resource tier,  though one level up: Survey & Analyse → Localise → Correct & Resolve → Track & Manage → Report → Close. This also describes the quantitative side of the model for the first time — individual resource-level KPIs (meaningful to an engineer) get aggregated into KQIs, service-level quality indicators that are meaningful to a customer and measurable against an SLA.

4.4 Customer Layer: Customer Problem

The final translation is the one the rest of the business actually cares about. A service problem becomes a customer problem — SID’s CustomerProblem entity, which can be raised reactively (a customer complains) or proactively (the provider’s own analytics catch it first). 

Critically, it supports hierarchical navigation both ways: many customer-reported problems point to a single shared root cause (fan-in), and one customer’s problem can decompose into multiple sub-problems if it turns out to point to many underlying issues (fan-out).

This is the tier where “a router failed” finally becomes “these twelve enterprise customers just breached their availability SLA” — the language the business, not the network, speaks. eTOM’s Problem Handling processes own this layer, and it’s measured against the KQIs, SLAs, and SLS thresholds that were aggregated on the way up.

5 From Standards to Computation: The Assurance Graph 

5.1 The Implementation Gap

TM Forum’s ABE/process pattern indicates that impact must propagate through a graph, but does not fully specify how to compute it algorithmically. IETF RFC 9417 (Service Assurance for Intent-Based Networking Architecture) fills that gap and is directly useful as an implementation model for the SID/eTOM pattern above — it even uses the term “Service Impact Analysis” by name to describe the bottom-up direction.

5.2 The assurance graph: dependency structure, no computation yet

RFC 9417 formalises the dependency picture from section 2.2 as an assurance graph: a Directed Acyclic Graph whose nodes are service instances and subservices — any independently assurable part of the system, such as an interface, a routing protocol instance, or a tunnel — and whose edges are typed dependencies.

An impacting dependency means the child’s health score directly reduces the parent’s, and the child’s symptoms are carried upward as the parent’s impacting reasons.

An informational dependency means the child’s health does not reduce the parent’s score, but its symptoms are still surfaced for context. A healthy backup path is informational to the primary path — right up until failover, at which point the relationship becomes load-bearing.

At this point, the graph is pure structure. It says what depends on what, and how hard each dependency is. It does not yet say how to turn any of that into a number. That is the next step.

5.3 The expression graph: turning structure into a health score

Once the dependency structure exists, RFC 9417 compiles it into a second DAG — the expression graph — in which the leaves are raw metrics collected from the network, the internal nodes are operations applied to them (threshold comparisons, aggregations such as MIN or weighted average, conditional branches), and the root is the health score of the service or subservice.

Each node carries a health score (0-100, or -1 if it cannot be computed, for example because the underlying metric is unavailable) and, wherever the score is below maximum, at least one symptom explaining why.

So the expression graph encodes all computations needed to determine the health status (a score from 0 to 100) of a service or sub-service. It is the operational translation of the assurance graph – converting dependency relationships and metric data into an executable health calculation. 

The pattern is Assurance Graph (service dependencies) → Expression Graph (computational recipe) → Health Score(numerical result).

5.4 Where the graph actually comes from: design-time structure vs runtime state

Sections 5.2 and 5.3 describe what the graph is and what it computes. Neither answers the question that decides whether any of it can be trusted: where does the graph come from, and what keeps it true? An assurance graph is not one dataset. It is a static skeleton with a live overlay, and the two are sourced from different systems on entirely different timescales.

Design time — does this edge exist?

  • Sourced from the Product Catalogue (productSpecificationRelationship, capturing prerequisite, dependency and mutual exclusivity between product specifications), the Service Catalogue (the CFS to RFS “vertical links” that decompose a customer-facing service into resource-facing services), and the Logical-to-Physical ResourceRelationship entities that bridge Service Inventory and Resource Inventory at the point an order is actually orchestrated.
  • TMF Open API surface: TMF620 (Product Catalogue), TMF633 (Service Catalogue), TMF638 (Service Inventory), TMF639 (Resource Inventory).
  • This data is comparatively stable. It changes when someone designs a product or fulfils an order — not second by second.

Run time — does this edge matter right now?

  • Sourced from the current status held in resource and network inventory, from controller and orchestrator state, and above all from active/standby status on redundant paths.
  • This data is volatile. A failover changes it in seconds, and nothing in a service catalogue knows it has happened.

Why the distinction is load-bearing rather than academic

  • The impacting versus informational classification introduced in 5.2 is not a design-time property. It is a runtime one. An edge to a backup path is informational only for as long as the primary is actually carrying traffic. The catalogue can tell you the backup path exists; only live state can tell you whether it is currently doing any work.
  • The failure mode is specific, and it is dangerous. Suppose a service is recorded as ACTIVE_STANDBY but has quietly degraded to SINGLE — the standby was consumed by earlier work and never restored. An impact record computed against the stale state reports “blast radius: minimal, the standby takes over.” The truth is “complete outage.” That is not simply stale data. It is wrong in exactly the direction that gets a maintenance window approved, which should have been refused.

What this means for the design

  • Do not resolve design-time and runtime in a joint query. Doing so inherits every upstream system’s downtime and latency, and an assurance graph that cannot answer while inventory is unreachable is precisely the thing it was needed for.
  • Hold one denormalised, versioned record per entity: the dependency structure as an array of typed edges, and the volatile state as fields on that same record, refreshed by change data capture from inventory and from the controllers.
  • Stamp every computed health score with the inventory version it was computed against, and compare the two at read time. Where the graph has moved on since the score was computed, the consumer should be told the answer is stale — not served a confident number quietly built on last week’s topology.
  • Treat the graph as a living, versioned asset, updated at order-orchestration time and at every topology change, rather than something reconstructed reactively in the middle of an incident.

5.5 When the binary edge is not enough: property graphs

  • The RFC 9417 assurance graph makes two simplifications. Both hold well enough for deterministic health scoring, and both start to strain elsewhere.
  • The edge carries too little information. Impacting versus informational is a single bit. Real dependency edges carry far more than that: the redundancy level sitting behind them, the latency each one contributes, the vendor and support regime of the component, the commercial cost of that particular path failing. A property graph generalises the edge into a typed, weighted, multi-attribute relationship, which lets severity and priority scoring vary by relationship type and weight rather than only by whether a path exists at all.
  • The acyclicity is a claim about the model, not about the network. Real topologies contain genuine cycles — rings, ECMP groups, mutually protecting paths, or simply two teams each modelling one direction of an interface-to-link dependency. RFC 9417’s answer is to detect strongly connected components and collapse each into a single synthetic node. That preserves reachability, so impact still propagates, but the dependency detail inside the cycle is lost. A property graph tolerates cycles natively and enforces acyclic traversal at query time instead. It also gives native path-existence queries — “is there still a path from A to B if I remove this node?” — which is precisely the redundancy question 5.4 showed you cannot afford to get wrong.
  • This is an enrichment, not a replacement. Keep the DAG as the deterministic, auditable backbone: when a health score has to justify an SLA credit, being able to trace it back through the graph to the raw metric that caused it is worth more than expressive power. Add property-graph semantics where the binary edge demonstrably fails — redundancy-aware blast radius, cost-weighted severity, cyclic topology — rather than as a default.
  • Further modelling techniques sit beyond this: Bayesian networks for probabilistic risk under partial evidence, fault trees for single-point-of-failure discovery, causality graphs and GNNs for learning undocumented dependencies, hypergraphs for shared-fate infrastructure such as a common power feed or duct. Each answers a different proactive question, and each carries real maintenance and skills cost. May be discussed in future posts.