NetGuardian
NetGuardian
Agentic Solutions
Autonomous network operations & service restoration

An AI workforce that resolves Rogers broadband incidents before customers need to call.

NetGuardian detects, diagnoses, coordinates, resolves and communicates across DOCSIS and fiber plant — ten specialised agents working a single incident under human supervision, from first alarm to regulator-ready report.

Broadband subscribers
14.2M
Fiber subscribers
3.6M
DOCSIS nodes
62,000
Fiber PONs
9,400
The executive problem

A congested node pulls in a dozen teams — and none of them share a view.

At Rogers, one incident touches NOC operators, plant technicians, headend engineers, CMTS and OLT vendors, contractors, customer care and business account teams. Coordination — not repair — is the bottleneck.

  • Engineers manually correlate hundreds of modem and node alarms
  • Incident bridges consume hundreds of hours per month
  • Customers report slow speeds before the NOC sees the service group degrade
  • Truck rolls are dispatched without plant-level root cause
  • Third-party coordination is fragmented across channels
  • Communications are inconsistent and written by hand
Live demo · three incident scenarios

Pick a plant scenario, then watch the agents work

Switch between DOCSIS node congestion, a fiber PON degradation and a predicted saturation event. Press Trigger incident and watch the agent team work — including the Teams approval gate.

Scenario
Rogers ZIP 28277 · Charlotte, NC · DOCSIS node CLT-N4471
08:03:00
Network health
98.9%
Correlated alarms
0
Customers impacted
0
Incidents open
0
Plant telemetry
Upstream utilization
41%
Downstream utilization
58%
MER
37.8 dB
SNR
38.2 dB
T3 / T4 timeouts
0 / 0
Active modems impacted
0
Customer tickets
3
Ballantyne hub · 14 nodes on the affected service group
Agent activity stream
01Network Monitoring AgentWorking08:03:04

Analyzing cmts telemetry

Executive walkthrough · Microsoft Customer Zero

Microsoft ran this on its own network first.

The Global Azure Networking team deployed NetAI inside their own NOC — three agents that detect, diagnose and autonomously drive fibre cut repairs with third-party suppliers. Here is the situation, what was deployed, how it works and the measured return.

Walkthrough
Azure regions
70+
Datacenters
400+
Points-of-presence
190+
km of fibre & subsea cable
600,000+

The Global Azure Networking NOC runs one of the largest private networks on earth.

The team is responsible for continuously improving the availability and throughput of the Azure backbone. At that scale, the constraint is not engineering talent — it is the volume of manual coordination each physical fault demands.

  • A global backbone spanning 600,000+ km of terrestrial and subsea fibre, where a single span cut degrades regional capacity.
  • Fibre repairs depend on third-party suppliers in every market — each with their own NOC, language, ticketing and escalation etiquette.
  • Engineers spent hours in email threads chasing root cause, ETRs and status updates instead of engineering the network.
  • Repair quality was unverifiable without manual OTDR and optical power checks, so failed splices were only found later.
  • Siloed EMS/NMS per supplier meant no unified view, and time-to-detect depended on who was watching which screen.
The agent team

Ten agents, one incident, one accountable record.

01

Network Monitoring Agent

Watches CMTS, DOCSIS nodes, HFC amplifiers, OLTs, ONTs and cable modems across the footprint.

CMTS telemetryDOCSIS node countersHFC amplifiersOLT/ONT telemetry

Upstream utilization on DOCSIS node CLT-N4471 at 94%. MER degraded to 24.1 dB. Service impact: HIGH.

97%
02

Incident Correlation Agent

Collapses modem, node and headend alarm storms into a single actionable incident.

Alarm streamService group topologyPlant digital twin

All alarms trace to one CMTS service group behind the Ballantyne hub. 1 incident, not 640 alerts.

96%
03

Root Cause Agent

Determines what actually happened in the plant, not just what alarmed.

PNM spectrum capturesT3/T4 timeout historyMaintenance windowsWeather feeds

Noise ingress on the upstream return path combined with a congested service group — not a modem-side fault.

93%
04

Impact Analysis Agent

Quantifies subscriber, service and financial exposure in real time.

CRMSLA contractsBillingPlant inventory

3,200 broadband subscribers · 180 SMB · 6 enterprise circuits. Revenue exposure $118K/day, SLA exposure $40K.

95%
05Human approval

Field Dispatch Agent

Finds the closest qualified plant technicians and drafts the dispatch request.

Workforce managementAzure MapsTechnician skills matrix

Maintenance tech MT-14 (17 min ETA) qualified for ingress sectorisation. Dispatch drafted — awaiting NOC manager approval.

91%
06

Vendor Coordination Agent

Engages CMTS/OLT vendors, contractors and the construction company.

Vendor portalsMicrosoft GraphServiceNowDynamics 365

CMTS vendor case opened, amplifier supplier notified, contractor case #CN-88213 created for the ingress source.

94%
07

Customer Communications Agent

Gets ahead of the call spike with tailored proactive messaging.

Subscriber baseBusiness account teamsNotification platform

3,200 broadband SMS queued, 6 enterprise emails drafted, executive brief prepared for VP Network Operations.

92%
08

Restoration Planning Agent

Evaluates remediation options and recommends the fastest safe path.

Capacity modelService group planningSpares inventory

Rebalance the service group and shift OFDMA profiles now (9 min); schedule a node split for permanent headroom.

96%
09

Recovery Verification Agent

Confirms the incident is genuinely resolved before closure.

MER/SNRThroughputT3/T4 countsCustomer tickets

MER back to 37.4 dB, T3 timeouts zero, throughput restored, ticket inflow down 88%. Closure recommended.

98%
10

Post-Incident Agent

Produces the paperwork that normally takes days.

Full incident recordAgent audit trailAgent365 governance log

Executive summary, RCA, node-split capital request, SLA report and regulator filing generated in 4 minutes.

97%
Human experience throughout the incident

Humans supervise the workforce. They no longer coordinate it.

NOC team

  • Teams alerts
  • Teams approvals
  • Recommended actions
  • Incident summaries

Plant technicians

  • Mobile work orders
  • Maps & routing
  • Ingress & splice instructions
  • Node and amplifier records

Vendors & contractors

  • Automated requests
  • Escalations
  • Priority levels
  • Evidence packages

Executives

  • Business impact
  • Subscriber impact
  • Financial exposure
  • Restoration timelines
Executive dashboard

Rogers broadband and fibre health in one view.

Broadband subscribers
14.2M
Fiber subscribers
3.6M
DOCSIS service availability
99.94%
Fiber service availability
99.98%
Average throughput
612 Mbps
Customer impacted events
38 / week
MTTR
72 min
Node congestion index
0.21
PON health score
94.6
Network reliability score
96.8
Service assurance copilot

Ask the network in plain language. The copilot reasons over CMTS, node, OLT and ONT telemetry and answers with the impacted service groups, PONs and customers.

  • Show me the worst performing DOCSIS nodes.
  • Why are customers in Charlotte experiencing slow speeds?
  • Which fiber PONs are at risk of service degradation?
  • What outages are impacting enterprise customers?
  • Predict tomorrow's network hotspots.

DOCSIS node congestion

A congested service group behind the Ballantyne hub degraded broadband for 3,200 subscribers. Agents isolated the node, sectorised the ingress and queued a node split before the evening peak.

3,200 subscribers protected

Fiber distribution issue

Optical power degradation on a PON splitter serving a business district was traced, rerouted and repaired before SLA credits were triggered on 22 commercial circuits.

22 business circuits held to SLA

Predicted degradation avoided

Forecast utilization trends showed a node reaching saturation within 6 days. Capacity was expanded during a maintenance window with zero customer-impacting minutes.

0 customer-impacting minutes
Executive ROI

Before and after NetGuardian at Rogers.

MetricTraditional NOCAI driven
Mean time to detect25 min< 1 min
Mean time to diagnose90 min3 min
Mean time to restore4.5 hrs60–90 min
Major incident bridge time150+ hrs/monthNear zero
NOC productivityBaseline+40%
Truck-roll accuracyBaseline+35%
Broadband call volumeVery high spike−30%
Annual benefits · example telco
  • Reduced outage costs$15M
  • Reduced SLA penalties$4M
  • Reduced truck rolls$8M
  • NOC efficiency$6M
  • Call centre avoidance$10M
Total annual benefit
$43M+

Fewer outages, faster restoration, lower operating cost and a measurably better customer experience — with a NOC that scales without adding headcount.

Microsoft architecture

Built on the stack the enterprise already runs.

Experience
Layer
Microsoft TeamsOperations DashboardMobile Technician AppExecutive Portal
AI Agent
Layer
Azure AI FoundryMonitoring · Correlation · RCADispatch · Vendor · CustomerRecovery · Post-Incident
Intelligence
Layer
Azure OpenAI — incident reasoningAzure AI Search — runbooksMicrosoft Graph — TeamsAzure Maps · AI Speech
Data
Layer
Azure Data ExplorerMicrosoft FabricEvent HubDigital Twins · PNM & OLT telemetry
Workflow
Layer
Power AutomateLogic AppsServiceNowDynamics 365
Governance
Layer
Agent365 human approvalsAudit trailsRole-based accessRegulatory compliance

NetGuardian is not a monitoring tool.

It is an autonomous workforce of network operations agents that detect DOCSIS and PON incidents, correlate alarms, diagnose plant root causes, coordinate technicians, engage vendors, inform customers, recommend remediation and generate compliance reports — while humans stay accountable for every decision.

Network reliabilityCustomer experienceOperational efficiencyFinancial performance