Network Monitoring Agent
Watches CMTS, DOCSIS nodes, HFC amplifiers, OLTs, ONTs and cable modems across the footprint.
Upstream utilization on DOCSIS node CLT-N4471 at 94%. MER degraded to 24.1 dB. Service impact: HIGH.

NetGuardian detects, diagnoses, coordinates, resolves and communicates across DOCSIS and fiber plant — ten specialised agents working a single incident under human supervision, from first alarm to regulator-ready report.
At Rogers, one incident touches NOC operators, plant technicians, headend engineers, CMTS and OLT vendors, contractors, customer care and business account teams. Coordination — not repair — is the bottleneck.
Switch between DOCSIS node congestion, a fiber PON degradation and a predicted saturation event. Press Trigger incident and watch the agent team work — including the Teams approval gate.
Analyzing cmts telemetry…
The Global Azure Networking team deployed NetAI inside their own NOC — three agents that detect, diagnose and autonomously drive fibre cut repairs with third-party suppliers. Here is the situation, what was deployed, how it works and the measured return.
The team is responsible for continuously improving the availability and throughput of the Azure backbone. At that scale, the constraint is not engineering talent — it is the volume of manual coordination each physical fault demands.
Watches CMTS, DOCSIS nodes, HFC amplifiers, OLTs, ONTs and cable modems across the footprint.
Upstream utilization on DOCSIS node CLT-N4471 at 94%. MER degraded to 24.1 dB. Service impact: HIGH.
Collapses modem, node and headend alarm storms into a single actionable incident.
All alarms trace to one CMTS service group behind the Ballantyne hub. 1 incident, not 640 alerts.
Determines what actually happened in the plant, not just what alarmed.
Noise ingress on the upstream return path combined with a congested service group — not a modem-side fault.
Quantifies subscriber, service and financial exposure in real time.
3,200 broadband subscribers · 180 SMB · 6 enterprise circuits. Revenue exposure $118K/day, SLA exposure $40K.
Finds the closest qualified plant technicians and drafts the dispatch request.
Maintenance tech MT-14 (17 min ETA) qualified for ingress sectorisation. Dispatch drafted — awaiting NOC manager approval.
Engages CMTS/OLT vendors, contractors and the construction company.
CMTS vendor case opened, amplifier supplier notified, contractor case #CN-88213 created for the ingress source.
Gets ahead of the call spike with tailored proactive messaging.
3,200 broadband SMS queued, 6 enterprise emails drafted, executive brief prepared for VP Network Operations.
Evaluates remediation options and recommends the fastest safe path.
Rebalance the service group and shift OFDMA profiles now (9 min); schedule a node split for permanent headroom.
Confirms the incident is genuinely resolved before closure.
MER back to 37.4 dB, T3 timeouts zero, throughput restored, ticket inflow down 88%. Closure recommended.
Produces the paperwork that normally takes days.
Executive summary, RCA, node-split capital request, SLA report and regulator filing generated in 4 minutes.
Ask the network in plain language. The copilot reasons over CMTS, node, OLT and ONT telemetry and answers with the impacted service groups, PONs and customers.
A congested service group behind the Ballantyne hub degraded broadband for 3,200 subscribers. Agents isolated the node, sectorised the ingress and queued a node split before the evening peak.
Optical power degradation on a PON splitter serving a business district was traced, rerouted and repaired before SLA credits were triggered on 22 commercial circuits.
Forecast utilization trends showed a node reaching saturation within 6 days. Capacity was expanded during a maintenance window with zero customer-impacting minutes.
| Metric | Traditional NOC | AI driven |
|---|---|---|
| Mean time to detect | 25 min | < 1 min |
| Mean time to diagnose | 90 min | 3 min |
| Mean time to restore | 4.5 hrs | 60–90 min |
| Major incident bridge time | 150+ hrs/month | Near zero |
| NOC productivity | Baseline | +40% |
| Truck-roll accuracy | Baseline | +35% |
| Broadband call volume | Very high spike | −30% |
Fewer outages, faster restoration, lower operating cost and a measurably better customer experience — with a NOC that scales without adding headcount.
It is an autonomous workforce of network operations agents that detect DOCSIS and PON incidents, correlate alarms, diagnose plant root causes, coordinate technicians, engage vendors, inform customers, recommend remediation and generate compliance reports — while humans stay accountable for every decision.