Knowledge hub

Practical guides for network operations teams

Field-tested advice on monitoring, observability, telemetry, automation, orchestration and security for data-center and service provider networks.

SNMP & Telemetry

SNMP Trap Design That Doesn't Flood Your NOC

A practical guide to designing SNMP trap handling (filtering, severity mapping, clear events and rate limiting) so your NOC sees real problems instead of noise.

5 min read
Observability

Cutting Alarm Noise: A Practical Guide to Event Correlation

How to reduce network alarm noise using deduplication, suppression, root-cause correlation, maintenance windows and better thresholds, with a step-by-step plan you can start this week.

4 min read
Observability

Five Grafana Dashboards Every Network Operations Team Needs

The five Grafana dashboards that give network teams the most value (NOC overview, interface health, WAN and internet edge, capacity planning and change correlation) with design tips for each.

4 min read
Observability

10 Splunk Searches Every Network Engineer Should Know

Ten practical Splunk SPL searches for network syslog (interface flaps, BGP changes, config changes, login failures, silent devices and more) with explanations you can adapt to your environment.

5 min read
SNMP & Telemetry

NetFlow vs sFlow vs IPFIX vs Streaming Telemetry: Which Should You Use?

A clear comparison of NetFlow, IPFIX, sFlow, SNMP polling and model-driven streaming telemetry, covering what each one measures, their trade-offs and how to combine them in a modern monitoring design.

4 min read
Data Center

Fault vs Performance Monitoring: How Spectrum, SevOne, Splunk and Grafana Fit Together

How fault management, performance monitoring, log analytics and visualization work together in a data-center monitoring stack, using Broadcom Spectrum, IBM SevOne, Splunk and Grafana as examples.

4 min read
Automation

Automating Device Onboarding into Monitoring with Python and REST APIs

A step-by-step approach to automating network device onboarding into monitoring platforms with Python and REST APIs, including a working script pattern with validation, idempotency and a dry-run mode.

5 min read
SNMP & Telemetry

Interface Errors and Discards Explained: What the Counters Really Mean

A practical guide to interface counters (input errors, CRC, discards, output drops, 32-bit vs 64-bit counters) covering what each one means, common causes and how to monitor them without false alarms.

5 min read
Orchestration

What Is Multi-Domain Service Orchestration? A Plain-English Guide

A plain-English explanation of multi-domain service orchestration (MDSO) covering what it is, how it works with SDN controllers, inventory and OSS/BSS, and what it takes to deploy it successfully.

4 min read
Security

Hardening Your Network Management Platform: A Security Checklist

A practical security checklist for network management and monitoring platforms covering access control, SNMPv3, credentials, segmentation, patching, logging and vulnerability scanning.

4 min read
Optical & Transport

Optical Network Alarms Explained for IP Engineers

A practical introduction to optical and OTN alarms (LOS, LOF, AIS, BDI, pre-FEC BER, OSNR, optical power) and performance monitoring for IP and data-center engineers who work alongside transport teams.

5 min read
Operations

How to Write a Method of Procedure (MOP) for Network Changes, with a Template

What a Method of Procedure (MOP) is, why it prevents outages, and a practical section-by-section template for network and monitoring platform changes, including pre-checks, rollback and verification.

4 min read