Most monitoring deployments fail for the same reason: everything is monitored and nothing is meaningful. I build NMS platforms around a service model, so alerts map to business impact rather than to interfaces.
Deployment covers collector sizing and placement, SNMP/ICMP/API polling, NetFlow and sFlow ingestion, syslog and trap normalisation, credential management, and the escalation matrix that decides who wakes up.
I tune thresholds against two weeks of real baseline data before handover. Noise is a design defect, not an operational fact of life.
- Actionable alert volume, typically 80 to 95 percent below the untuned baseline
- Mean time to detect under 60 seconds for link and node loss
- Capacity trending your finance team can read
Which NMS do you recommend?+
It depends on estate size and who operates it. Zabbix for mixed infrastructure with a small ops team, Prometheus and Grafana where the estate is container heavy, LibreNMS where the priority is fast, low effort network coverage.
Can you deploy on premise only?+
Yes. Fully air gapped deployments are supported, including local repository mirroring.