Skip to content
Netlume

Product

The investigation layer for infrastructure engineers.

Netlume combines a live model of your infrastructure with AI agents that investigate like experienced engineers: follow the topology, collect the right evidence, test hypotheses and explain the result — with every step visible and every change approved.

Incident lifecycle

Detect to resolve, with a clear hand-off to people.

Netlume handles the repetitive, evidence-gathering stages of an incident at machine speed. Decisions that change production stay with your operators.

  1. 1

    Detect

    Anomaly from telemetry, alarms, probes or a question from an engineer.

  2. 2

    Collect

    Targeted, read-only collection from devices on the affected topology.

  3. 3

    Correlate

    Align routing events, interface counters, logs and changes on one timeline.

  4. 4

    Investigate

    Test hypotheses hop by hop; discard what the evidence rules out.

  5. 5

    Explain

    Root cause candidate with confidence, evidence and reasoning.

  6. 6

    Recommend

    Remediation, blast radius, verification plan and rollback.

  7. 7

    Approve

    An authorized operator approves — or rejects — the exact change.

  8. 8

    Verify

    Re-run evidence checks and confirm the fix holds.

  9. 9

    Resolve

    Close with a complete, auditable record of what happened.

Netlume performs Operator decides

Investigation Engine

Hypotheses, tested against evidence.

The engine treats every incident as a set of competing explanations. It plans targeted, read-only collection to confirm or reject each one, weighs the evidence and keeps going until a candidate clearly stands out — or tells you that it doesn't.

Results always include the evidence trail, so an engineer can verify the conclusion in minutes.

  • Plans what to inspect, and where
  • Ranks competing explanations
  • Records what was ruled out
  • Confidence tied to evidence

Hypotheses · INC-4127

4 tested
  1. H1ECMP redistribution → TC3 congestion on leaf-07 Et12

    92%
    supported · 7 evidence
  2. H2Degraded optic or cabling on leaf-07 Et12

    4%
    ruled out · 2 evidence
  3. H3Host NIC firmware regression on GPU-WORKER-042

    2%
    ruled out · 2 evidence
  4. H4PFC storm / pause deadlock

    2%
    ruled out · 3 evidence

Root cause candidate

evidence-linked

East-west congestion following ECMP path redistribution resulted in queue saturation on leaf-07.

Confidence

92%

Agent Architecture

Collectors near the network. Reasoning in the platform.

Lightweight collectors run close to your devices and speak SSH, NETCONF, RESTCONF, gNMI, SNMP, syslog and vendor or cloud APIs. Specialised agents handle topology modelling, investigation and verification, each with a narrow, auditable set of tools.

Collectors only use the commands and paths you allow — there is no general-purpose shell access.

  • Deployed per site, region or cloud
  • Outbound-only connectivity
  • Read-only credentials by default
  • Specialised agents with narrow tools

Agents & collectors

42 / 42 online
  • collector-dc-east-1a

    212 devices · AI Fabric A

    Collector
  • collector-dc-east-1b

    486 devices · DC East EVPN

    Collector
  • collector-backbone

    138 devices · SR-MPLS core

    Collector
  • collector-cloud

    4 accounts · 3 providers

    Collector
  • topology

    21,904 edges · 26 networks

    Model
  • investigator

    3 investigations running

    Reasoning
  • verifier

    1 check scheduled

    Verification

All collectors read-only · outbound connections only

Topology Context

Every investigation starts from how traffic actually flows.

Netlume continuously builds a graph of your environment: physical adjacency, IGP and BGP, EVPN/VXLAN and MPLS overlays, VRFs, ECMP groups and the services that depend on them.

When something fails, the investigation is scoped automatically to the devices that can actually affect the symptom — not every device that raised an alarm.

  • LLDP and physical links
  • Underlay and overlay routing
  • ECMP groups and MLAG pairs
  • Service and workload dependencies

Incident scope · AI Fabric A

auto-scoped
SPECTRUM-01ECMP set changedSPECTRUM-02nominalLEAF-05nominalLEAF-06CNP ↑LEAF-07Et12 q3 saturatedLEAF-08nominalGPU-041nominalGPU-042step +23%GPU-043nominal
  • Healthy
  • Warning
  • Critical
  • Incident path

Telemetry Correlation

Signals from different systems, on one timeline.

Netlume aligns telemetry, routing events, logs and probes from different devices and vendors on a shared timeline, then looks for the causal order: what changed first, and what followed.

That is how a 20:27 ECMP change gets linked to a 20:31 packet-loss alarm four layers away.

  • Streaming telemetry and counters
  • Routing and control-plane events
  • Syslog and alarms
  • Probes and service metrics

Correlated telemetry · 20:15 – 20:45 UTC

4 sources
ECMP members · SPECTRUM-01 → leaf tierECN-marked % · leaf-07 Et12 TC3PFC pause ms/s · leaf-07 Et12Probe loss % · GPU-042 ↔ GPU-04120:27 BGP path change20:31 loss detected

Configuration Intelligence

Know what changed, and whether it matters.

Netlume keeps configuration history across vendors and understands it semantically — policies, neighbors, VRFs, ACLs and QoS — rather than as plain text.

It detects drift from intended state, highlights risky changes and connects recent changes to the incidents that follow them.

  • Multi-vendor configuration parsing
  • Drift from golden or intended state
  • Change-to-incident correlation
  • Pre-change impact review

Configuration drift · campus access

3 devices drifted
  • ARUBA-ACC-114
  • ARUBA-ACC-117
  • ARUBA-ACC-121
# intended (golden: campus-access-v14) vs. running
  ntp server 10.10.0.10
- ntp server 10.10.0.11
+ ntp server 10.10.0.12
  logging 10.10.5.20 severity warning
- aaa authentication login default group tacacs local
+ aaa authentication login default local

Introduced

manual CLI · 2 days ago · no change record

Risk

TACACS bypass · time skew in logs

Network Path Analysis

Trace any flow, hop by hop, across vendors.

Resolve the real forwarding path for a flow across leaf, spine, border, WAN, DCI and cloud — using RIB, FIB, overlay and ECMP state from each device — then inspect every hop for errors, drops and congestion.

  • Source to destination across domains
  • ECMP-aware path resolution
  • Per-hop counters, queues and optics
  • VRF, overlay and policy aware

10.20.4.15 → 10.44.18.0/24 · VRF PROD

6 hops
  1. Source

    app-​web-​17

    10.20.4.15

  2. Leaf

    ARISTA-​LEAF-​03

    EOS · VTEP 10.0.250.3

  3. Spine

    CISCO-​N9K-​SPINE-​02

    NX-OS

  4. Border Leaf

    ARISTA-​BL-​01

    EOS · VRF PROD

  5. Router

    JUN-​MX-​EDGE-​01

    Junos · et-0/0/2

    Suspect hop
  6. Destination

    10.44.18.0/24

    dc-west-2 · payments-db

Incident Correlation

Many alarms. One incident. A clear blast radius.

Symptoms that share a cause are grouped into a single incident using topology and timing, not just text similarity. Each incident shows which devices, paths, services and workloads are affected — and which are not.

  • Topology-aware de-duplication
  • Service and job impact
  • Linked to changes and events
  • Fewer pages, better context

Merged signals

  • BGP path change detected

    NVIDIA-SPECTRUM-01 · bgp

    20:27:14
  • ECMP distribution changed 4 minutes before incident

    fabric · ecmp-groups

    20:27:15
  • ECN queue saturation detected

    ARISTA-LEAF-07 · Et12 · TC3

    20:33:02
  • PFC pause duration increased

    ARISTA-LEAF-07 · Et12 · prio 3

    20:33:08
  • Packet drops detected on leaf-07

    ARISTA-LEAF-07 · counters

    20:33:41

Incident INC-4127 · blast radius

5 signals merged
Devices
23
Paths
18
Services
1
Jobs
1
  • GPU Training Clusterdegraded
  • train-llm-run-0412step time +23%
  • Storage replication · NVMe-oFnot affected
  • DC East EVPN tenantsnot affected

Human-in-the-loop Remediation

Recommend first. Change only with approval.

Netlume never applies destructive changes on its own. Each deployment chooses how far automation may go — and every step is recorded.

  1. 01Read Only

    Default for every connector

    Collect state, logs, telemetry and configuration. No writes to any device.

  2. 02Recommend

    Netlume generates

    Propose a remediation with evidence, blast radius and a rollback plan.

  3. 03Approval Required

    Operator decides

    A named operator with the right role must approve the exact change.

  4. 04Execute Approved Action

    Scoped credentialoff by default

    Apply only the approved change, scoped to the approved devices and window.

  5. 05Verify Result

    Netlume verifies

    Re-run the evidence checks and confirm the symptom is gone.

  6. 06Rollback Guidance

    Operator decides

    If verification fails, present the prepared rollback and its expected effect.

Approval required

rb-2291 · INC-4126

Revert CHG-2291 on JUN-MX-EDGE-01

Restore local-preference 100 for DCI-EXPORT so 10.44.18.0/24 prefers JUN-MX-EDGE-02 while the et-0/0/2 optic is replaced.

[edit policy-options policy-statement DCI-EXPORT term PREFER-EDGE-01 then]
-    local-preference 200;
+    local-preference 100;
Blast radius
1 device · 1 term
Prefixes moved
1 (/24)
Method
commit confirmed 5
Rollback
automatic if unconfirmed

Verification plan

  • Best path for 10.44.18.0/24 via 172.16.0.9
  • Probe loss < 0.1% for 5 minutes
  • No new BGP flaps on ARISTA-BL-01

Requires role: network-lead · 1 of 1 approvals

Reject Approve change

Audit log · INC-4126

immutable · exportable
  1. 14:03:52

    investigator · Generated recommendation

    Revert CHG-2291 on JUN-MX-EDGE-01

  2. 14:04:30

    netops-oncall · Requested approval

    Rollback candidate rb-2291

  3. 14:06:12

    network-lead · Approved change

    Rollback candidate rb-2291 · window 5 min

  4. 14:06:20

    executor (scoped) · Applied commit confirmed 5

    JUN-MX-EDGE-01 · 1 policy term

  5. 14:07:41

    investigator · Verified result

    Best path → EDGE-02 · probe loss 0.0%

  6. 14:08:02

    netops-oncall · Resolved incident

    INC-4126 · root cause linked

Operator view

The whole estate at a glance.

Health across networks, active incidents, device inventory and agent activity — the starting point for every shift.

  1. Overview
  2. /All networks
OC
OverviewTopologyDevicesIncidentsInvestigationsAgentsChangesConfigurationsAI Fabric

Infrastructure overview

Last 24h26 networks

Infrastructure Health

98.7%

+0.2 pts 24h

Managed Devices

3,842

+36 this week

Active Incidents

7

2 high · 5 medium

Networks

26

9 sites · 4 clouds

AI Fabric Health

Healthy

1 leaf under watch

Investigations Today

148

median 3m 40s to RCA

Active incidents

7 open
  • High

    Packet loss · GPU workers ↔ leaf fabric

    INC-4127 · 6m ago

  • High

    Loss to 10.44.18.0/24 via DCI

    INC-4126 · 38m ago

  • Medium

    BFD flaps · PE-04 ↔ P-02

    INC-4121 · 2h ago

  • Medium

    Config drift · 3 campus access switches

    INC-4118 · 5h ago

  • Medium

    Elevated optics temp · BL-02 Et7

    INC-4115 · 9h ago

Network health

devices · score
  • AI Fabric A

    RoCEv2 · BGP unnumbered

    21296.1%
  • DC East EVPN

    EVPN/VXLAN · MP-BGP

    48699.4%
  • Backbone

    SR-MPLS · IS-IS · L3VPN

    13899.1%
  • Internet Edge

    BGP · 6 transit peers

    1299.8%
  • Campus HQ

    OSPF · MC-LAG

    64498.9%

Devices

3,842
Managed devices (demo data)
DevicePlatformRoleSiteCollectionCPUHealth
ARISTA-LEAF-07Arista · EOSGPU leafdc-east-1gNMI · SSH · syslog
38%
Critical
NVIDIA-SPECTRUM-01NVIDIA · Cumulus LinuxAI spinedc-east-1gNMI · SSH
22%
Warning
JUN-MX-EDGE-01Juniper · JunosEdge routerdc-east-1NETCONF · SNMP
21%
Warning
CISCO-N9K-SPINE-01Cisco · NX-OSSpinedc-east-1gNMI · NX-API
14%
Healthy
NOKIA-CORE-01Nokia · SR OSDCI corepop-sgngNMI · NETCONF
17%
Healthy
CISCO-ASR-PE-04Cisco · IOS-XRPE routerpop-hangNMI · SSH
26%
Healthy
SONIC-LEAF-09SONiC · SONiCCompute leafdc-east-1gNMI · SSH
9%
Healthy

Agent activity

  • Ran 4 read-only commands on ARISTA-LEAF-07

    collector-dc-east-1a · 20:35:02

  • Ranked 4 hypotheses for INC-4127 · top 92%

    investigator · 20:34:47

  • Subscribed gNMI queue stats on 6 interfaces

    collector-dc-east-1b · 20:33:10

  • ECMP group change on SPECTRUM-01 recorded

    topology · 20:27:15

  • Drift detected on 3 campus access switches

    config-watch · 20:12:09

Illustrative product UI with simulated demo data. Not real customer data or statistics.

Turn infrastructure complexity into clear answers.

See Netlume investigate a real-world failure scenario on a multi-vendor fabric, and talk with the engineers building it about your environment.