Skip to content
Netlume

AI Infrastructure

Built for the networks behind AI.

GPU clusters are only as fast as the fabric between them. Netlume understands GPU cluster networking — rail-optimized topologies, RoCEv2 and InfiniBand, ECN, PFC and DCQCN — and connects fabric behaviour to the workloads that depend on it.

AI Fabric A· rail-optimized · RoCEv2 · 400G
Fabric healthy1 leaf under watch
spine-01spine-02spine-03spine-04leaf-00rail 0leaf-01rail 1leaf-02rail 2leaf-03rail 3leaf-04rail 4leaf-05rail 5leaf-06rail 6leaf-07rail 7GPU-WORKER-041GPU-WORKER-042GPU-WORKER-043GPU-WORKER-044GPU n → rail leaf n · 1 NIC per GPU

Congestion signalled to RoCEv2 senders (DCQCN)

leaf-00
leaf-01
leaf-02
leaf-03
leaf-04
leaf-05
leaf-06
leaf-07
Et1Et12Et24
lowhigh · % of TC3 packets marked

Illustrative product UI with simulated demo data. Not real customer data or statistics.

Why it is different

AI fabrics fail in ways traditional tools miss.

Most monitoring was designed for best-effort networks with averaged counters. AI fabrics need lossless behaviour, microsecond visibility and an understanding of the jobs on top.

Lossless by design

RDMA over Converged Ethernet expects a near-lossless fabric. PFC and ECN must be tuned together — get them wrong and you trade drops for pause storms.

Synchronized traffic

Collective operations make thousands of GPUs talk at once. Incast and microbursts appear in microseconds and vanish before a poll interval ends.

Tail latency is job latency

A training step waits for the slowest flow. One congested port or flapping NIC can slow an entire job without any alarm firing.

Scale and symmetry

Rail-optimized fabrics have thousands of identical high-speed links. Problems hide in small asymmetries — one ECMP member, one rail, one optic.

Congestion control

ECN, PFC and DCQCN — understood as one loop.

Lossless Ethernet relies on several mechanisms interacting correctly. Netlume models the whole loop, so it can tell whether a slowdown comes from load, thresholds, a misbehaving NIC or a physical fault.

RoCEv2 congestion control at a switch egress queueTraffic from a GPU sender fills the egress queue of a leaf switch. Between Kmin and Kmax packets are ECN marked; the receiver returns congestion notification packets to the sender, which reduces its rate (DCQCN). Above the Xoff threshold the switch sends PFC pause frames upstream.GPU senderRoCEv2 NIC · DCQCNleaf-07 · Et12 · TC3egress queue (lossless class)KminKmaxXoffECN marking probability ↑queue depth → shared bufferGPU receiverRoCEv2 NICdataECN-CE markedCNP → sender reduces rate (DCQCN)PFC pause (priority 3) above Xoff
Netlume watches every stage of this loop — queue depth, ECN marks, CNPs, PFC pause frames and sender rate — and correlates them with topology and jobs to tell load-driven congestion apart from misconfiguration or faults.

Coverage

What Netlume understands in an AI fabric.

GPU cluster networking

Back-end fabrics that connect GPU servers for collective communication, typically separate from front-end and storage networks.

Watches · Rail mapping, NIC-to-leaf adjacency, per-rail utilization, job placement.

RoCE / RoCEv2

RDMA over Converged Ethernet. RoCEv2 runs over UDP/IP (port 4791), so it is routable across leaf-spine fabrics.

Watches · Lossless traffic class, DSCP/priority mapping, retransmits, out-of-sequence events.

InfiniBand

A purpose-built, credit-based lossless interconnect widely used for HPC and AI clusters.

Watches · Port state, link errors, congestion counters and topology from fabric management.

Ethernet AI fabrics

High-radix 400G/800G leaf-spine Ethernet built for AI, using RoCEv2, ECN and PFC with careful buffer tuning.

Watches · ECMP balance, buffer profiles, queue occupancy, link and optic health.

ECN

Explicit Congestion Notification. Switches mark packets as queues grow instead of dropping them, signalling senders to slow down.

Watches · Marking rate per port and class against Kmin/Kmax thresholds.

PFC

Priority Flow Control (802.1Qbb). Pauses a single priority on a link to prevent drops in the lossless class.

Watches · Pause frames and duration, pause propagation, storm and deadlock patterns.

DCQCN

Data Center Quantized Congestion Notification. The rate-control loop combining ECN marks with CNPs returned to RoCEv2 senders.

Watches · CNP rates, sender rate reductions and recovery behaviour.

Congestion & microbursts

Short-lived queue build-ups from synchronized flows, often invisible in averaged counters.

Watches · High-resolution queue telemetry, buffer watermarks, drop counters.

Fabric topology

Leaf-spine and rail-optimized designs where GPU n of every server connects to rail leaf n.

Watches · Cabling against intended design, missing or asymmetric links, ECMP group membership.

Latency

Collective completion time is bounded by the slowest path across the fabric.

Watches · Per-path latency, queueing delay, job step time correlation.

Packet loss

Even small loss in the lossless class triggers retransmits that stall collectives.

Watches · Drops per class, FCS/CRC errors, optics, NIC retransmit counters.

AI workload impact

The link between a network symptom and the job, tenant or model run it slows down.

Watches · Job-to-host mapping, step time, collective duration alongside fabric telemetry.

Investigation

From a slower training step to a specific queue.

When step time rises, Netlume works down from the job to the hosts, rails, leaves and queues involved — and shows which explanations it ruled out along the way.

Correlated telemetry · 20:15 – 20:45 UTC

4 sources
ECMP members · SPECTRUM-01 → leaf tierECN-marked % · leaf-07 Et12 TC3PFC pause ms/s · leaf-07 Et12Probe loss % · GPU-042 ↔ GPU-04120:27 BGP path change20:31 loss detected

Hypotheses · INC-4127

4 tested
  1. H1ECMP redistribution → TC3 congestion on leaf-07 Et12

    92%
    supported · 7 evidence
  2. H2Degraded optic or cabling on leaf-07 Et12

    4%
    ruled out · 2 evidence
  3. H3Host NIC firmware regression on GPU-WORKER-042

    2%
    ruled out · 2 evidence
  4. H4PFC storm / pause deadlock

    2%
    ruled out · 3 evidence

Root cause candidate

evidence-linked

East-west congestion following ECMP path redistribution resulted in queue saturation on leaf-07.

Confidence

92%

Ruled out

  • Optic / physical layer fault (Rx power nominal)
  • Host NIC firmware regression
  • PFC storm / deadlock

Illustrative product UI with simulated demo data. Not real customer data or statistics.

For AI infrastructure teams

Questions you can ask about your fabric.

Netlume is designed for heterogeneous AI infrastructure: Ethernet and InfiniBand fabrics, multi-vendor switching, SmartNICs and DPUs, and the Linux hosts at the edge of it all.

  • ›Why did step time for run-0412 increase 23% at 20:31?
  • ›Which rails are seeing PFC pauses above baseline right now?
  • ›Is this all-reduce slowdown caused by the network or the hosts?
  • ›Which GPU NICs show rising retransmits or link flaps this week?
  • ›Did the last ECMP change unbalance traffic across spines?
  • ›Are buffer profiles consistent on every leaf in AI Fabric A?

Make your AI fabric explainable.

Walk through a GPU-fabric investigation with our engineers and discuss how Netlume would connect to your environment.