Network Detection and Response: Where NDR Gets Its Traffic

Six ways NDR acquires traffic, what each method misses, and what AWS, Azure and Google document about their own mirroring limits.

The Port0 owl on a large floating island collecting glowing lines from three snowy islands while two faded islands stay unconnected

Network detection and response products get judged on detection quality. The ceiling on that quality is set earlier, by how traffic reaches the sensor. No detector finds lateral movement in a segment nobody copied.

MITRE ATT&CK defines network traffic as data transmitted across a network, either summarized as flow records or captured as raw packets. NDR products read that data and decide what it means. The analysis is the product. The acquisition is the project.

This post covers the six ways NDR gets traffic, and what each method sees and misses. It then covers what AWS, Azure and Google document about their own mirroring services. It ends with the costs that appear after the proof of concept.

What the acquisition layer decides

Three properties of a deployment are fixed before any detection logic runs.

Coverage. Which segments produce data at all. A segment with no TAP, no mirror session and no logging device on the path is dark. Detection content cannot recover it.

Fidelity. What arrives. Full packets carry payload. Metadata carries protocol detail without payload. Flow records carry counters and tuples. Each step down removes a class of detection.

Cost. Sensor hardware, switch ports, broker licenses, cloud egress and storage. This is the number that grows after year one.

A vendor demo runs on a segment chosen because it was easy to mirror. Production rarely looks like that segment. Ask which of the three properties the demo was optimizing.

Six ways NDR gets traffic

Inline network TAP

A physical device sits in the link and copies every frame to a monitor port. Passive optical TAPs split light, so the copy path needs no power and no forwarding decision. Fidelity is the highest available.

The cost is physical. You need a link to break into, a maintenance window, rack space and a monitor port for every TAP. The method does not exist in public cloud.

SPAN and port mirroring

SPAN is a switch feature. It copies traffic from chosen ports or VLANs to a destination port. No new hardware, no downtime, a config change on gear you already own. That is why it dominates proofs of concept.

The limit is arithmetic. A destination port has a fixed line rate. Twenty-four gigabit access ports at half load produce more copies than one gigabit destination port can carry. The excess is dropped, and the sensor never learns which packets it lost. Garland Technology describes SPAN ports as commonly known to drop packets when a port is oversubscribed. Garland sells TAPs, so read the framing with that in mind. The oversubscription arithmetic holds either way.

Network packet broker

A broker aggregates many TAP and SPAN feeds, filters them, removes duplicates and load balances the result across tool ports. It turns an unmanageable number of feeds into a fixed number of tool inputs.

It also adds a vendor, a license and a device that can drop traffic of its own. Brokers stay relevant in cloud deployments. Microsoft's own virtual network TAP documentation lists Gigamon and Keysight as packet broker partners for Azure traffic.

Flow records

NetFlow, IPFIX and sFlow come from routers and switches you already run. Records carry addresses, ports, protocol, byte and packet counts, and timestamps. No payload.

Flow answers reachability and volume questions well. It answers protocol questions badly. sFlow samples, so a short port scan can fall entirely between samples. Treat flow as a coverage floor, not a detection source for anything that depends on protocol content.

Cloud traffic mirroring

Each major cloud provider offers a service that copies VM traffic to a collector you run. The three implementations differ enough that a single architecture rarely ports across them. Details are in the next section.

Telemetry from devices already on the path

Firewalls, proxies, DNS resolvers, load balancers and VPC flow logs all describe traffic they already handled. Nothing new gets deployed, and coverage reaches places a sensor never would, including networks you do not control.

The tradeoff is format and depth. Every device describes sessions its own way, and none of them give you packets. Coverage follows your device footprint rather than your segment map.

Reading those sources where they already sit is a federated query problem across existing tools. The normalization work moves from a sensor into the query layer.

What the three clouds actually document

Cloud mirroring is sold as the easy answer for east-west visibility. The provider documentation is more specific than the marketing.

AWS VPC Traffic Mirroring

Mirrored traffic is encapsulated in VXLAN on UDP port 4789. The target security group must allow that traffic, and the collector has to parse VXLAN before it sees a packet.

Two details change what your detections can do:

  • AWS states that inbound traffic dropped at the mirror source by security group or network ACL rules "is not mirrored". Traffic your controls already blocked never reaches the NDR. External scan and brute-force detection has to come from flow logs or the control itself.
  • With a Network Load Balancer target, mirroring requires a UDP listener on port 4789. Remove that listener and, in AWS's wording, "Traffic Mirroring fails without an error indication". A silent failure mode belongs in your monitoring.

Packet size matters for Gateway Load Balancer endpoint targets. AWS documents a maximum MTU of 8500, with mirroring adding 54 bytes of headers for IPv4 and 74 for IPv6. That puts the largest packet deliverable without truncation at 8446 bytes and 8426 bytes. AWS also warns that load balancer targets can deliver packets out of order.

Azure virtual network TAP

Azure virtual network TAP is in public preview, not general availability, and Microsoft lists 15 supported regions. Sources must be VM network interfaces. Destinations must be a load balancer or a VM network interface.

The documented limitations shape the design:

  • VMs in a virtual network with encryption enabled cannot be mirroring sources.
  • IPv6 is not supported.
  • Virtual WAN peering between source and destination virtual networks is not supported. Direct peering is required.
  • Inbound traffic from Private Link Service is not mirrored.

The preview limitations shape the change request. Each source VM has to be stopped, deallocated and started once before it can be mirrored. Adding or removing a VM as a source can cause up to 60 seconds of network downtime. Live Migration is disabled on any VM configured as a source.

Read that last point against your availability commitments before you scope coverage.

Google Cloud Packet Mirroring

Google Cloud Packet Mirroring captures full packet data, including payloads and headers. Scope is tighter than the other two. All mirrored sources must sit in the same project, VPC network and region.

Quotas cap a single policy at 5 subnets, 5 tags and 50 instances, with a maximum of 30 mirroring filters. Large estates need many policies.

The bandwidth note is the figure to take to a budget meeting. Google's own example describes an instance with 1 Gbps of ingress and 1 Gbps of egress. Mirroring adds 2 Gbps of egress, for 3 Gbps total on that VM. Mirroring triples the egress the machine has to carry.

Google now recommends Network Security Integration Packet Mirroring for new deployments, so a design built on the older service starts from a superseded path.

Side by side

AWS Traffic MirroringAzure virtual network TAPGoogle Packet Mirroring
StatusGenerally availablePublic preview, 15 regionsSuperseded for new deployments
SourceElastic network interfaceVM network interface onlyVM instances, by subnet, tag or name
DeliveryVXLAN, UDP 4789Load balancer or VM network interfaceForwarded to an instance group
Documented scope limitInbound traffic dropped by security group or network ACL is not mirroredNo IPv6, no encrypted virtual networks, no vWAN peeringSame project, VPC network and region
Operational cost in the docsSilent failure if the UDP 4789 listener is removedUp to 60 seconds downtime when adding or removing a source VMAdds 2 Gbps egress to a 1 Gbps in, 1 Gbps out VM

The cost curve that appears after the proof of concept

Sensor count follows the segment map

Budgets get built per site. Coverage gets consumed per segment. Take a data center with separated production, management, storage and user VLANs. That site needs a TAP and sensor per segment, or a broker to aggregate them. East-west coverage is where the sensor count climbs.

Mirrored traffic is real traffic

Google documents mirroring tripling a VM's egress. That bandwidth is produced, carried and, where it crosses a zone, charged. Model mirroring volume as a share of production volume before you commit to a mirroring-based design. Then check that figure against the segments you actually need.

Retention outlives the sensor

Packet capture at any useful retention is the largest line in many NDR deployments. Metadata and flow shrink it by orders of magnitude and remove detections at the same time. Decide the retention window per zone, then pick the fidelity that fits it.

Questions to put to an NDR vendor

  1. Which acquisition methods do your sensors support, and which method did this demo use?
  2. What happens to detection coverage when the input is flow records rather than packets? Name the detections that stop working.
  3. In AWS, how do you detect activity that security groups already block, given that it is never mirrored?
  4. Is your Azure support built on virtual network TAP? If so, how do you handle the preview limitations and the source VM downtime?
  5. What is the expected mirrored volume as a share of production traffic for the segments in scope?
  6. How do you detect and alert on sensor input loss, including a dropped SPAN session or an oversubscribed destination port?
  7. Which segments in our current architecture stay dark after deployment?

Question seven carries the most weight. A vendor who answers it honestly is describing the project you are actually buying.

Matching the method to the zone

No single acquisition method covers a modern estate. Practical deployments mix them, and the mix follows what each zone can support.

  • Data center core and crown jewel segments. Inline TAP into a broker. The fidelity justifies the hardware.
  • Distributed campus and branch. SPAN where oversubscription is measured, flow records where it is not. Instrument the aggregation layer rather than every access switch.
  • Cloud workloads. Provider mirroring for the segments that justify the egress cost. Flow logs everywhere else.
  • Remote users, SaaS, OT and unmanaged devices. Telemetry from devices on the path. Firewall sessions, proxy logs and DNS resolver records. A sensor was never going to sit there.

We built our network detection to run in either mode. Use a lightweight sensor in zones that can support it. Elsewhere, run network detection on the firewall, proxy and cloud flow log signal your environment already produces. The tradeoffs against an appliance-based product are set out in our capability-by-capability comparison with traditional NDR.

Map acquisition per zone before you compare detection catalogues. A detection catalogue you cannot feed is a list of things that will not fire.

See Port0 on your own data.

Bring your noisiest alert queue. Watch Soc0 investigate it live.

Book a Demo

Never Miss an Insight

Subscribe to get the latest posts delivered to your inbox.