Amazon ENA Deep Dive: Safe Upgrades, Linearization, and DMA Mapping
Amazon ENA Deep Dive: Safe Upgrades, Linearization, and DMA Mapping
ENA incidents are often reduced to “upgrade the driver and reboot.” Packet loss, Tx failures, or a missing interface after hot replacement require a more precise model of both the PCI device lifecycle and the Linux transmit path.
This guide combines two related topics: upgrading ENA safely and interpreting linearize_failed, dma_mapping_err, and related counters.
Device lifecycle and upgrade risk
After module load, the PCI core matches the ENA device and calls ena_probe(). The driver validates versions, negotiates features, builds queues and interrupts, and registers the netdev.
During removal, ena_remove() reverses that process. rmmod ena therefore dismantles the active network driver; it does not merely replace a file. SSH and SSM may disappear with the primary NIC, so hot replacement requires EC2 Serial Console or another tested recovery channel.
Record a baseline first:
uname -r
modinfo ena | grep -E '^(version|srcversion|vermagic|filename):'
ethtool -i eth0
ethtool -S eth0 > /var/tmp/ena-stats.before
ip -details link show eth0
dmesg -T | grep -i ena | tail -100
test -e "/lib/modules/$(uname -r)/build" || echo 'missing kernel headers'Unknown symbol and Invalid module format usually indicate a module/kernel ABI mismatch rather than broken ENA hardware.
Preferred upgrade order
From lowest to highest operational risk:
- Upgrade the distribution kernel and its in-tree ENA driver.
- Use DKMS to build the module for every target kernel and reboot in a window.
- Use
rmmod/modprobehot replacement only with an independent recovery path.
After the change:
modinfo ena | grep -E '^(version|filename|vermagic):'
ethtool -i eth0
ip link show eth0
ethtool -S eth0 | grep -iE 'reset|timeout|dma|linear|drop|error'
dmesg -T | grep -i ena | tail -100skb linearization
A transmit skb can contain a linear head and multiple fragments. ENA has a finite descriptor budget. If the fragment count exceeds the supported limit, the driver attempts skb_linearize() to combine data into a contiguous buffer.
That operation allocates memory and copies data. Failure becomes more likely under memory pressure, fragmentation, unusual skb layouts, or changed GSO/TSO behavior.
ethtool -S eth0 | grep -i linear
cat /proc/buddyinfo
vmstat 1
slabtop -o
ethtool -k eth0 | grep -E 'scatter-gather|tcp-segmentation|generic-segmentation'A high linearization count without failures is not automatically a fault. A continuously increasing linearize_failed counter indicates that the memory consolidation itself is failing and packets may be dropped.
Do not disable offloads merely to reduce a counter. Run controlled A/B traffic tests; otherwise increased CPU cost may create a different bottleneck.
DMA mapping failures
After the skb fits the descriptor limits, every buffer must be mapped to a DMA address accessible by the device. An error from dma_map_single() or dma_map_page() increments dma_mapping_err, and the packet cannot be submitted to the hardware queue.
ethtool -S eth0 | grep -i dma
dmesg -T | grep -iE 'dma|swiotlb|iommu|ena'
grep -i swiotlb /proc/meminfo 2>/dev/null || truePotential causes include IOMMU/SWIOTLB exhaustion, kernel defects, and extreme memory pressure. This is not ordinary queue congestion, and simply increasing queue length is unlikely to repair it.
Map counters back to stages
| Counter or symptom | Stage | First checks |
|---|---|---|
| High linearize count | skb fragment consolidation | Offloads, encapsulation, fragment count |
Increasing linearize_failed | Memory consolidation | Memory pressure, fragmentation, slab |
Increasing dma_mapping_err | DMA address mapping | IOMMU, SWIOTLB, kernel logs |
| Tx preparation errors | Descriptor/device command preparation | Queue state, driver logs |
| Reset/keep-alive growth | Device health and recovery | dmesg, host events, driver version |
Measure rates rather than lifetime totals:
ethtool -S eth0 > /var/tmp/ena-stats.1
sleep 60
ethtool -S eth0 > /var/tmp/ena-stats.2
diff -u /var/tmp/ena-stats.1 /var/tmp/ena-stats.2Upgrade acceptance
A successful upgrade means more than a new modinfo version. Confirm that the driver loads after reboot, addresses/routes/DNS return, SSH or SSM reconnects, workload throughput returns to baseline, error counters stop growing abnormally, and DKMS covers current and planned kernels.
If modprobe ena fails after hot replacement, preserve serial output and inspect vermagic, missing symbols, and initramfs before repeatedly unloading and loading modules.
Summary
An ENA upgrade is a network-device lifecycle change and must be designed for temporary loss of connectivity. Tx troubleshooting should follow the data path: fragment limits and linearization, memory consolidation, and finally DMA mapping. Mapping each counter back to its source stage prevents every send failure from being misdiagnosed as “an old driver.”
