Why was Linux SRv6 H.Encaps.L2 still slow with tunsrc set?
Note: the text of this post was written by Claude and then checked and edited by me.
In the previous post, setting tunsrc made both H.Encaps.V6 and H.Encaps.L2 faster. Even so, H.Encaps.L2 stayed at 660.7 kpps against 976.4 kpps for H.Encaps.V6 (both PDR, with tunsrc set). For the same input packets, H.Encaps.L2 additionally carries the 14-byte inner Ethernet header, which should not cost a third of the throughput. It turned out to be an unnecessary memory allocation on every packet, and the fix is now in net-next.
Test setup
Everything below was measured on net-next, just before and at the fix, built from the same configuration (based on Ubuntu’s, with CONFIG_INIT_ON_ALLOC_DEFAULT_ON=y). The SUT (the machine under test) has a Xeon E5-2650 v3 at 2.30 GHz and forwards 64-byte frames on a single core, connected back to back to a T-Rex tester with Intel 82599ES (ixgbe) NICs, with iommu=pt and tunsrc set. Throughput is measured with my fork of SRPerf (updated to run on Ubuntu 24.04 and Python 3.12) as in the previous post: PDR (Partial Drop Rate, the highest offered rate with at most 0.5% packet loss) and MRR (Maximum Receive Rate, the receive rate when sending at line rate).
What the flame graph showed
I took a flame graph of the forwarding core while sending H.Encaps.L2 traffic at a fixed 500 kpps, which the kernel before the fix forwards without loss.
Including their callees, pskb_expand_head() takes 21.1% of the sampled cycles, kmalloc_reserve() 18.7% and memset_orig() 14.1%. Every memset_orig() sample is on the path seg6_do_srh() → pskb_expand_head() → kmalloc_reserve(), so the zeroing comes from reallocating the skb head. The code shows that this happens on every packet.
Why the skb head was reallocated on every packet
seg6_do_srh() in net/ipv6/seg6_iptunnel.c pushes the Ethernet header back in front of the packet before building the outer IPv6 header, and in the L2 encapsulation modes it called pskb_expand_head() unconditionally to make room for it. pskb_expand_head() always allocates a new head and copies the headroom and the linear data into it, even when the existing headroom is already large enough. The IPv6 encapsulation modes use skb_cow_head() instead, which reallocates only when the headroom is too small or the head is shared with a clone.
In this setup, the skb reaches seg6_do_srh() with 206 bytes of headroom, while a single-segment L2 encapsulation needs 94 bytes (14 for the Ethernet header, 40 for the IPv6 header, 24 for the SRH and 16 reserved for the outgoing link layer). The skbs were not cloned either, so the reallocation was not needed.
The cost is made larger by CONFIG_INIT_ON_ALLOC_DEFAULT_ON. This is a hardening option, added in Linux 5.3, that zero-fills page and slab allocations to reduce the risk of exposing stale data. Ubuntu and Debian enable it by default, and it was enabled in the kernels here, so every reallocated head was also cleared with memset(). It can be turned off with init_on_alloc=0, but turning off a security feature to win back performance is the wrong way around; not reallocating is the better fix.
The unconditional reallocation itself dates back to when H.Encaps.L2 was added in 2017, so the kernel used in the SRPerf paper (net-next during the 5.2 cycle) had it as well. That kernel predates init_on_alloc, which was merged in 5.3, though, and in the paper H.Encaps.L2 was much closer to H.Encaps.V6 (about 828 against 978 kpps, 0.85). To see how much of the gap came from the zeroing, I booted the kernel before the fix with init_on_alloc=0:
| init_on_alloc | H.Encaps.V6 | H.Encaps.L2 | L2/V6 |
|---|---|---|---|
| 1 (default) | 937.7 | 644.0 | 0.69 |
| 0 | 938.0 | 773.9 | 0.83 |
(MRR in kpps, mean of 10 runs, before the fix)
Without the zeroing, L2/V6 goes from 0.69 to 0.83, close to the paper.
The fix
The fix asks skb_cow_head() for the whole encapsulation up front:
- if (pskb_expand_head(skb, skb->mac_len, 0, GFP_ATOMIC) < 0)
- return -ENOMEM;
+ headroom = skb->mac_len + sizeof(struct ipv6hdr) +
+ ipv6_optlen(tinfo->srh) +
+ dst_dev_overhead(cache_dst, skb);
+
+ err = skb_cow_head(skb, headroom);
+ if (unlikely(err))
+ return err;
When the headroom is large enough and the skb is not cloned, nothing is reallocated. When the headroom really is too small, the head is reallocated once, and the later skb_cow_head() in the IPv6 encapsulation code then finds the room it needs instead of reallocating a second time.
After the fix, the allocation no longer appears in the flame graph under the same load:
pskb_expand_head(), kmalloc_reserve() and memset_orig() have no samples left, and seg6_do_srh() falls from 28.2% to 5.6% of the cycles. Dividing the cycles of each 30-second capture by the 15 million packets forwarded in it, the core spends 2602 cycles per packet instead of 3862.
Measurement
| H.Encaps.L2 PDR | |
|---|---|
| before the fix | 654.6 kpps |
| after the fix | 965.7 kpps (+47.5%) |
(mean of 10 PDR searches, each using 10-second trials)
H.Encaps.V6 does not take this code path, and its MRR stays at about 938 kpps on both kernels. With init_on_alloc=0, the kernel with the fix gives an L2/V6 ratio of 1.02 (956.8 against 939.1 kpps MRR): H.Encaps.L2 is as fast as H.Encaps.V6, so the gap that was left after removing the zeroing was the reallocation.
Getting it upstream
The first version only asked skb_cow_head() for skb->mac_len. Eric Dumazet pointed out that a header-cloned skb could then still be reallocated twice, once to unclone it and again for the outer header, so the second version asked for the whole encapsulation as above.
While reviewing the second version, Sashiko, the AI review bot on netdev, found a different bug. dst_dev_overhead() reserves room for the egress device’s link-layer header, but the code rebuilds the ingress MAC header in front of the packet. When the ingress MAC header is longer, for example a VLAN-tagged frame received with reorder_hdr off, the offset of the rebuilt header wraps around and the copy lands about 64 KB past the buffer. This was not specific to the L2 modes, so I fixed it separately in dst_dev_overhead() and sent it to the net tree. That fix was merged first, and the third version of this patch went into net-next, which currently tracks 7.3-rc, so it should be in Linux 7.4.
References
- seg6: reallocate the skb head on L2 encapsulation only when needed (net-next, bf515c9fae87)
- net: ipv6: keep room for the mac header in dst_dev_overhead() (net, 87cd6b717e40)
- Patch threads: v1, v3
- A. Abdelsalam, P. L. Ventre, C. Scarpitta, A. Mayer, S. Salsano, P. Camarillo, F. Clad, C. Filsfils, “SRPerf: A Performance Evaluation Framework for IPv6 Segment Routing”, IEEE Transactions on Network and Service Management, vol. 18, no. 2, pp. 2320-2333, 2021
- The kernel’s command-line parameters: init_on_alloc
- SRPerf (GitHub) and my fork
- How much slower is Linux SRv6 H.Encaps without tunsrc?
- Brendan Gregg’s FlameGraph