Linux 6.6 UDP GRO 收包调用链源码分析

1. 概述

GRO(Generic Receive Offload,通用接收卸载)是 Linux 内核网络栈中的一项关键性能优化技术。
其核心思想是在数据包进入协议栈入口处,
将属于同一条流/五元组的多个小包合并为一个大的“超级包”(super-skb),
从而显著减少上层协议栈的处理次数,提升网络吞吐量。

GRO 的处理是分层进行的:
在网口驱动的napi处理函数里,依次经过链路层 → IP 层 → 传输层(TCP/UDP)的逐层 GRO 回调,
最终决定数据包是合并入已有流、被消费还是作为普通包送往协议栈。

本文以 Linux 6.6 内核源码为基础,分析 UDP 数据包在 GRO 处理路径上的完整调用链。

1.1. 函数调用链

完整的 UDP GRO 调用链如下:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
napi_gro_receive()
├─ dev_gro_receive()
│ ├─ 计算 hash bucket,定位 gro_list
│ ├─ 遍历 offload_base,匹配 skb->protocol
│ └─ inet_gro_receive()
│ ├─ 检查 IP 层信息(源目 IP、TTL、TOS、分片)
│ ├─ 查找 inet_offloads[proto]
│ └─ udp4_gro_receive()
│ ├─ udp4_gro_lookup_skb() 查找 socket
│ └─ udp_gro_receive()
│ ├─ 端口匹配
│ ├─ 长度校验
│ └─ skb_gro_receive() / skb_gro_receive_list()
└─ napi_skb_finish()
├─ GRO_MERGED / GRO_MERGED_FREE → 合并完成,释放 skb
├─ GRO_CONSUMED → 已被消费
└─ GRO_NORMAL → 走普通协议栈路径
1.1.0.0.1. 分层解析

UDP GRO 的调用链在内核中经历了一条清晰的分层路径

  • 入口: napi_gro_receive()

    • 由网卡驱动调用,计算数据包哈希值
    • 通过 skb_get_hash_raw() & (GRO_HASH_BUCKETS - 1) 定位到对应的哈希桶 gro_hash[bucket]下的链表入口,
    • 然后调用 dev_gro_receive() 进行协议层分发
  • 协议分发:dev_gro_receive() -

    • 根据 skb->protocol 找到匹配的协议。 遍历 offload_base 链表,找到IP协议的注册的回调函数inet_gro_receive,
    • 通过 INDIRECT_CALL_INET 宏进行间接调用;
  • IP 层:inet_gro_receive()

    • 检查 IP 头部(源目 IP、TTL、TOS、分片信息等),
    • 在哈希桶内的 gro_list 中查找可合并的流。
    • 如果 TTL 或 TOS 不一致、是分片包等场景,则启动 flush 操作;
  • 传输层:udp4_gro_receive()

    • 查找对应的 socket(用于检查 udp_sk(sk)->gro_enabled),这sk在udp tunnel类的场景(如vxlan)下有用,其他场景可忽略。
    • 然后由 udp_gro_receive() 执行端口匹配和长度校验,
    • 最终通过 skb_gro_receive() 完成数据合并;
  • 完成层:napi_skb_finish()

    • 根据 dev_gro_receive() 的返回值决定数据包的最终去向
    • ——合并:有两种方式,详见后续专门讲述。
      • GRO_MERGED:把当前udp报文合并到之前的udp报文里。简单,效率低。
      • GRO_MERGED_FREE:保留报文大小。复杂,高效。
    • 消费(GRO_CONSUMED)
    • 或走普通路径(GRO_NORMAL)。

2. 数据结构

2.1. 核心: struct offload_callbacks

内核里的GRO是分层处理的, 普通报文包含两层

  • 网络层:IPv4、IPv6。
  • 传输层:TCP、UDP。
    尽管层级不一样,协议不一样,但是他们的GRO处理都有个共同的抽象结构体offload_callbacks
1
2
3
4
5
6
7
2704 struct offload_callbacks {
2705 struct sk_buff *(*gso_segment)(struct sk_buff *skb,
2706 netdev_features_t features);
2707 struct sk_buff *(*gro_receive)(struct list_head *head,
2708 struct sk_buff *skb);
2709 int (*gro_complete)(struct sk_buff *skb, int nhoff);
2710 };
  • gso_segment:GRO报文穿过协议栈,到达socket时候,如何还原切分?
  • gro_receive:RX软中断在协议出口,尝试组装GRO报文时,判断报文是否可以GRO。
  • gro_complete:如果GRO解释,。。。

2.2. IP层

Linux 内核 inet_offloads 数组结构

  • 每个IP层的offload/GRO操作都汇总到packet_offload结构体中。
    • type: 对应eth头里的proto字段,如IPv4,IPv6.
    • struct offload_callbacks callbacks: 协议对应的GRO处理函数集。
  • 所有的结构体通过一个list 挂载到offload_base链表下。

2.2.1. offload_base

1
13 struct list_head offload_base __read_mostly = LIST_HEAD_INIT(offload_base);

2.2.2. packet_offload

1
2
3
4
5
6
2712 struct packet_offload {
2713 __be16 type; /* This is really htons(ether_type). */
2714 u16 priority;
2715 struct offload_callbacks callbacks;
2716 struct list_head list;
2717 };

2.2.2.1. IPv4: ip_packet_offload

1
2
3
4
5
6
7
8
9
1902 static struct packet_offload ip_packet_offload __read_mostly = {
1903 .type = cpu_to_be16(ETH_P_IP),
1904 .callbacks = {
1905 .gso_segment = inet_gso_segment,
1906 .gro_receive = inet_gro_receive,
1907 .gro_complete = inet_gro_complete,
1908 },
1909 };
1910

2.2.2.2. IPv6: ipv6_packet_offload

1
2
3
4
5
6
7
8
394 static struct packet_offload ipv6_packet_offload __read_mostly = {
395 .type = cpu_to_be16(ETH_P_IPV6),
396 .callbacks = {
397 .gso_segment = ipv6_gso_segment,
398 .gro_receive = ipv6_gro_receive,
399 .gro_complete = ipv6_gro_complete,
400 },
401 };

2.3. TCP/UDP层

  • `` 与Ip层处理类似,也有一个传输层的结构体存放TCP/UDP的GRO操作。
  • inet_offloads

Linux 内核 inet_offloads 数组结构

2.3.1. inet_offloads

1
29 const struct net_offload __rcu *inet_offloads[MAX_INET_PROTOS] __read_mostly;

2.3.2. net_offload

1
2
3
4
68 struct net_offload {
69 struct offload_callbacks callbacks;
70 unsigned int flags; /* Flags used by IPv6 for now */
71 };

2.3.3. TCP4: tcpv4_offload

1
2
3
4
5
6
7
347 static const struct net_offload tcpv4_offload = {
348 .callbacks = {
349 .gso_segment = tcp4_gso_segment,
350 .gro_receive = tcp4_gro_receive,
351 .gro_complete = tcp4_gro_complete,
352 },
353 };

2.3.4. UDP4: udpv4_offload

1
2
3
4
5
6
7
740 static const struct net_offload udpv4_offload = {
741 .callbacks = {
742 .gso_segment = udp4_ufo_fragment,
743 .gro_receive = udp4_gro_receive,
744 .gro_complete = udp4_gro_complete,
745 },
746 };

3. 内核函数

3.1. 入口函数 napi_gro_receive

1
2
3
4
5
6
7
8
600 gro_result_t napi_gro_receive(struct napi_struct *napi, struct sk_buff *skb)
601 {
...
609 ret = napi_skb_finish(napi, skb, dev_gro_receive(napi, skb));
611
612 return ret;
613 }
614 EXPORT_SYMBOL(napi_gro_receive);

napi_gro_receive() 是 GRO 处理的统一入口,由网卡驱动在 NAPI poll 函数中调用(驱动需设置 NETIF_F_GRO 特性)。
其核心逻辑在dev_gro_receive()里。
napi_skb_finish() 根据返回值完成最终处理。

3.2. 核心分发函数 dev_gro_receive

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
437 static enum gro_result dev_gro_receive(struct napi_struct *napi, struct sk_buff *skb)
438 {
439 u32 bucket = skb_get_hash_raw(skb) & (GRO_HASH_BUCKETS - 1);
440 struct gro_list *gro_list = &napi->gro_hash[bucket];
441 struct list_head *head = &offload_base;
442 struct packet_offload *ptype;
443 __be16 type = skb->protocol;
...
454 list_for_each_entry_rcu(ptype, head, list) {
455 if (ptype->type == type && ptype->callbacks.gro_receive)
456 goto found_ptype;
457 }
...
461 found_ptype:
462 skb_set_network_header(skb, skb_gro_offset(skb));
...
490 pp = INDIRECT_CALL_INET(ptype->callbacks.gro_receive,
491 ipv6_gro_receive, inet_gro_receive,
492 &gro_list->list, skb);
...
501 same_flow = NAPI_GRO_CB(skb)->same_flow;
502 ret = NAPI_GRO_CB(skb)->free ? GRO_MERGED_FREE : GRO_MERGED;
503
504 if (pp) {
505 skb_list_del_init(pp);
506 napi_gro_complete(napi, pp);
507 gro_list->count--;
508 }
...
537 return ret;
543 }

dev_gro_receive() 是 GRO 处理的核心分发函数,其关键步骤如下:

计算哈希链入口:skb_get_hash_raw() 获取数据包的哈希值(由网卡RSS或软件计算),找到hash桶里的链表入口。

查找网络层处理协议:遍历全局 offload_base 链表,根据 skb->protocol 匹配对应的 packet_offload 类型。

  • IPv4 对应 inet_gro_receive。
  • IPv6 对应 ipv6_gro_receive。

调用协议处理GRO:通过 INDIRECT_CALL_INET 宏调用协议层的 GRO 处理函数。
处理协议返回结果:处理协议层返回 pp。

  • 如果pp 非空:有一个流的GRO缓存内容,需要被flush。调用 napi_gro_complete() 将其送往协议栈,并减少桶计数。

3.3. IP 层处理 inet_gro_receive

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
1471 struct sk_buff *inet_gro_receive(struct list_head *head, struct sk_buff *skb)
1472 {
1473 const struct net_offload *ops;
...
1489 proto = iph->protocol;
1490
1491 ops = rcu_dereference(inet_offloads[proto]);
1492 if (!ops || !ops->callbacks.gro_receive)
...
1509 list_for_each_entry(p, head, list) {
// 检查是否是同一个flow, IP层,源目IP
// ttl ,tos,不一致,分片等场景下启动flush操作。
1562 }
...
1577 pp = indirect_call_gro_receive(tcp4_gro_receive, udp4_gro_receive,
1578 ops->callbacks.gro_receive, head, skb);
1579
1580 out:
1581 skb_gro_flush_final(skb, pp, flush);
1582
1583 return pp;
1584 }

inet_gro_receive() 负责 IP 层的流匹配和 flush 判断:

  • 确定传输层协议的GRO:根据iph->protocol 从 inet_offloads[] 数组查找对应的传输层协议回调
    • TCP 对应 tcp4_gro_receive。
    • UDP 对应 udp4_gro_receive。
  • 流匹配:遍历当前哈希桶内的 skb 列表,
    • 检查 IP 层信息是否一致——源目 IP 地址、TTL、TOS、是否分片等。
    • 任何一项不匹配都会触发 flush 标志,停止当前流的GSO缓存并flush。
  • 调用传输层: 调用具体的传输层 GRO 函数。

3.4. UDP层处理入口:udp4_gro_receive

1
2
3
4
5
6
7
8
9
10
11
12
621 INDIRECT_CALLABLE_SCOPE
622 struct sk_buff *udp4_gro_receive(struct list_head *head, struct sk_buff *skb)
623 {
624 struct udphdr *uh = udp_gro_udphdr(skb);
...
644 if (static_branch_unlikely(&udp_encap_needed_key))
645 sk = udp4_gro_lookup_skb(skb, uh->source, uh->dest);
646
647 pp = udp_gro_receive(head, skb, uh, sk);
648 return pp;
...
653 }

udp4_gro_receive() 是 UDP 协议的 GRO 入口,无实质意义,核心函数在udp_gro_receive。
注:

  • Socket 查找(L644-L645):在 udp_encap_needed_key 静态分支开启时,调用 udp4_gro_lookup_skb() 根据源/目的端口查找对应的 UDP socket。这个 socket 用于检查 udp_sk(sk)->gro_enabled 是否开启,以决定是否允许 GRO 合并。

udp_gro_receive() 内部执行端口匹配(比较源/目的端口是否一致)、长度校验(防止 GRO 合并过大),以及决定使用普通合并(skb_gro_receive)还是 fraglist 合并(skb_gro_receive_list)。

3.4.1. UDP层处理函数:udp_gro_receive

1
2
3
4
5
6
7
8
9
10
545 struct sk_buff *udp_gro_receive(struct list_head *head, struct sk_buff *skb,
546 struct udphdr *uh, struct sock *sk)
547 {
...
562 if ((!sk && (skb->dev->features & NETIF_F_GRO_UDP_FWD)) ||
563 (sk && udp_sk(sk)->gro_enabled) || NAPI_GRO_CB(skb)->is_flist)
564 return call_gro_receive(udp_gro_receive_segment, head, skb);
...
604 }
605 EXPORT_SYMBOL(udp_gro_receive);

3.4.2. UDP层核心处理函数:udp_gro_receive_segment

  • 比对源目port,确保在一个五元组后缓存。 continue是考虑有多个流的场景。
    • TODO:支持多流会不会引入乱序等问题。
  • list 和非list模式不混用。
  • 报文大小要一致, 方便以后切割。 list模式可以接受报文不一样大小。
  • checksum不一致,也会触发flush。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
464 static struct sk_buff *udp_gro_receive_segment(struct list_head *head,
465 struct sk_buff *skb)
466 {
467 struct udphdr *uh = udp_gro_udphdr(skb);
468 struct sk_buff *pp = NULL;
...
489 list_for_each_entry(p, head, list) {
... // 确保是同一个udp流,这里用u32检查源目端口。
495 /* Match ports only, as csum is always non zero */
496 if ((*(u32 *)&uh->source != *(u32 *)&uh2->source)) {
497 NAPI_GRO_CB(p)->same_flow = 0;
498 continue;
499 }
... //下面是核心代码。
506 /* Terminate the flow on len mismatch or if it grow "too much".
507 * Under small packet flood GRO count could elsewhere grow a lot
508 * leading to excessive truesize values.
509 * On len mismatch merge the first packet shorter than gso_size,
510 * otherwise complete the GRO packet.
511 */
512 if (ulen > ntohs(uh2->len)) { // 如果当前报文比以前的大, 不缓存并启动flush。
513 pp = p;
514 } else { // list模式和非list模式后面专题解释。
515 if (NAPI_GRO_CB(skb)->is_flist) {
516 if (!pskb_may_pull(skb, skb_gro_offset(skb))) {
517 NAPI_GRO_CB(skb)->flush = 1;
518 return NULL;
519 }
520 if ((skb->ip_summed != p->ip_summed) ||
521 (skb->csum_level != p->csum_level)) {
522 NAPI_GRO_CB(skb)->flush = 1;
523 return NULL;
524 }
525 ret = skb_gro_receive_list(p, skb);
526 } else {
527 skb_gro_postpull_rcsum(skb, uh,
528 sizeof(struct udphdr));
529
530 ret = skb_gro_receive(p, skb);
531 }
532 }

4. NAPI相关部分

4.1. 数据结构

4.1.1. napi_struct里GRO相关字段)

1
2
3
4
5
6
352 struct napi_struct {
...
373 struct gro_list gro_hash[GRO_HASH_BUCKETS];
374 struct sk_buff *skb;
...
383 };

4.1.2. GRO_HASH_BUCKETS

1
347 #define GRO_HASH_BUCKETS        8

4.1.3. gro_list

1
2
3
4
338 struct gro_list {
339 struct list_head list;
340 int count;
341 };
  • list:list_head 链表头,用于挂载属于同一哈希桶的所有流(每个流对应一个正在 GRO 合并的 skb)。
  • count:当前桶内挂载的 skb 数量,用于控制流数量上限(MAX_GRO_SKBS)。
  • gro_hash[8]:8 个 GRO 哈希桶组成的数组。
    • 每个桶独立维护自己的流列表。
  • skb:当前正在处理的 skb 指针,用于 NAPI poll 上下文。

在 Linux 5.x 之后的内核中,gro_hash 数组的引入使得 GRO 在多个并行流同时活跃时,每个数据包只需在其哈希值对应的桶中查找,而无需遍历整个 per-NAPI 流列表,显著提升了多流场景的性能。

5. skb_gro_receive与skb_gro_receive_list的区别

5.1. 核心区别:融合 (Frags) vs. 串联 (Fraglist)

特性 skb_gro_receive() skb_gro_receive_list()
合并方式 数据融合(拷贝进 frags[]) 数据包串联(挂到 frag_list)
数据拷贝 需要(或内存重映射) 无需
大小限制 受 MAX_SKB_FRAGS 限制 受 64KB 总长度限制
主要场景 普通接收路径 转发、支持 GRO_FRAGLIST 的设备
性能开销 较高 较低(尤其对转发)

因此,代码中根据 is_flist 标志来选择调用哪个函数,本质上是在数据拷贝开销、内存布局限制和后续处理效率之间做出的权衡。

5.2. skb_gro_receive()

skb_gro_receive() 试图将新 skb 的载荷数据真正“融合”进主 skb p 的存储空间中。它会尝试将数据追加到 p 已有的线性数据区或者其 frags[] 数组里。

这种方式需要处理内存分配、拷贝和 frags[] 数组的管理,开销相对较大,并且受限于 MAX_SKB_FRAGS(通常为 17 个),单个 GRO 包的大小有限。

5.3. skb_gro_receive_list()

skb_gro_receive_list() 不融合载荷数据,而是通过 frag_list 指针,将整个 skb 以链表的形式串联起来。它做的事情非常轻量:

  • 将新 skb 挂到主 skb p 的 skb_shinfo(p)->frag_list 链表尾部。
  • 更新链表指针 NAPI_GRO_CB(p)->last 指向新的尾部。
  • 直接累加 p->len、p->data_len 和 p->truesize 等元数据。

6. is_flist 标志与设计动机

代码中通过 NAPI_GRO_CB(skb)->is_flist 来区分这两种模式。这个标志通常由网络设备驱动或上层配置决定,例如当设备支持 NETIF_F_GRO_FRAGLIST 特性时,is_flist 会被设置为 1。

区分这两种方式的核心动机是性能和场景适配:

  1. 性能优化(尤其转发场景)
    fraglist 方式避免了数据拷贝,在大包转发场景下优势明显。当 GRO 聚合后的数据需要再次以 GSO 形式发送时,fraglist 结构可以直接被 skb_segment_list() 高效处理,无需重新拆分和拷贝,大幅降低了 CPU 开销。

  2. 突破 frags[] 数量限制
    传统 skb_gro_receive() 受 MAX_SKB_FRAGS(约 17 个)限制,构建的 GRO 包大小有限。fraglist 方式通过串联 skb,可以构建更大的 GRO 包(上限由 65536 字节限制),支持更高效的聚合。

  3. 适配硬件/驱动卸载
    某些网卡或驱动可能已经将数据组织成 frag_list 形式,GRO 层直接沿用这种结构可以避免不必要的重构。

7. 注意事项

fraglist 方式虽然高效,但有其复杂性。例如,它需要防止 GSO skb 被再次嵌套聚合,否则可能导致后续分段路径处理异常。因此,代码中在进入 skb_gro_receive_list() 前会检查 flush 标志等条件。


Linux 6.6 UDP GRO 收包调用链源码分析
https://martinbj2008.github.io/2026/09/29/udp-GRO-summary/
Author
Martinbj2008
Posted on
September 29, 2026
Licensed under