1. 概述
GRO(Generic Receive Offload,通用接收卸载)是 Linux 内核网络栈中的一项关键性能优化技术。
其核心思想是在数据包进入协议栈入口处,
将属于同一条流/五元组的多个小包合并为一个大的“超级包”(super-skb),
从而显著减少上层协议栈的处理次数,提升网络吞吐量。
GRO 的处理是分层进行的:
在网口驱动的napi处理函数里,依次经过链路层 → IP 层 → 传输层(TCP/UDP)的逐层 GRO 回调,
最终决定数据包是合并入已有流、被消费还是作为普通包送往协议栈。
本文以 Linux 6.6 内核源码为基础,分析 UDP 数据包在 GRO 处理路径上的完整调用链。
1.1. 函数调用链
完整的 UDP GRO 调用链如下:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
| napi_gro_receive() ├─ dev_gro_receive() │ ├─ 计算 hash bucket,定位 gro_list │ ├─ 遍历 offload_base,匹配 skb->protocol │ └─ inet_gro_receive() │ ├─ 检查 IP 层信息(源目 IP、TTL、TOS、分片) │ ├─ 查找 inet_offloads[proto] │ └─ udp4_gro_receive() │ ├─ udp4_gro_lookup_skb() 查找 socket │ └─ udp_gro_receive() │ ├─ 端口匹配 │ ├─ 长度校验 │ └─ skb_gro_receive() / skb_gro_receive_list() └─ napi_skb_finish() ├─ GRO_MERGED / GRO_MERGED_FREE → 合并完成,释放 skb ├─ GRO_CONSUMED → 已被消费 └─ GRO_NORMAL → 走普通协议栈路径
|
1.1.0.0.1. 分层解析
UDP GRO 的调用链在内核中经历了一条清晰的分层路径
入口: napi_gro_receive()
- 由网卡驱动调用,计算数据包哈希值
- 通过
skb_get_hash_raw() & (GRO_HASH_BUCKETS - 1) 定位到对应的哈希桶 gro_hash[bucket]下的链表入口,
- 然后调用
dev_gro_receive() 进行协议层分发
协议分发:dev_gro_receive() -
- 根据
skb->protocol 找到匹配的协议。 遍历 offload_base 链表,找到IP协议的注册的回调函数inet_gro_receive,
- 通过
INDIRECT_CALL_INET 宏进行间接调用;
IP 层:inet_gro_receive()
- 检查 IP 头部(源目 IP、TTL、TOS、分片信息等),
- 在哈希桶内的
gro_list 中查找可合并的流。
- 如果 TTL 或 TOS 不一致、是分片包等场景,则启动 flush 操作;
传输层:udp4_gro_receive()
- 查找对应的 socket(用于检查
udp_sk(sk)->gro_enabled),这sk在udp tunnel类的场景(如vxlan)下有用,其他场景可忽略。
- 然后由
udp_gro_receive() 执行端口匹配和长度校验,
- 最终通过
skb_gro_receive() 完成数据合并;
完成层:napi_skb_finish()
- 根据
dev_gro_receive() 的返回值决定数据包的最终去向
- ——合并:有两种方式,详见后续专门讲述。
GRO_MERGED:把当前udp报文合并到之前的udp报文里。简单,效率低。
GRO_MERGED_FREE:保留报文大小。复杂,高效。
- 消费(
GRO_CONSUMED)
- 或走普通路径(
GRO_NORMAL)。
2. 数据结构
2.1. 核心: struct offload_callbacks
内核里的GRO是分层处理的, 普通报文包含两层
- 网络层:IPv4、IPv6。
- 传输层:TCP、UDP。
尽管层级不一样,协议不一样,但是他们的GRO处理都有个共同的抽象结构体offload_callbacks
1 2 3 4 5 6 7
| 2704 struct offload_callbacks { 2705 struct sk_buff *(*gso_segment)(struct sk_buff *skb, 2706 netdev_features_t features); 2707 struct sk_buff *(*gro_receive)(struct list_head *head, 2708 struct sk_buff *skb); 2709 int (*gro_complete)(struct sk_buff *skb, int nhoff); 2710 };
|
- gso_segment:GRO报文穿过协议栈,到达socket时候,如何还原切分?
- gro_receive:RX软中断在协议出口,尝试组装GRO报文时,判断报文是否可以GRO。
- gro_complete:如果GRO解释,。。。
2.2. IP层

- 每个IP层的offload/GRO操作都汇总到packet_offload结构体中。
- type: 对应eth头里的proto字段,如IPv4,IPv6.
- struct offload_callbacks callbacks: 协议对应的GRO处理函数集。
- 所有的结构体通过一个list 挂载到
offload_base链表下。
2.2.1. offload_base
1
| 13 struct list_head offload_base __read_mostly = LIST_HEAD_INIT(offload_base);
|
2.2.2. packet_offload
1 2 3 4 5 6
| 2712 struct packet_offload { 2713 __be16 type; /* This is really htons(ether_type). */ 2714 u16 priority; 2715 struct offload_callbacks callbacks; 2716 struct list_head list; 2717 };
|
2.2.2.1. IPv4: ip_packet_offload
1 2 3 4 5 6 7 8 9
| 1902 static struct packet_offload ip_packet_offload __read_mostly = { 1903 .type = cpu_to_be16(ETH_P_IP), 1904 .callbacks = { 1905 .gso_segment = inet_gso_segment, 1906 .gro_receive = inet_gro_receive, 1907 .gro_complete = inet_gro_complete, 1908 }, 1909 }; 1910
|
2.2.2.2. IPv6: ipv6_packet_offload
1 2 3 4 5 6 7 8
| 394 static struct packet_offload ipv6_packet_offload __read_mostly = { 395 .type = cpu_to_be16(ETH_P_IPV6), 396 .callbacks = { 397 .gso_segment = ipv6_gso_segment, 398 .gro_receive = ipv6_gro_receive, 399 .gro_complete = ipv6_gro_complete, 400 }, 401 };
|
2.3. TCP/UDP层
- `` 与Ip层处理类似,也有一个传输层的结构体存放TCP/UDP的GRO操作。
inet_offloads

2.3.1. inet_offloads
1
| 29 const struct net_offload __rcu *inet_offloads[MAX_INET_PROTOS] __read_mostly;
|
2.3.2. net_offload
1 2 3 4
| 68 struct net_offload { 69 struct offload_callbacks callbacks; 70 unsigned int flags; /* Flags used by IPv6 for now */ 71 };
|
2.3.3. TCP4: tcpv4_offload
1 2 3 4 5 6 7
| 347 static const struct net_offload tcpv4_offload = { 348 .callbacks = { 349 .gso_segment = tcp4_gso_segment, 350 .gro_receive = tcp4_gro_receive, 351 .gro_complete = tcp4_gro_complete, 352 }, 353 };
|
2.3.4. UDP4: udpv4_offload
1 2 3 4 5 6 7
| 740 static const struct net_offload udpv4_offload = { 741 .callbacks = { 742 .gso_segment = udp4_ufo_fragment, 743 .gro_receive = udp4_gro_receive, 744 .gro_complete = udp4_gro_complete, 745 }, 746 };
|
3. 内核函数
3.1. 入口函数 napi_gro_receive
1 2 3 4 5 6 7 8
| 600 gro_result_t napi_gro_receive(struct napi_struct *napi, struct sk_buff *skb) 601 { ... 609 ret = napi_skb_finish(napi, skb, dev_gro_receive(napi, skb)); 611 612 return ret; 613 } 614 EXPORT_SYMBOL(napi_gro_receive);
|
napi_gro_receive() 是 GRO 处理的统一入口,由网卡驱动在 NAPI poll 函数中调用(驱动需设置 NETIF_F_GRO 特性)。
其核心逻辑在dev_gro_receive()里。
napi_skb_finish() 根据返回值完成最终处理。
3.2. 核心分发函数 dev_gro_receive
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31
| 437 static enum gro_result dev_gro_receive(struct napi_struct *napi, struct sk_buff *skb) 438 { 439 u32 bucket = skb_get_hash_raw(skb) & (GRO_HASH_BUCKETS - 1); 440 struct gro_list *gro_list = &napi->gro_hash[bucket]; 441 struct list_head *head = &offload_base; 442 struct packet_offload *ptype; 443 __be16 type = skb->protocol; ... 454 list_for_each_entry_rcu(ptype, head, list) { 455 if (ptype->type == type && ptype->callbacks.gro_receive) 456 goto found_ptype; 457 } ... 461 found_ptype: 462 skb_set_network_header(skb, skb_gro_offset(skb)); ... 490 pp = INDIRECT_CALL_INET(ptype->callbacks.gro_receive, 491 ipv6_gro_receive, inet_gro_receive, 492 &gro_list->list, skb); ... 501 same_flow = NAPI_GRO_CB(skb)->same_flow; 502 ret = NAPI_GRO_CB(skb)->free ? GRO_MERGED_FREE : GRO_MERGED; 503 504 if (pp) { 505 skb_list_del_init(pp); 506 napi_gro_complete(napi, pp); 507 gro_list->count--; 508 } ... 537 return ret; 543 }
|
dev_gro_receive() 是 GRO 处理的核心分发函数,其关键步骤如下:
计算哈希链入口:skb_get_hash_raw() 获取数据包的哈希值(由网卡RSS或软件计算),找到hash桶里的链表入口。
查找网络层处理协议:遍历全局 offload_base 链表,根据 skb->protocol 匹配对应的 packet_offload 类型。
- IPv4 对应
inet_gro_receive。
- IPv6 对应
ipv6_gro_receive。
调用协议处理GRO:通过 INDIRECT_CALL_INET 宏调用协议层的 GRO 处理函数。
处理协议返回结果:处理协议层返回 pp。
- 如果
pp 非空:有一个流的GRO缓存内容,需要被flush。调用 napi_gro_complete() 将其送往协议栈,并减少桶计数。
3.3. IP 层处理 inet_gro_receive
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22
| 1471 struct sk_buff *inet_gro_receive(struct list_head *head, struct sk_buff *skb) 1472 { 1473 const struct net_offload *ops; ... 1489 proto = iph->protocol; 1490 1491 ops = rcu_dereference(inet_offloads[proto]); 1492 if (!ops || !ops->callbacks.gro_receive) ... 1509 list_for_each_entry(p, head, list) { 1562 } ... 1577 pp = indirect_call_gro_receive(tcp4_gro_receive, udp4_gro_receive, 1578 ops->callbacks.gro_receive, head, skb); 1579 1580 out: 1581 skb_gro_flush_final(skb, pp, flush); 1582 1583 return pp; 1584 }
|
inet_gro_receive() 负责 IP 层的流匹配和 flush 判断:
- 确定传输层协议的GRO:根据
iph->protocol 从 inet_offloads[] 数组查找对应的传输层协议回调
- TCP 对应
tcp4_gro_receive。
- UDP 对应
udp4_gro_receive。
- 流匹配:遍历当前哈希桶内的 skb 列表,
- 检查 IP 层信息是否一致——源目 IP 地址、TTL、TOS、是否分片等。
- 任何一项不匹配都会触发
flush 标志,停止当前流的GSO缓存并flush。
- 调用传输层: 调用具体的传输层 GRO 函数。
3.4. UDP层处理入口:udp4_gro_receive
1 2 3 4 5 6 7 8 9 10 11 12
| 621 INDIRECT_CALLABLE_SCOPE 622 struct sk_buff *udp4_gro_receive(struct list_head *head, struct sk_buff *skb) 623 { 624 struct udphdr *uh = udp_gro_udphdr(skb); ... 644 if (static_branch_unlikely(&udp_encap_needed_key)) 645 sk = udp4_gro_lookup_skb(skb, uh->source, uh->dest); 646 647 pp = udp_gro_receive(head, skb, uh, sk); 648 return pp; ... 653 }
|
udp4_gro_receive() 是 UDP 协议的 GRO 入口,无实质意义,核心函数在udp_gro_receive。
注:
- Socket 查找(L644-L645):在
udp_encap_needed_key 静态分支开启时,调用 udp4_gro_lookup_skb() 根据源/目的端口查找对应的 UDP socket。这个 socket 用于检查 udp_sk(sk)->gro_enabled 是否开启,以决定是否允许 GRO 合并。
udp_gro_receive() 内部执行端口匹配(比较源/目的端口是否一致)、长度校验(防止 GRO 合并过大),以及决定使用普通合并(skb_gro_receive)还是 fraglist 合并(skb_gro_receive_list)。
3.4.1. UDP层处理函数:udp_gro_receive
1 2 3 4 5 6 7 8 9 10
| 545 struct sk_buff *udp_gro_receive(struct list_head *head, struct sk_buff *skb, 546 struct udphdr *uh, struct sock *sk) 547 { ... 562 if ((!sk && (skb->dev->features & NETIF_F_GRO_UDP_FWD)) || 563 (sk && udp_sk(sk)->gro_enabled) || NAPI_GRO_CB(skb)->is_flist) 564 return call_gro_receive(udp_gro_receive_segment, head, skb) ... 604 } 605 EXPORT_SYMBOL(udp_gro_receive)
|
3.4.2. UDP层核心处理函数:udp_gro_receive_segment
- 比对源目port,确保在一个五元组后缓存。 continue是考虑有多个流的场景。
- list 和非list模式不混用。
- 报文大小要一致, 方便以后切割。 list模式可以接受报文不一样大小。
- checksum不一致,也会触发flush。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41
| 464 static struct sk_buff *udp_gro_receive_segment(struct list_head *head, 465 struct sk_buff *skb) 466 { 467 struct udphdr *uh = udp_gro_udphdr(skb) 468 struct sk_buff *pp = NULL ... 489 list_for_each_entry(p, head, list) { ... // 确保是同一个udp流,这里用u32检查源目端口。 495 /* Match ports only, as csum is always non zero */ 496 if ((*(u32 *)&uh->source != *(u32 *)&uh2->source)) { 497 NAPI_GRO_CB(p)->same_flow = 0 498 continue 499 } ... //下面是核心代码。 506 /* Terminate the flow on len mismatch or if it grow "too much". 507 * Under small packet flood GRO count could elsewhere grow a lot 508 * leading to excessive truesize values. 509 * On len mismatch merge the first packet shorter than gso_size, 510 * otherwise complete the GRO packet. 511 */ 512 if (ulen > ntohs(uh2->len)) { // 如果当前报文比以前的大, 不缓存并启动flush。 513 pp = p 514 } else { // list模式和非list模式后面专题解释。 515 if (NAPI_GRO_CB(skb)->is_flist) { 516 if (!pskb_may_pull(skb, skb_gro_offset(skb))) { 517 NAPI_GRO_CB(skb)->flush = 1 518 return NULL 519 } 520 if ((skb->ip_summed != p->ip_summed) || 521 (skb->csum_level != p->csum_level)) { 522 NAPI_GRO_CB(skb)->flush = 1 523 return NULL 524 } 525 ret = skb_gro_receive_list(p, skb) 526 } else { 527 skb_gro_postpull_rcsum(skb, uh, 528 sizeof(struct udphdr)) 529 530 ret = skb_gro_receive(p, skb) 531 } 532 }
|
4. NAPI相关部分
4.1. 数据结构
4.1.1. napi_struct里GRO相关字段)
1 2 3 4 5 6
| 352 struct napi_struct { ... 373 struct gro_list gro_hash[GRO_HASH_BUCKETS]; 374 struct sk_buff *skb; ... 383 };
|
4.1.2. GRO_HASH_BUCKETS
1
| 347 #define GRO_HASH_BUCKETS 8
|
4.1.3. gro_list
1 2 3 4
| 338 struct gro_list { 339 struct list_head list; 340 int count; 341 };
|
list:list_head 链表头,用于挂载属于同一哈希桶的所有流(每个流对应一个正在 GRO 合并的 skb)。
count:当前桶内挂载的 skb 数量,用于控制流数量上限(MAX_GRO_SKBS)。
gro_hash[8]:8 个 GRO 哈希桶组成的数组。
skb:当前正在处理的 skb 指针,用于 NAPI poll 上下文。
在 Linux 5.x 之后的内核中,gro_hash 数组的引入使得 GRO 在多个并行流同时活跃时,每个数据包只需在其哈希值对应的桶中查找,而无需遍历整个 per-NAPI 流列表,显著提升了多流场景的性能。
5. skb_gro_receive与skb_gro_receive_list的区别
5.1. 核心区别:融合 (Frags) vs. 串联 (Fraglist)
| 特性 |
skb_gro_receive() |
skb_gro_receive_list() |
| 合并方式 |
数据融合(拷贝进 frags[]) |
数据包串联(挂到 frag_list) |
| 数据拷贝 |
需要(或内存重映射) |
无需 |
| 大小限制 |
受 MAX_SKB_FRAGS 限制 |
受 64KB 总长度限制 |
| 主要场景 |
普通接收路径 |
转发、支持 GRO_FRAGLIST 的设备 |
| 性能开销 |
较高 |
较低(尤其对转发) |
因此,代码中根据 is_flist 标志来选择调用哪个函数,本质上是在数据拷贝开销、内存布局限制和后续处理效率之间做出的权衡。
5.2. skb_gro_receive()
skb_gro_receive() 试图将新 skb 的载荷数据真正“融合”进主 skb p 的存储空间中。它会尝试将数据追加到 p 已有的线性数据区或者其 frags[] 数组里。
这种方式需要处理内存分配、拷贝和 frags[] 数组的管理,开销相对较大,并且受限于 MAX_SKB_FRAGS(通常为 17 个),单个 GRO 包的大小有限。
5.3. skb_gro_receive_list()
skb_gro_receive_list() 不融合载荷数据,而是通过 frag_list 指针,将整个 skb 以链表的形式串联起来。它做的事情非常轻量:
- 将新 skb 挂到主 skb
p 的 skb_shinfo(p)->frag_list 链表尾部。
- 更新链表指针
NAPI_GRO_CB(p)->last 指向新的尾部。
- 直接累加
p->len、p->data_len 和 p->truesize 等元数据。
6. is_flist 标志与设计动机
代码中通过 NAPI_GRO_CB(skb)->is_flist 来区分这两种模式。这个标志通常由网络设备驱动或上层配置决定,例如当设备支持 NETIF_F_GRO_FRAGLIST 特性时,is_flist 会被设置为 1。
区分这两种方式的核心动机是性能和场景适配:
性能优化(尤其转发场景)
fraglist 方式避免了数据拷贝,在大包转发场景下优势明显。当 GRO 聚合后的数据需要再次以 GSO 形式发送时,fraglist 结构可以直接被 skb_segment_list() 高效处理,无需重新拆分和拷贝,大幅降低了 CPU 开销。
突破 frags[] 数量限制
传统 skb_gro_receive() 受 MAX_SKB_FRAGS(约 17 个)限制,构建的 GRO 包大小有限。fraglist 方式通过串联 skb,可以构建更大的 GRO 包(上限由 65536 字节限制),支持更高效的聚合。
适配硬件/驱动卸载
某些网卡或驱动可能已经将数据组织成 frag_list 形式,GRO 层直接沿用这种结构可以避免不必要的重构。
7. 注意事项
fraglist 方式虽然高效,但有其复杂性。例如,它需要防止 GSO skb 被再次嵌套聚合,否则可能导致后续分段路径处理异常。因此,代码中在进入 skb_gro_receive_list() 前会检查 flush 标志等条件。