<feed xmlns='http://www.w3.org/2005/Atom'>
<title>linux-toradex.git/net/sched, branch master</title>
<subtitle>Linux kernel for Apalis and Colibri modules</subtitle>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/'/>
<entry>
<title>Merge tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf</title>
<updated>2026-10-01T10:00:40+00:00</updated>
<author>
<name>Paolo Abeni</name>
<email>pabeni@redhat.com</email>
</author>
<published>2026-10-01T10:00:40+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=e23a64eb244356ee47c0620f0722d51bd88db522'/>
<id>e23a64eb244356ee47c0620f0722d51bd88db522</id>
<content type='text'>
Pablo Neira Ayuso says:

====================
Netfilter/IPVS fixes for net

The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:

1) Expand existing ipset fix for bitmap sets to disallow comments
   updates from kernel-side adds, from Florian Westphal.

2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
   flowtable cannot ever be removed, from Aohan Mei.

3) nft_rbtree GC should collect end elements that contained in
   this transaction batch, new or deleted elements are never
   expired. From Weiming Shi.

4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
   values, from Fernando F. Mancera.

5) Flowtable GC must skip flows that are pending hardware updates,
   generalize the PENDING flag and use it to inhibit GC.

6) Restore flowtable with ieee80211 which broke due to a relatively
   recent commit, which was pulled in by -stable, causing a regression
   in 6.18 kernels.

And the following IPVS fixes:

1) Fix accounting of cache entries in IPVS LBLC for destinations,
   which eventually fills up the table and trigger recurrent
   resizing, from Julian Anastasov.

2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
   from Zhiling Zou.

3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
   do not allow to use it with templates. Also from Julian.

4) Sanitize flags in IPVS sync messages received in the backup.
   From Julian Anastasov.

netfilter pull request 26-09-30

* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
  netfilter: flowtable: restore ieee80211 forward path
  netfilter: flowtable: generalize pending status bit
  netfilter: bpf: reject invalid NAT manipulation types
  netfilter: nft_set_rbtree: skip transaction elements during GC
  ipvs: filter some flags received in the backup server
  ipvs: do not create invisible templates
  ipvs: bound LBLCR and LBLC cache growth
  ipvs: fix missing counter decrement in lblc
  netfilter: nft_flow_offload: drop flowtable reference on init error path
  netfilter: ipset: do not update comments from kernel-side adds
====================

Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Pablo Neira Ayuso says:

====================
Netfilter/IPVS fixes for net

The following batch contains Netfilter fixes for net. This batch
fixes crashes as recent feature regression, one of the due to a
dependency that has been pulled into -stable:

1) Expand existing ipset fix for bitmap sets to disallow comments
   updates from kernel-side adds, from Florian Westphal.

2) Drop flowtable reference if nf_ct_netns_get() fails, otherwise
   flowtable cannot ever be removed, from Aohan Mei.

3) nft_rbtree GC should collect end elements that contained in
   this transaction batch, new or deleted elements are never
   expired. From Weiming Shi.

4) Restrict nf_nat_bpf so it does not set unknown NF_NAT_MANIP_*
   values, from Fernando F. Mancera.

5) Flowtable GC must skip flows that are pending hardware updates,
   generalize the PENDING flag and use it to inhibit GC.

6) Restore flowtable with ieee80211 which broke due to a relatively
   recent commit, which was pulled in by -stable, causing a regression
   in 6.18 kernels.

And the following IPVS fixes:

1) Fix accounting of cache entries in IPVS LBLC for destinations,
   which eventually fills up the table and trigger recurrent
   resizing, from Julian Anastasov.

2) Limit IPVS cache growth for LBLCR and LBLC schedulers,
   from Zhiling Zou.

3) Restrict IP_VS_CONN_F_ONE_PACKET for normal connections,
   do not allow to use it with templates. Also from Julian.

4) Sanitize flags in IPVS sync messages received in the backup.
   From Julian Anastasov.

netfilter pull request 26-09-30

* tag 'nf-26-09-30' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
  netfilter: flowtable: restore ieee80211 forward path
  netfilter: flowtable: generalize pending status bit
  netfilter: bpf: reject invalid NAT manipulation types
  netfilter: nft_set_rbtree: skip transaction elements during GC
  ipvs: filter some flags received in the backup server
  ipvs: do not create invisible templates
  ipvs: bound LBLCR and LBLC cache growth
  ipvs: fix missing counter decrement in lblc
  netfilter: nft_flow_offload: drop flowtable reference on init error path
  netfilter: ipset: do not update comments from kernel-side adds
====================

Link: https://patch.msgid.link/20260930074142.298353-1-pablo@netfilter.org
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>netfilter: flowtable: generalize pending status bit</title>
<updated>2026-09-30T07:35:20+00:00</updated>
<author>
<name>Pablo Neira Ayuso</name>
<email>pablo@netfilter.org</email>
</author>
<published>2026-09-23T08:53:02+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=7c549fb7eecd01fdba7c0d12c01da876ef8a3214'/>
<id>7c549fb7eecd01fdba7c0d12c01da876ef8a3214</id>
<content type='text'>
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.

Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.

And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.

Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.

Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso &lt;pablo@netfilter.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.

Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.

And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.

Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.

Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso &lt;pablo@netfilter.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net: cap skb-&gt;queue_mapping when the tx queue is picked</title>
<updated>2026-09-30T00:49:35+00:00</updated>
<author>
<name>Jamal Hadi Salim</name>
<email>jhs@mojatatu.com</email>
</author>
<published>2026-09-28T12:46:16+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=ea4d4b5dddb5121e64f56fd8e0720bd8b447b63d'/>
<id>ea4d4b5dddb5121e64f56fd8e0720bd8b447b63d</id>
<content type='text'>
skbedit can set skb-&gt;queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.

A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb-&gt;queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q-&gt;qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.

We (ab)use the skb-&gt;nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb-&gt;nf_skip_egress bit/flag to skb-&gt;skip_egress

Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb-&gt;nf_skip_egress. An egress qdisc
classifier runs in q-&gt;enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).

Store the value netdev_cap_txqueue() selected back into skb-&gt;queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.

netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb-&gt;nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.

A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".

Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.

Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.

Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative &lt;zdi-disclosures@trendmicro.com&gt;
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet &lt;edumazet@google.com&gt;
Suggested-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
Tested-by: hybris &lt;hybris@mojatatu.ai&gt;
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Eric Dumazet &lt;edumazet@google.com&gt;
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
skbedit can set skb-&gt;queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.

A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb-&gt;queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q-&gt;qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.

We (ab)use the skb-&gt;nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb-&gt;nf_skip_egress bit/flag to skb-&gt;skip_egress

Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb-&gt;nf_skip_egress. An egress qdisc
classifier runs in q-&gt;enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).

Store the value netdev_cap_txqueue() selected back into skb-&gt;queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.

netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb-&gt;nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.

A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".

Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.

Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.

Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative &lt;zdi-disclosures@trendmicro.com&gt;
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet &lt;edumazet@google.com&gt;
Suggested-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
Tested-by: hybris &lt;hybris@mojatatu.ai&gt;
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Eric Dumazet &lt;edumazet@google.com&gt;
Link: https://patch.msgid.link/QDISC-9R8V.v4.20260928081529@mojatatu.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: sch_codel: match the no-drop threshold to the packet size</title>
<updated>2026-09-29T10:49:13+00:00</updated>
<author>
<name>Jamal Hadi Salim</name>
<email>jhs@mojatatu.com</email>
</author>
<published>2026-09-26T18:03:28+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=54518e0e827f4ca9229ae657022c60bf60f5c1bf'/>
<id>54518e0e827f4ca9229ae657022c60bf60f5c1bf</id>
<content type='text'>
commit 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid
disabling CoDel") clamped q-&gt;params.mtu to [256, 1 &lt;&lt; 20]. The value is
the CoDel no-drop threshold (codel_impl.h "*backlog &lt;= params-&gt;mtu"), so
on a link whose maximum transmitted packet size is below 256 the floor
extends CoDel's minimum-backlog exemption beyond one packet and delays
drop or mark eligibility by several small packets.

Keep the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device) but drop the 256 floor, so the
threshold tracks the real device packet size.

Conditions to recreate the bug: attach a codel qdisc on a link whose
MTU plus hard_header_len is below 256 (e.g. a CAN interface). At that
MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Requires CAP_NET_ADMIN in a user namespace.

Fixes: 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid disabling CoDel")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com.2
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
commit 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid
disabling CoDel") clamped q-&gt;params.mtu to [256, 1 &lt;&lt; 20]. The value is
the CoDel no-drop threshold (codel_impl.h "*backlog &lt;= params-&gt;mtu"), so
on a link whose maximum transmitted packet size is below 256 the floor
extends CoDel's minimum-backlog exemption beyond one packet and delays
drop or mark eligibility by several small packets.

Keep the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device) but drop the 256 floor, so the
threshold tracks the real device packet size.

Conditions to recreate the bug: attach a codel qdisc on a link whose
MTU plus hard_header_len is below 256 (e.g. a CAN interface). At that
MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Requires CAP_NET_ADMIN in a user namespace.

Fixes: 6439461f1618 ("net/sched: sch_codel: clamp default mtu to avoid disabling CoDel")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com.2
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: fq_codel: match the no-drop threshold to the packet size</title>
<updated>2026-09-29T10:49:07+00:00</updated>
<author>
<name>Jamal Hadi Salim</name>
<email>jhs@mojatatu.com</email>
</author>
<published>2026-09-26T18:03:27+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=d031465acc362866f27e4f8f79cfecf1037fe2e5'/>
<id>d031465acc362866f27e4f8f79cfecf1037fe2e5</id>
<content type='text'>
commit d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
clamped both q-&gt;quantum and q-&gt;cparams.mtu to [256, FQ_CODEL_QUANTUM_MAX].
The two fields mean different things: quantum is a DRR credit that wants
the 256 floor, but cparams.mtu is the CoDel no-drop threshold
(codel_impl.h "*backlog &lt;= params-&gt;mtu"). On a link whose maximum
transmitted packet size is below 256, the floor extends CoDel's
minimum-backlog exemption beyond one packet and delays drop or mark
eligibility by several small packets.

Split the clamp. quantum keeps [256, FQ_CODEL_QUANTUM_MAX]; cparams.mtu
tracks psched_mtu() (the device MTU plus its hard-header length) with
only the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device).

Conditions to recreate the bug: attach an fq_codel qdisc on a link
whose MTU plus hard_header_len is below 256 (e.g. a CAN interface). At
that MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Basic Testing done: with dev-&gt;mtu=100 and hard_header_len=14, a
return probe on fq_codel_init() observed cparams.mtu change from 256 to 114

Fixes: d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
commit d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
clamped both q-&gt;quantum and q-&gt;cparams.mtu to [256, FQ_CODEL_QUANTUM_MAX].
The two fields mean different things: quantum is a DRR credit that wants
the 256 floor, but cparams.mtu is the CoDel no-drop threshold
(codel_impl.h "*backlog &lt;= params-&gt;mtu"). On a link whose maximum
transmitted packet size is below 256, the floor extends CoDel's
minimum-backlog exemption beyond one packet and delays drop or mark
eligibility by several small packets.

Split the clamp. quantum keeps [256, FQ_CODEL_QUANTUM_MAX]; cparams.mtu
tracks psched_mtu() (the device MTU plus its hard-header length) with
only the upper bound that guards the original overflow (psched_mtu()
wrapping to ~2 GiB on a huge-MTU device).

Conditions to recreate the bug: attach an fq_codel qdisc on a link
whose MTU plus hard_header_len is below 256 (e.g. a CAN interface). At
that MTU the no-drop threshold must equal the device MTU plus its
hard-header length; before this patch it was forced to 256.
Basic Testing done: with dev-&gt;mtu=100 and hard_header_len=14, a
return probe on fq_codel_init() observed cparams.mtu change from 256 to 114

Fixes: d9ebd8f9aa8b ("net/sched: fq_codel: clamp default quantum and mtu")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260819143136.57350-1-jhs@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/QDISC-34MS.v1.20260925165535@mojatatu.com
Signed-off-by: Paolo Abeni &lt;pabeni@redhat.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: cls_api: reclaim an empty proto on the error path</title>
<updated>2026-09-26T00:19:23+00:00</updated>
<author>
<name>Jamal Hadi Salim</name>
<email>jhs@mojatatu.com</email>
</author>
<published>2026-09-24T08:32:49+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=0bb8eab29b55045281a4963d4557f2776cdf8cfa'/>
<id>0bb8eab29b55045281a4963d4557f2776cdf8cfa</id>
<content type='text'>
Two racing tc filter add requests on the same chain/prio of an
unlocked classifier both run change() on the shared proto and both
can fail: the winner's tcf_chain_tp_delete_empty() attempt gives up
because the loser's handle is still in the idr, and the loser's
error path drops only its own reference without a second reclamation
attempt. The empty proto stays linked in the chain, holding the
chain reference, a block reference and the classifier module
reference until the chain or block is torn down.

Reclaim the proto on the error path of any failed request that holds
a proto reference. The reclamation is emptiness-gated:
tcf_chain_tp_delete_empty() unlinks the proto only when
delete_empty() admits it is empty, so a live shared proto is never
unlinked. A proto the request created is reclaimed unconditionally -
it is the only owner, so marking it for deletion is safe even without
a delete_empty callback. Classifiers without one (the check marks the
proto unconditionally) are rtnl-serialized, so the raced window this
guard closes cannot arise for them.

This is a follow-up to commit d4e359b3608a ("net/sched: cls_api: fix
teardown of an adopted proto on insert-race loss"), which stopped the
loser of the insert race from unlinking the winner's live proto but
left the empty-proto residual in place.

Conditions to recreate:
- CONFIG_NET_CLS_FLOWER=y; veth pair
- tc qdisc add dev veth0 ingress
- two concurrent `tc filter add dev veth0 ingress protocol ip pref 1
  flower skip_sw ... action drop` (both fail in fl_hw_replace_filter
  after publishing their handle in the idr); repeat in a loop
- an empty flower tp stays linked after both requests fail; visible
  as a bare `filter protocol ip pref 1 flower chain 0` header in
  `tc filter show` with no filter entries
- CAP_NET_ADMIN (namespace-local via unshare -Urn suffices)

Fixes: 8b64678e0af8 ("net: sched: refactor tp insert/delete for concurrent execution")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260805134049.927864-1-victor@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Simon Horman &lt;horms@kernel.org&gt;
Link: https://patch.msgid.link/QDISC-JCOT.v1.20260910090924@mojatatu.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Two racing tc filter add requests on the same chain/prio of an
unlocked classifier both run change() on the shared proto and both
can fail: the winner's tcf_chain_tp_delete_empty() attempt gives up
because the loser's handle is still in the idr, and the loser's
error path drops only its own reference without a second reclamation
attempt. The empty proto stays linked in the chain, holding the
chain reference, a block reference and the classifier module
reference until the chain or block is torn down.

Reclaim the proto on the error path of any failed request that holds
a proto reference. The reclamation is emptiness-gated:
tcf_chain_tp_delete_empty() unlinks the proto only when
delete_empty() admits it is empty, so a live shared proto is never
unlinked. A proto the request created is reclaimed unconditionally -
it is the only owner, so marking it for deletion is safe even without
a delete_empty callback. Classifiers without one (the check marks the
proto unconditionally) are rtnl-serialized, so the raced window this
guard closes cannot arise for them.

This is a follow-up to commit d4e359b3608a ("net/sched: cls_api: fix
teardown of an adopted proto on insert-race loss"), which stopped the
loser of the insert race from unlinking the winner's live proto but
left the empty-proto residual in place.

Conditions to recreate:
- CONFIG_NET_CLS_FLOWER=y; veth pair
- tc qdisc add dev veth0 ingress
- two concurrent `tc filter add dev veth0 ingress protocol ip pref 1
  flower skip_sw ... action drop` (both fail in fl_hw_replace_filter
  after publishing their handle in the idr); repeat in a loop
- an empty flower tp stays linked after both requests fail; visible
  as a bare `filter protocol ip pref 1 flower chain 0` header in
  `tc filter show` with no filter entries
- CAP_NET_ADMIN (namespace-local via unshare -Urn suffices)

Fixes: 8b64678e0af8 ("net: sched: refactor tp insert/delete for concurrent execution")
Reported-by: Sashiko (nipa) &lt;sashiko-bot@kernel.org&gt;
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260805134049.927864-1-victor@mojatatu.com
Signed-off-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Simon Horman &lt;horms@kernel.org&gt;
Link: https://patch.msgid.link/QDISC-JCOT.v1.20260910090924@mojatatu.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: sch_teql: fix shadowed err in __teql_resolve()</title>
<updated>2026-09-24T18:00:59+00:00</updated>
<author>
<name>Eric Dumazet</name>
<email>edumazet@google.com</email>
</author>
<published>2026-09-24T08:29:50+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=907b978e82cb4c1c245fc2985bb27c5d5c88c8f6'/>
<id>907b978e82cb4c1c245fc2985bb27c5d5c88c8f6</id>
<content type='text'>
__teql_resolve() declares an inner 'int err;' inside the
'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
'int err = 0;'. As a result, a negative return from dev_hard_header()
is written to the inner variable and __teql_resolve() still returns 0.

Remove the shadowed variable and set the outer err to -EINVAL when
dev_hard_header() returns a negative error.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Cc: Jiri Pirko &lt;jiri@resnulli.us&gt;
Signed-off-by: Eric Dumazet &lt;edumazet@google.com&gt;
Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
__teql_resolve() declares an inner 'int err;' inside the
'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
'int err = 0;'. As a result, a negative return from dev_hard_header()
is written to the inner variable and __teql_resolve() still returns 0.

Remove the shadowed variable and set the outer err to -EINVAL when
dev_hard_header() returns a negative error.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Cc: Jiri Pirko &lt;jiri@resnulli.us&gt;
Signed-off-by: Eric Dumazet &lt;edumazet@google.com&gt;
Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: act_ct: fix helper UAF due to extensions realloc</title>
<updated>2026-09-24T16:56:02+00:00</updated>
<author>
<name>Ilya Maximets</name>
<email>i.maximets@ovn.org</email>
</author>
<published>2026-09-21T14:55:48+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=dad19b59da050cb60d3f7023dac2a042a84bf0bd'/>
<id>dad19b59da050cb60d3f7023dac2a042a84bf0bd</id>
<content type='text'>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:

  -&gt; nf_ct_helper()
   -&gt; helper-&gt;help()
    -&gt; nf_ct_expect_related_report()
     -&gt; nf_ct_expect_insert()
      -&gt; hlist_add_head_rcu(&amp;exp-&gt;lnode, &amp;master_help-&gt;expectations)

In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.

Make sure that helpers are called at the end after all the other
extensions are already added.

Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure.  And
there are no atomicity guarantees provided by the API anyway.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk &lt;axel.mierczuk@1password.com&gt;
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:

  -&gt; nf_ct_helper()
   -&gt; helper-&gt;help()
    -&gt; nf_ct_expect_related_report()
     -&gt; nf_ct_expect_insert()
      -&gt; hlist_add_head_rcu(&amp;exp-&gt;lnode, &amp;master_help-&gt;expectations)

In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.

Make sure that helpers are called at the end after all the other
extensions are already added.

Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure.  And
there are no atomicity guarantees provided by the API anyway.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk &lt;axel.mierczuk@1password.com&gt;
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: act_ct: remove 'add_helper' dead code</title>
<updated>2026-09-24T16:56:02+00:00</updated>
<author>
<name>Ilya Maximets</name>
<email>i.maximets@ovn.org</email>
</author>
<published>2026-09-21T14:55:47+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=00df72e39f306e2f7adb68528a5c92109da6a0a9'/>
<id>00df72e39f306e2f7adb68528a5c92109da6a0a9</id>
<content type='text'>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed.  So, it can be
treated as being always false and just removed.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed.  So, it can be
treated as being always false and just removed.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>net/sched: act_ct: avoid modifying shared unconfirmed ct entry</title>
<updated>2026-09-24T16:56:02+00:00</updated>
<author>
<name>Ilya Maximets</name>
<email>i.maximets@ovn.org</email>
</author>
<published>2026-09-21T14:55:46+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=f85009dfcd65e5969526b0db7a49b5413746e630'/>
<id>f85009dfcd65e5969526b0db7a49b5413746e630</id>
<content type='text'>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.

The series of events:

 1. The first clone wants to commit and runs the helpers wiring up
    the extension pointer into the expectation list.
 2. Then it looses the confirmation keeping the entry unconfirmed.
 3. Second clone now wants to commit labels or run NAT and adds the
    new extension for that breaking the pointer in the expectation
    list causing UAF on the destruction path later.

While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone.  So, let's just reset the entry in
case for some reason we got an skb with a shared one.  This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.

Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break.  So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space.  This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.

The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk &lt;axel.mierczuk@1password.com&gt;
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.

The series of events:

 1. The first clone wants to commit and runs the helpers wiring up
    the extension pointer into the expectation list.
 2. Then it looses the confirmation keeping the entry unconfirmed.
 3. Second clone now wants to commit labels or run NAT and adds the
    new extension for that breaking the pointer in the expectation
    list causing UAF on the destruction path later.

While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone.  So, let's just reset the entry in
case for some reason we got an skb with a shared one.  This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.

Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break.  So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space.  This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.

The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.

Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk &lt;axel.mierczuk@1password.com&gt;
Signed-off-by: Ilya Maximets &lt;i.maximets@ovn.org&gt;
Reviewed-by: Aaron Conole &lt;aconole@redhat.com&gt;
Reviewed-by: Xin Long &lt;lucien.xin@gmail.com&gt;
Reviewed-by: Jamal Hadi Salim &lt;jhs@mojatatu.com&gt;
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski &lt;kuba@kernel.org&gt;
</pre>
</div>
</content>
</entry>
</feed>
