| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from wireless, wireguard, CAN and Bluetooth.
We have one known regression to wrap up in VLAN handling.
Current release - regressions:
- Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex
Previous releases - regressions:
- can: fix regression in handling RPS after migrating metadata to skb_ext
- eth:
- iavf: fix regressions in reconfig impacting bonding
- mana: fix packet forwarding performance regression
- stmmac: remove buggy VLAN acceleration support
Previous releases - always broken:
- a few high prio fixes for tun, and af_packet
- amt: fix a UaF on tunnel teardown
- eth:
- bnxt: fix PCIe AER recovery and FLR handling issues
- macb: don't modify Tx skbs before taking ownership
- axienet: don't leak Tx skbs on interface stop
- wifi:
- nxpwifi: number of LLM-ish fixes
- assorted mt76 fixes"
* tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits)
net: macb: copy shared skbs before appending the FCS
net: macb: check TX ring before modifying skb
vsock: Fix memory leak in vmci_transport_recv_dgram_cb()
wireguard: noise: reject response consumption after intermediate initiation
wireguard: queueing: preserve tstamp_type when encapsulating packet
net: openvswitch: validate transport header presence in set_ipv6_addr
net/smc: protect clcsock lifetime in smc_getname
ipv6: do not warn on route notification size race
ipv4: do not warn on route notification size race
ipv4: validate checksum_start before completing checksum
ptp: ocp: fix PCIe delay estimation calculation
xen/netfront: don't leak the skb when xennet_fill_frags() fails
net/packet: call packet_parse_headers after virtio_net_hdr_to_skb
xen/netfront: drop RX packets with a short Ethernet header
net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()
net: sparx5: free the matchall entry on destroy
selftests: mlxsw: Test port range occupancy on template create
mlxsw: spectrum_flower: Fix port range register leak in tmplt_create()
net: dsa: microchip: fix KSZ8765 fiber detection
net/mlx5e: Order ICOSQ cc update after CQ doorbell
...
|
|
macb_pad_and_fcs() appends the FCS in place when the skb has tailroom.
A shared skb, as pktgen sends in clone_skb mode, grows by one FCS per
transmit. BQL then completes more bytes than were queued and
dql_completed() hits its BUG_ON.
On a Raspberry Pi CM5 (RP1 GEM) pktgen with clone_skb 1000 burst 32 at
60 bytes kills the box within seconds.
Copy shared skbs before appending the FCS. Clearing IFF_TX_SKB_SHARING
would also fix it but makes pktgen refuse clone_skb on macb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-2-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
macb_pad_and_fcs() replaces or extends the skb before the ring space
check. On NETDEV_TX_BUSY the stack requeues an skb that is already freed
or grown.
Check the ring first, using the padded length for the descriptor count.
Nonlinear skbs always take the copy path so the count can assume a
linear skb.
Fixes: 653e92a9175e ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20261006-nb-macb-shared-skb-net-v1-1-a80641479041@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Two threads begin processing the identical response message, received
twice. The first thread, A, runs. While it's running, the second one,
B, gets partway through, and during that slow calculation, or even while
blocking on down_write(), A completes and then also a handshake
initiation that's already been queued up runs in thread C, which itself
takes that same down_write(). The handshake initiation creation
succeeds, and sets the state back to waiting-for-response, and calls
up_write(), at which point thread B resumes, because either its finished
its calculations or was finally allowed to acquire down_write(). Thread
B then copies the state back to the peer, and begins a new session,
using that state, which is the same session as the one made in thread A.
Thread A Thread B Thread C
down_read()
sA = handshake->state
memcpy(cA, handshake->crypto)
up_read()
if (sA != 1)
goto fail
slow_crypto(cA)
down_read()
sB = handshake->state
memcpy(cB, handshake->crypto)
up_read()
if (sB != 1)
goto fail
slow_crypto(cB)
down_write()
if (sA != handshake->state)
goto fail
memcpy(handshake->crypto, cA)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
down_write()
slow_crypto(handshake->crypto)
handshake->state = 1
up_write()
down_write()
if (sB != handshake->state)
goto fail
memcpy(handshake->crypto, cB)
handshake->state = 2
up_write()
down_write()
if (handshake->state != 2)
goto fail
derive_session(handshake->crypto)
up_write()
This seems basically impossible to hit in a meaningful way in practice,
but ensure that it absolutely cannot happen by comparing the ephemeral
private key that's on the stack with the latest one that the peer's
handshake state has.
Cc: stable@vger.kernel.org
Fixes: e7096c131e51 ("net: WireGuard secure network tunnel")
Reported-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-4-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Sending traffic through a wireguard tunnel on a host using the fq
qdisc fills the log with:
fq: likely mono tstamp with tstamp_type 0
An skb carries a timestamp in skb->tstamp and, separately, a
skb->tstamp_type field recording which clock that timestamp came from.
The two have to agree.
When wireguard encapsulates a packet it calls wg_reset_packet(), which
clears the fields that must not leak from the inner packet into the
tunnel packet. It does so in two steps:
skb_scrub_packet(skb, true);
memset(&skb->headers, 0, sizeof(skb->headers));
skb_scrub_packet() deliberately keeps skb->tstamp when it holds a
monotonic timestamp: that value is the time the packet is scheduled to
be sent, and the qdisc still needs it. The memset then zeroes
skb->tstamp_type, because that field sits inside the headers group
while skb->tstamp does not. The packet therefore leaves wireguard
carrying a monotonic timestamp labelled as a realtime one.
Nothing noticed until commit c4f796c4f16b ("net_sched: sch_fq: convert
skb->tstamp if not monotonic"): fq used to assume every timestamp was
monotonic. It now consults tstamp_type, spots the mismatch, warns, and
falls back to treating the value as monotonic. Pacing still ends up
correct, so the log spam is the actual problem.
Save tstamp_type before the memset and restore it when encapsulating,
next to the hash fields that are already carried over this way. When
decapsulating it stays zeroed, which is right: an incoming packet's
timestamp is a realtime receive timestamp.
Fixes: d98d58a00261 ("net: Set skb->mono_delivery_time and clear it after sch_handle_ingress()")
Signed-off-by: Ramses de Norre <ramses@well-founded.dev>
Reviewed-by: Toke Høiland-Jørgensen <toke@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
Link: https://patch.msgid.link/20261008130124.724119-3-Jason@zx2c4.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The commit in fixes introduced a high cap for delayas U64_MAX value
while ktime_t is actually s64. This is wrong cap as it becomes negative
value and any comparison to a real delay will fail to update delay
value. Use KTIME_MAX constant as correct max cap for PCIe delay.
The issue was hit in production (a negative value is observed):
# cat /sys/class/timecard/ocp0/ts_window_adjust
-3
Fixes: aa05fe67bcd64 ("ptp: ocp: Improve PCIe delay estimation")
Signed-off-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261007203359.417270-1-vadim.fedorenko@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When a response chain has more slots than fit in the skb's frags,
xennet_fill_frags() returns an error and xennet_poll() jumps to its
error path. That path moves what's left on tmpq to errq to be freed,
but the skb being filled was already dequeued from tmpq, so it's never
freed. Each chain that overflows leaks the skb and the pages attached
to it as frags, and the backend decides how many slots it sends.
Put the skb back on tmpq before taking the error path, like the
xennet_set_skb_gso() failure just above it does.
Fixes: ad4f15dc2c70 ("xen/netfront: don't bug in case of too many frags")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Reviewed-by: Juergen Gross <jgross@suse.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-fill-frags-leak-v1-1-a8a01ff9cd52@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
handle_incoming_queue() pulls pull_to bytes into the head before
calling eth_type_trans(). pull_to is the length of the first RX slot,
capped at RX_COPY_THRESHOLD, and that length comes from the backend.
Nothing checks it against ETH_HLEN.
If the first slot is shorter than ETH_HLEN and more slots follow, the
head ends up shorter than an Ethernet header while skb->len is longer,
and eth_type_trans() BUG()s in __skb_pull(). If the whole packet is
shorter than ETH_HLEN, eth_type_trans() reads the header past the end
of the data instead.
Pull at least ETH_HLEN, and drop the packet if that fails, which also
drops packets too short to hold an Ethernet header. This also checks
the return value of the pull, which was ignored.
Fixes: 0d160211965b ("xen: add virtual network device driver")
Cc: stable@vger.kernel.org
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Link: https://patch.msgid.link/20261007-b4-xen-netfront-short-head-v1-1-12d7113a7e4e@toxicpanda.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
sparx5_tc_matchall_replace() allocates a struct sparx5_mall_entry for
every offloaded matchall filter and adds it to sparx5->mall_entries.
sparx5_tc_matchall_destroy() removes the entry from the list, but never
frees it, so the entry of every deleted mirror and goto matchall filter
is leaked.
Free the entry after unlinking it.
The leak was discovered by an AI code review agent, and reproduced with
kmemleak on a lan969x EV board (EV23X71A) by repeatedly adding and
deleting matchall mirror and goto filters. With the fix, kmemleak no
longer reports the leak.
Cc: stable@vger.kernel.org
Fixes: 1ede4acf045c ("net: sparx5: add bookkeeping code for matchall rules")
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261007-sparx5-matchall-kfree-net-v1-1-c8918685b337@microchip.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
mlxsw_sp_flower_tmplt_create() parses a flow_cls_offload template
into a stack-local struct mlxsw_sp_acl_rule_info purely to compute
rulei.values.elusage. Parsing can acquire port range registers via
mlxsw_sp_flower_parse_ports_range(), but since this rulei never goes
through mlxsw_sp_acl_rulei_destroy(), those registers were never
released, including through chain template deletion.
Factor out of mlxsw_sp_acl_rulei_destroy() the code to actually
release the necessary resources and call from
mlxsw_sp_flower_tmplt_create() to plug the leak.
The issue was found during a review of Wentao Liang's patch referenced
below.
Fixes: fe22f7410527 ("mlxsw: spectrum_flower: Add ability to match on port ranges")
Reported-by: Wentao Liang <vulab@iscas.ac.cn>
Closes: https://lore.kernel.org/netdev/20260917113236.2149095-1-vulab@iscas.ac.cn/
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Petr Machata <petrm@nvidia.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/95339cf970ce78e1a74aa86ffcc104ecf7c7762c.1791294384.git.petrm@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The KSZ8765 is similar to the KSZ8795 but has two fiber ports.
KSZ8_PORT_STATUS_0 is used to detect fiber mode. It is currently
defined as 0x08, which is the Global Control 6 MIB Control register
and has bit 7 defined as Flush Counter.
Set KSZ8_PORT_STATUS_0 to 0x18, which is the Port 1 Status 0 register
where bit 7 is the Fiber Mode bit.
This issue was discovered when upgrading an embedded device from kernel
5.10 to 6.16. The device was using a KSZ8765 device tree configuration
and the corresponding hardware, but during boot the kernel incorrectly
detected it as a KSZ8795. The relevant boot messages were:
ksz-switch spi0.0: found switch: KSZ8795, rev 0
ksz-switch spi0.0: Device tree specifies chip KSZ8765 but found KSZ8795, please fix it!
Cc: stable@vger.kernel.org
Fixes: 91a98917a883 ("net: dsa: microchip: move switch chip_id detection to ksz_common")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Sebastien Royen <sebastien.royen@armadeus.com>
Link: https://patch.msgid.link/20261007080557.15719-1-sebastien.royen@armadeus.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
mlx5e_poll_ico_cq() requires sq->cc to be updated only after
mlx5_cqwq_update_db_record(), otherwise a CQ overrun may occur.
The current implementation updates sq->cc before the CQ doorbell
record, violating this ordering requirement.
Update the CQ doorbell record first and use dma_wmb() before updating
sq->cc. This ensures that the CQ space is released to the device
before the corresponding ICOSQ consumer index is updated by software.
Fixes: fd9b4be8002c ("net/mlx5e: RX, Support multiple outstanding UMR posts")
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261006105820.257208-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A veth device advertises NETDEV_XDP_ACT_NDO_XMIT only if its peer has an
XDP program attached or GRO enabled, that is, only if the peer will have
NAPI to receive the frames.
veth_set_features() updates the peer's flag when GRO is toggled, but
returns early if the device is down, and veth_open() only refreshes the
flags of the device being opened. Toggling GRO while the device is down
therefore leaves the peer's flag stale after the device comes up.
If GRO was enabled while down, the device comes up with NAPI but the
peer does not advertise NDO_XMIT, and devmap rejects redirects to the
peer with -EOPNOTSUPP. If GRO was disabled while down, the device comes
up without NAPI but the peer still advertises NDO_XMIT, so redirects are
accepted and then dropped in veth_xdp_xmit() with -ENXIO.
Commit 7a6102aa6df0 ("veth: Update XDP feature set when bringing up
device") made veth_open() refresh the device's own flags. Refresh the
peer's flags there too. The peer's flag depends on this device's XDP
program and GRO setting, not on whether the peer is up, so it is
correct to set it even if the peer is down.
Fixes: 8267fc71abb2 ("veth: take into account peer device for NETDEV_XDP_ACT_NDO_XMIT xdp_features flag")
Signed-off-by: Tianyi Gao <tianyi@cloudflare.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20261006173241.65945-2-tianyi@cloudflare.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers
instead of full pages to improve memory efficiency.") started handing
out RX buffers with zero headroom so that two buffers fit into one page
at the default MTU.
The MANA TX path, however, stores the per scatter-gather entry DMA
mappings in `struct mana_skb_head` at skb->head, and mana_start_xmit()
therefore calls skb_cow_head(skb, MANA_HEADROOM). The port advertises
this requirement as ndev->needed_headroom = MANA_HEADROOM.
As a result every packet that is received and then forwarded out of a
MANA port fails the skb_cow() in ip_forward() and gets reallocated and
copied by pskb_expand_head(). This is invisible to a plain RX or TX
workload, but it puts a full skb reallocation plus memcpy on the hot
path of every single forwarded packet, which is exactly what a
router/NVA workload does.
Restore the headroom. Note that reserving MANA_HEADROOM (232) is not
enough: ip_forward() asks for LL_RESERVED_SPACE(dev), which rounds
hard_header_len + needed_headroom up to HH_DATA_MOD and is 256 bytes on
ethernet. Use LL_RESERVED_SPACE() directly so the value keeps tracking
both constants. Also, since LL_RESERVED_SPACE() tracks MANA_HEADROOM,
it grows with MAX_SKB_FRAGS and for MAX_SKB_FRAGS >= 19 it is greater
than 256, so we have to account for that by using the headroom the RX
queue actually uses (instead of assuming XDP_PACKET_HEADROOM) and
turning MANA_XDP_MTU_MAX into MANA_XDP_MTU_MAX(ndev) (note that at the
default CONFIG_MAX_SKB_FRAGS=17 they are equivalent).
At the default MTU on a 4K page this means a buffer no longer fits twice
into a page (SKB_DATA_ALIGN(1500 + MANA_RXBUF_PAD + 256) = 2112), so the
frag-vs-single decision is now made by computing the real buffer size
instead of comparing the MTU against PAGE_SIZE / 2. The page_pool
fragment path is still used wherever at least two buffers genuinely fit,
e.g. on 16K and 64K page sizes.
Measured on an Azure VM with a MANA NIC acting as a forwarding NVA (UDP,
1400 byte payload, 4 streams, 8 Gbps offered, only the forwarding
node's kernel differs), 8 runs each, median:
forwarded pps throughput
before 272,830 3.06 Gbps
after 390,560 4.37 Gbps (+43%)
perf on the forwarding node, same workload:
memset_orig __pi_memcpy pskb_expand_head
before 10.07% 3.96% present
after 0.94% 0.64% gone
Cc: stable@vger.kernel.org
Fixes: 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers instead of full pages to improve memory efficiency.")
Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20261003013647.2051416-1-hamzamahfooz@linux.microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull power sequencing fixes from Bartosz Golaszewski:
- add missing PCI device IDs for Thinkpad T14s gen6 which should have
been part of commit a39ac4651e3b ("power: sequencing: pcie-m2: Match
WCN6855 and WCN7851 UART BT variants by subdevice ID") in
pwrseq-pcie-m2
- fix memory leak in pwrseq-pcie-m2 (leaking the array returned by
of_regulator_bulk_get_all())
* tag 'pwrseq-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
power: sequencing: pcie-m2: Fix leaking array from of_regulator_bulk_get_all()
power: sequencing: pcie-m2: Add Lenovo ThinkPad T14s gen6 WCN7850 subsystem PCI ids
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull gpio fixes from Bartosz Golaszewski:
- fix runtime PM leak in error path in gpio-xilinx
- fix race when arming the IRQ poll worker in gpio-mpsse
- fix devres cleanup path on probe error in gpio-exar
* tag 'gpio-fixes-for-v7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
gpio: mpsse: fix race when arming the IRQ poll worker
gpio: exar: initialize the ID before registering its cleanup
gpio: xilinx: fix runtime PM leak on request error path
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
- Update .mailmap entries for Andy Yan and John Garry
- Fix read-only MAP_SHARED /dev/zero mappings so they retain
shared-file semantics instead of being treated as anonymous memory,
also avoiding a CONFIG_DEBUG_VM assertion
- Fix 32-bit build warnings in the hugetlb-mmap selftest caused by
using the wrong printf format for size_t values
- Fix a boot-time crash when early function tracing causes CPA to free
kernel page tables before the workqueues used for deferred freeing
are available
- Fix two MREMAP_DONTUNMAP locked_vm accounting leaks: one caused by
an mlock-on-fault VMA self-merging, and one caused by partially
remapping a locked VMA
* tag 'mm-hotfixes-stable-2026-10-07-21-48' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
mailmap: update entry for Andy Yan
drivers/char/mem: mmap readonly MAP_SHARED-/dev/zero correctly
selftests/mm: cleanup -Wformat issues in hugetlb-mmap
mm: don't schedule deferred kernel page table freeing while booting
mailmap: update addresses for John Garry
mm/mremap: fix locked_vm leak by splitting VMA for MREMAP_DONTUNMAP
mm/mremap: fix locked_vm leak from MREMAP_DONTUNMAP self-merge
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc
Pull SoC fixes from Arnd Bergmann:
- Four distinct issues in TEE firmware, all fairly minor
- Five devicetree mistakes on NXP i.MX8, lx2160a and Qualcomm
based machines, one of these may cause file system corruption
from an incorrect SD card supply voltage
- Five fixes for clk drivers on new Qualcomm platforms,
addressing issues with incorrect enable states
* tag 'soc-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc:
MAINTAINERS: update my email address
tee: optee: ffa: support shared memory offsets on large-page kernels
optee: register TEE devices only once fully initialized
tee: shm: reject zero-sized allocations in tee_dyn_shm_alloc_helper()
arm64: dts: imx8mp-var-dart-sonata: Fix Sonata SD I/O supply
arm64: dts: lx2160a: fix the iic5 spi3 pinmux offset and value
arm64: dts: lx2160a: fix IIC1 pinmux submask rejected by pinctrl-single
arm64: dts: lx2160a: fix incorrect pinmux
clk: qcom: gpucc-kaanapali: Mark the GPU CX GDSC as votable
clk: qcom: gpucc-glymur: Mark the GPU CX GDSC as votable
clk: qcom: gcc-kaanapali: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-hawi: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-eliza: Fix always-enabling PCIE_RSCC clocks
arm64: dts: qcom: x1-denali: Fix microphone distortion
tee: qcomtee: fix kernel-doc warnings
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu
Pull tracing fix from Paul McKenney:
"Fix a double-dereference splat in TP_printk in usb mtu3. This was a
pre-existing bug that can corrupt tracing output, but which was
exposed this cycle by additional checking that was added to the
tracing subsystem.
This splat is reporting a double-dereference that can result in
garbage traces being dumped due to the possibility of the TP_printk()
being executed without the benefit of main memory being present.
The fix is to move the extra dereference from TP_printk() time to
TP_fast_assign() time"
* tag 'usb-mtu3.2026.10.06a' of git://git.kernel.org/pub/scm/linux/kernel/git/paulmck/linux-rcu:
usb: mtu3: Fix double dereference in TP_printk
|
|
bareudp_xmit_skb() sets the inner protocol of the skb to the configured
ethertype, and bareudp6_xmit_skb() does not set it at all. The inner
protocol is used by skb_udp_tunnel_segment() to segment the inner packet
when GSO has to be done in software, for instance when the lower device
does not offload the checksum.
With multiproto, an IPv4 device also carries IPv6, so the IPv6 header
is parsed as an IPv4 one. With IPv6 underlay, the inner protocol is
whatever the skb had before. In both cases segmentation fails and the
packets are dropped. With veth tx checksum offload disabled, iperf3 TCP
over bareudp gets about 10 Mbit/s with thousands of retransmits instead
of about 700 Mbit/s for IPv6 over IPv4 and for both families over IPv6,
while IPv4 over IPv4 works.
At this point skb->protocol is the protocol of the inner packet, which
bareudp_xmit() has already checked against the configuration, so use it
on both paths.
Fixes: 571912c69f0e ("net: UDP tunnel encapsulation module for tunnelling different protocols like MPLS, IP, NSH etc.")
Fixes: 4b5f67232d95 ("net: Special handling for IP & MPLS.")
Signed-off-by: Haishuang Yan <yanhaishuang@cmss.chinamobile.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20261002153039.462663-1-yanhaishuang@cmss.chinamobile.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Javier reports seeing packet loss issues related to the re-enablement
of K1; disabling it makes the issues go away. Add the reported system
to have K1 disabled by default.
Reproduction steps (from the Link):
Plug the cable and ping the default gateway on the LAN. No suspend/resume
involved, machine freshly booted and on AC power. Nothing e1000e-related in
dmesg apart from the link up/down messages; no "Hardware Unit Hang".
Default (K1 enabled):
$ ping -c 30 192.168.10.1
--- 192.168.10.1 ping statistics ---
30 packets transmitted, 18 received, 40% packet loss, time 29734ms
rtt min/avg/max/mdev = 0.211/0.315/0.384/0.046 ms
The lost packets are spread over the run (seq 6, 8, 9, 11, 13, 14, 16, 19,
20, 24, 25, 30), not a single burst. DHCP on this link also takes 1-2
minutes to get a lease, consistent with incoming packets being dropped.
With K1 disabled at runtime:
# ethtool --set-priv-flags enp0s31f6 disable-k1 on
$ ethtool --show-priv-flags enp0s31f6
Private flags for enp0s31f6:
s0ix-enabled: on
disable-k1 : on
$ ping -c 30 192.168.10.1
--- 192.168.10.1 ping statistics ---
30 packets transmitted, 30 received, 0% packet loss, time 29713ms
rtt min/avg/max/mdev = 0.169/0.251/0.468/0.054 ms
$ ping -c 30 192.168.10.1
--- 192.168.10.1 ping statistics ---
30 packets transmitted, 30 received, 0% packet loss, time 29677ms
rtt min/avg/max/mdev = 0.101/0.222/0.360/0.064 ms
I have only compile tested this, but Javier has tested that disabling K1
on his system (Lenovo ThinkPad P14s Gen 5 (Intel), I219-LM [8086:550a])
resolves his issues.
Cc: stable@vger.kernel.org
Fixes: 578294b8b60d ("e1000e: Reconfigure PLL clock gate timeout and re-enable K1 on Meteor Lake")
Reported-by: Javier Herrera <javier.herrera@afronta.com>
Link: https://lore.kernel.org/intel-wired-lan/CAKRuCQLWEcci9hPgNUmvq44Bd2H30T5_vP=exMjsknXMN5Yh9A@mail.gmail.com/
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
Link: https://patch.msgid.link/20261001222443.3500206-7-anthony.l.nguyen@intel.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
[fault reason 0x06] PTE Read access is not set
bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
...
NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
The fault address (0xfc499000) is on an exact 4KB page boundary,
pointing to a DMA read buffer overrun.
In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
such as 42-byte untagged ARP frames, are padded:
if (length < BNXT_MIN_PKT_SIZE) {
pad = BNXT_MIN_PKT_SIZE - length;
if (skb_pad(skb, pad))
goto tx_kick_pending;
length = BNXT_MIN_PKT_SIZE;
}
mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
...
dma_unmap_len_set(tx_buf, len, len);
However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
and is left unadjusted after padding. Consequently, dma_map_single() and
dma_unmap_len_set() map and track only 42 bytes.
Later, the hardware TX buffer descriptor is programmed with the padded length:
txbd->tx_bd_len_flags_type =
cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
TX_BD_FLAGS_PACKET_END);
The NIC DMA engine is thus instructed to read 52 bytes from a region where
only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
page (within 'pad' bytes of the next page), the hardware DMA read overruns
into the unmapped adjacent page, triggering an IOMMU fault.
This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
and races when toggling HW VLAN offload") because reserving extra VLAN
headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
offsets and potentially causing small frames to land right against page
boundaries.
Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in
bnxt_start_xmit(), before skb_shinfo(skb)->nr_frags is sampled,
because padding might linearize the skb.
This ensures skb->len and skb_headlen(skb) reflect the padded size so
that dma_map_single() maps the full buffer and the descriptor length is
consistent. This also removes the temporary 'pad' variable and masking logic.
The padding is done after the SW USO branch, otherwise
bnxt_sw_udp_gso_xmit() would account the padding as UDP payload
and send it in extra segments.
Note that SW USO segments are still not padded: with IPv4, a segment
carrying less than 10 bytes of UDP payload (small gso_size, or a short
last segment) is sent as a frame shorter than BNXT_MIN_PKT_SIZE.
This is a separate issue, left for a followup patch.
Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
Cc: stable@vger.kernel.org
Reported-by: Stefan Fleischmann <sfle@kth.se>
Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Michael Chan <michael.chan@broadcom.com>
Link: https://patch.msgid.link/20261007051747.455273-1-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Currently the driver zeroes the BARs only when fatal PCIe errors
are reported so that pci_restore_state() restores it. However
firmware handles both fatal and non-fatal errors the same way when
it sees the slot reset resulting from the PCI_ERS_RESULT_NEED_RESET
return code from the driver. This means that we must re-write the
BARs post recovery even during non-fatal errors. Otherwise we will
see that every MMIO access returns all-ones and the firmware appears
dead.
Zero-out the BARs during PCIe error recovery regardless of type of
PCIe error, and make the wait after the hot reset unconditional.
Disable memory decode and bus mastering before rewriting the BARs so
the device doesn't decode a half-updated address, bailing out if
config space is still inaccessible. Defer pci_enable_device() until
after the BAR rewrite and restore, so the device isn't re-enabled
while its BARs are still being rewritten, then re-enable the device
and re-assert bus mastering. Guard the same Command register cleanup
on the re-enable failure path against an inaccessible device.
Skip re-enabling the device in bnxt_io_slot_reset() if it is already
enabled, so enable_cnt does not go unbalanced. A concurrent
bnxt_fw_reset_task() can also be re-enabling the same device in its
ENABLE_DEV state, so guard that call the same way.
Fixes: f75d9a0aa967 ("bnxt_en: Re-write PCI BARs after PCI fatal error.")
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Reviewed-by: Scott Branden <scott.branden@broadcom.com>
Signed-off-by: Pavan Chebbi <pavan.chebbi@broadcom.com>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/20261005204246.3822563-4-michael.chan@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Fix and strengthen the FLR sequence when initializing in the kdump
kernel. If the NIC is behind a PCIe switch in synthetic (smart)
mode, the switch may need to see that the BARs have been initialized
before it will pass mem read/write TLPs to the NIC.
On a Dell system with a PEX89144 PCIe switch, echo c > /proc/sysrq-trigger
will trigger fatal AER without this patch:
bnxt_en 0000:67:00.0: enabling device (0000 -> 0002)
[Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 5
[Hardware Error]: event severity: recoverable
[Hardware Error]: Error 0, type: fatal
[Hardware Error]: section_type: PCIe error
[Hardware Error]: port_type: 5, upstream switch port
...
Add a new bnxt_kdump_reset() to do the expanded FLR sequence in the
kdump kernel. We now disable bus master and memory, save the PCI
state, do the FLR, clear the BARs, and restore the PCI state. The
BARs have to be cleared to ensure that they get re-initialized and
visible to the PCIe switch.
Since it is the kdump kernel, we make every effort to continue in
the best possible way even if pci_save_state() or pcie_flr() returns
error. After FLR, we poll for an additional 5 seconds before
aborting in case the device is not properly returning CRS. This is
similar to the 5-second wait in bnxt_io_slot_reset().
Fixes: 8743db4a9acf ("bnxt_en: Issue PCIe FLR in kdump kernel to cleanup pending DMAs.")
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Signed-off-by: Pavan Chebbi <pavan.chebbi@broadcom.com>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/20261005204246.3822563-3-michael.chan@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
In bnxt_io_slot_reset(), we clear the 6 BAR registers. Add a helper
function to do that. The helper will be used again in the next patch.
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Reviewed-by: Somnath Kotur <somnath.kotur@broadcom.com>
Reviewed-by: Joe Damato <joe@dama.to>
Signed-off-by: Pavan Chebbi <pavan.chebbi@broadcom.com>
Signed-off-by: Michael Chan <michael.chan@broadcom.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/20261005204246.3822563-2-michael.chan@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 4ce62d5b2f7a ("net: usb: ax88179_178a: stop lying about
skb->truesize") replaced skb_clone() with a copy into a new skb for
every packet of a bulk-in transfer except the last one. The copy skips
the 2-byte IP alignment pseudo header but still copies pkt_len bytes,
and pkt_len includes that header. Each of these frames is therefore
delivered with two extra trailing bytes: the 0xeeee pseudo header of
the next slot when the frame ends on an 8-byte boundary, or slot
padding otherwise. The last packet is still trimmed to pkt_len - 2.
IPv4 and IPv6 trim the extra bytes, so most traffic is unaffected.
Anything that relies on the frame length is not. MACsec takes the ICV
from the last 16 bytes and drops every such frame as InPktsNotValid. In
our setup that was 0.2-1.2% of received MACsec frames, depending on how
many frames shared a transfer, including service discovery multicast.
Allocate and copy pkt_len - 2 bytes, like the last-packet path.
Fixes: 4ce62d5b2f7a ("net: usb: ax88179_178a: stop lying about skb->truesize")
Cc: stable@vger.kernel.org
Signed-off-by: Fredrik Nyberg <fredrik.nyberg@volvocars.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20261002-ax88179-rx-len-v1-1-f01a2a295079@volvocars.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless
Johannes Berg says:
====================
Some more fixes:
- ath12k: revert broken panic handler
- mt76: SAR/MLO/STA/NAPI fixes
- nxpwifi: race fixes, validation
- libertas: revert botched frame size fix
- brcmfmac: fix crash on P2P removal
- mac80211: fix WARN on downlink OFDMA frames
* tag 'wireless-2026-10-07' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (25 commits)
Revert "wifi: libertas: reject short monitor TX frames"
wifi: brcmfmac: fix P2P device removal race in brcmf_detach()
wifi: mac80211: don't estimate airtime for unsupported rate widths
Revert "wifi: ath12k: add panic handler"
wifi: nxpwifi: do not delete Rx reorder entries under RCU
wifi: nxpwifi: fix inverted check in Tx BA stream entry deletion
wifi: nxpwifi: handle authentication frame allocation failures
wifi: nxpwifi: fix the authentication frame length handling
wifi: nxpwifi: zero the channel statistics array
wifi: nxpwifi: free the aggregation buffer when the RA list disappears
wifi: nxpwifi: wait for the wakeup timer before the adapter is freed
wifi: nxpwifi: delete the station entry on the uAP deauth event
wifi: nxpwifi: protect sta_list against concurrent add and delete
wifi: mt76: mt7996: fix uninitialized buf read in mt7996_variant_fem_init()
wifi: mt76: mt792x: pick the SAR table layout from the table itself
wifi: mt76: mt792x: treat 0xff ACPI SAR entries as no limit
wifi: mt76: mt7603: initialize global station WCID
wifi: mt76: mt7925: don't put the band auto marker in TGID
wifi: mt76: mt7915: disable rx napi when removing device
wifi: mt76: mt7996: fix struct mt7996_mcu_wed_rro_ba_delete_event layout
...
====================
Link: https://patch.msgid.link/20261007205143.464714-3-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This reverts commit 13ff543e0b2c713aedeaadadde686686e949dc78.
The 'goto free' just made things worse, causing a spinlock
to be handled incorrectly. Revert this for now.
Reported-by: Takashi Iwai <tiwai@suse.de>
Cc: stable@vger.kernel.org
Fixes: 13ff543e0b2c ("wifi: libertas: reject short monitor TX frames")
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
When the driver is removed while the user space process that created
the P2P device (e.g. wpa_supplicant) is still exiting, the P2P device
interface is removed twice and the kernel crashes.
The two paths are:
- wpa_supplicant, on exit, sends NL80211_CMD_DEL_INTERFACE for the P2P
device. nl80211_del_interface() holds RTNL and the wiphy mutex and
calls brcmf_p2p_del_vif(), which disables discovery and waits up to
1.5 s for BRCMF_E_IF_DEL. Then it calls brcmf_remove_interface(),
which unregisters the wdev and frees the vif and the ifp.
- brcmf_detach() sets the bus down, so BRCMF_E_IF_DEL never arrives,
and calls brcmf_remove_interface(ifp, false) for the same ifp.
brcmf_p2p_ifp_removed() reads ifp->vif and then blocks on
rtnl_lock(). When brcmf_p2p_del_vif() releases RTNL, it calls
cfg80211_unregister_wdev() on the wdev that was already unregistered
and freed.
The second list_del_rcu() of wdev->list faults on LIST_POISON2:
brcmfmac: brcmf_p2p_del_vif delete P2P vif
brcmfmac: brcmf_p2p_deinit_discovery enter
brcmfmac: brcmf_p2p_del_vif P2P: GO_NEG_PHASE status cleared
brcmfmac: brcmf_detach Enter
brcmfmac: brcmf_bus_change_state 1 -> 0
brcmfmac: brcmf_remove_interface Enter, bsscfgidx=1, ifidx=0
brcmfmac: brcmf_del_if Enter, bsscfgidx=1, ifidx=0
brcmfmac: brcmf_p2p_ifp_removed P2P: device interface removed
[1.5 s later, brcmf_p2p_del_vif() timed out]
brcmfmac: brcmf_remove_interface Enter, bsscfgidx=1, ifidx=0
brcmfmac: brcmf_del_if Enter, bsscfgidx=1, ifidx=0
brcmfmac: brcmf_p2p_ifp_removed P2P: device interface removed
Unable to handle kernel paging request at virtual address dead000000000122
Internal error: Oops: 96000044 [#1] PREEMPT_RT SMP
pc : _cfg80211_unregister_wdev+0x6c/0x264 [cfg80211]
Call trace:
_cfg80211_unregister_wdev+0x6c/0x264 [cfg80211]
cfg80211_unregister_wdev+0x10/0x1c [cfg80211]
brcmf_p2p_ifp_removed+0x84/0xb0 [brcmfmac]
brcmf_remove_interface+0x1f4/0x254 [brcmfmac]
brcmf_detach+0x88/0x180 [brcmfmac]
brcmf_sdio_remove+0x90/0x7b0 [brcmfmac]
brcmf_sdiod_remove+0x20/0xa0 [brcmfmac]
brcmf_ops_sdio_remove+0xbc/0x12c [brcmfmac]
sdio_bus_remove+0x38/0x144
device_release_driver_internal+0x1e8/0x2c0
device_release_driver+0x14/0x20
brcmf_sdio_bus_remove+0x2c/0x40 [brcmfmac]
brcmf_fwvid_unregister_vendor+0xd8/0x170 [brcmfmac]
__exit_compat+0x18/0x418 [brcmfmac_cyw]
__arm64_sys_delete_module+0x1ac/0x244
The log was taken with debug=0x406 on an i.MX 8M Plus board with a
CYW55513 (Sona IF513) on SDIO. The driver is the Ezurio backport of
brcmfmac from v6.18.22 on a 5.4-rt kernel. The code involved is the same
in mainline.
To reproduce:
1. Start wpa_supplicant on wlan0 and let it associate. It creates the
P2P device (iw dev shows "type P2P-device").
2. Stop wpa_supplicant and remove the driver without waiting for
wpa_supplicant to exit:
kill $(pidof wpa_supplicant)
rmmod brcmfmac_cyw brcmfmac
The crash is reproducible on that board.
For the interface without a netdev (the P2P device), take RTNL and the
wiphy mutex in brcmf_detach() before reading iflist, and remove it with
locked=true. If brcmf_p2p_del_vif() runs first, brcmf_detach() finds the
slot empty. If brcmf_detach() runs first, nl80211 does not find the wdev
any more. The interfaces with a netdev are removed as before, because
brcmf_del_if() takes RTNL on its own for the primary interface.
Fixes: 9831bcb987df ("brcmfmac: Deleting of p2p device is leaking memory.")
Cc: stable@vger.kernel.org
Assisted-by: claude-opus-5-5
Signed-off-by: Michele Dionisio <michele.dionisio@gmail.com>
Acked-by: Arend van Spriel <arend.vanspriel@broadcom.com>
Link: https://patch.msgid.link/20260924123027.4122909-2-michele.dionisio@gmail.com
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ath/ath
Jeff Johnson says:
==================
ath.git update for v7.3-rc7
In ath12k, revert the panic handler that landed in v6.11 since it can
perform a voluntary context switch within an RCU read-side critical
section.
==================
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
https://github.com/nxp-upstream/nxpwifi-next
Jeff Chen says:
===============
nxpwifi updates for 7.3
This update contains fixes for station list synchronization,
aggregation and reorder buffer cleanup, authentication frame
handling, and several memory safety issues.
Highlights:
- fix concurrent sta_list add/delete races
- clean up station entries on uAP deauthentication
- fix Tx BA stream entry deletion logic
- avoid deleting Rx reorder entries under RCU
- improve aggregation buffer cleanup
- fix authentication frame length validation
- handle authentication frame allocation failures
- fix wakeup timer shutdown ordering
===============
Signed-off-by: Johannes Berg <johannes.berg@intel.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/device-mapper/linux-dm
Pull device mapper fixes from Mikulas Patocka:
"dm-integrity:
- validate the superblock after re-reading it
- fix buffer overflow if tag size > 64
dm:
- fix reading up to 7 bytes beyond the end of block in dm-ioctl
- fix reading free memory if the ioctls are called concurrently
dm-crypt:
- fix a crash on invalid table line
dm-snap:
- fix a crash on invalid table line"
* tag 'for-7.3/dm-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/device-mapper/linux-dm:
dm-integrity: validate the superblock on resume
dm-snap: reject the transient exception store for snapshot-merge
dm: fix reading free memory in do_resume
dm ioctl: don't let data_start run past the buffer
dm-crypt: reject the lmk IV mode with AEAD ciphers
dm-integrity: fix buffer overflow in inline mode with large tag size
|
|
pfcp_encap_recv() only makes sure that the UDP header and the 4 byte
PFCP header are in the linear area. When the S flag is set,
pfcp_session_recv() then reads the 8 byte SEID that follows, which is
not covered by the pskb_may_pull() check.
A short PFCP packet with the S flag set, or one whose session header
lies in a fragment, therefore makes pfcp_session_recv() read beyond the
end of the packet data, and whatever it finds there is stored in the
tunnel metadata that flower later classifies on.
Pull up to the end of the SEID before reading it, and drop packets
that are too short to contain it. Reload the header pointer afterwards
since pskb_may_pull() may reallocate the skb head.
Use offsetofend() rather than sizeof(struct pfcphdr_session): the
structure is not packed, so its size is 16 bytes because of the
alignment of the __be64 member, while the header on the wire is only 12
bytes, and valid session messages would be dropped.
Fixes: 6dd514f48110 ("pfcp: always set pfcp metadata")
Signed-off-by: Haishuang Yan <yanhaishuang@cmss.chinamobile.com>
Link: https://patch.msgid.link/20260930124818.114224-3-yanhaishuang@cmss.chinamobile.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
When the revised suspend/resume sequence was introduced this led to an edge
case where the TX is left disabled in the following sequence.
1. phy link is down, so UMAC is held in reset and then network interface
is WoL enabled
2. Enter suspend, bcmgenet_wol_power_down_cfg() enables UMAC_RX since MAC
is in SW_RESET
4. Enter resume, UMAC_RX is left enabled. Since we only enable UMAC_TX
and UMAC_RX in SW_RESET. The UMAC_TX is never enabled again on link up.
The issue was hit in production and I verified the sequence keeps
the TX disabled in my standalone test. I tested on an internal
development board with bcm77122, but should be reproducible
on any genet HW that supports power management.
Fixes: 254f3239dd07 ("net: bcmgenet: revise suspend/resume")
Signed-off-by: Justin Chen <justin.chen@broadcom.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260930230634.3132729-1-justin.chen@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/mkl/linux-can
Marc Kleine-Budde says:
====================
pull-request: can 2026-10-01
The first patch is by zjamg and restores the skb header initialization
lost during the v7.0 release cycle.
Oliver Hartkopp contributes a patch for the CAN net layer to fix the
unique skb identifier regression under RPS, introduced in the v7.0
release cycle, which causes lost packets.
The last patch is by Ji-Ze Hong and fixes a struct size mismatch in
the f81604 CAN driver, which results in a TX starvation.
* tag 'linux-can-fixes-for-7.3-20261001' of git://git.kernel.org/pub/scm/linux/kernel/git/mkl/linux-can:
usb: f81604: fix struct f81604_int_data size mismatch
can: fix unique skb identifier regression under RPS
can: dev: init_can_skb(): restore skb header initialization
====================
Link: https://patch.msgid.link/20261001151905.1556270-1-mkl@pengutronix.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This reverts commit ad0ae7aefa7a ("net/mlx5: E-Switch, preserve max tx
speed on vport state modification").
mlx5_modify_vport_admin_state() and mlx5_esw_adj_vport_modify() query
the vport's current max_tx_speed before modifying its state, and
write that value back, to avoid resetting it to 0 as a side effect
of an unrelated admin-state change.
That's unnecessary: max_tx_speed is optional in MODIFY_VPORT_STATE -
FW skips writing it whenever it's 0, treating that as "not provided"
rather than "reset to zero". Leaving it unset already preserves FW's
current value, with no query needed.
Fixes: ad0ae7aefa7a ("net/mlx5: E-Switch, preserve max tx speed on vport state modification")
Signed-off-by: Or Har-Toov <ohartoov@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20261004083531.216988-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_sync: Fix command skb lifetime during scan setup
- hci_sock: Serialize dead-device detachment
- RFCOMM: free the skb when the DLC has no owner
- RFCOMM: connect the session socket without rfcomm_mutex
- RFCOMM: Fix NULL tty_dev dereference in rfcomm_dev_shutdown
- SCO: serialise sco_conn lifetime against sco_recv_scodata()
- ISO: Reject concurrent BIS listener setup
- ISO: Serialize concurrent connect calls
- MGMT: Fix status of pending commands flushed on power off
Driver:
- btintel_pcie: fix TX descriptor bounds check
- btintel_pcie: use managed IRQ teardown
- btintel_pcie: Add shared HW reset for clean start
* tag 'for-net-2026-10-05' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: ISO: Serialize concurrent connect calls
Bluetooth: ISO: Reject concurrent BIS listener setup
Bluetooth: hci_sync: Fix command skb lifetime during scan setup
Bluetooth: hci_sock: Serialize dead-device detachment
Bluetooth: RFCOMM: Fix NULL tty_dev dereference in rfcomm_dev_shutdown
Bluetooth: btintel_pcie: use managed IRQ teardown
Bluetooth: MGMT: Fix status of pending commands flushed on power off
Bluetooth: btintel_pcie: Add shared HW reset for clean start
Bluetooth: RFCOMM: connect the session socket without rfcomm_mutex
Bluetooth: btintel_pcie: fix TX descriptor bounds check
Bluetooth: SCO: serialise sco_conn lifetime against sco_recv_scodata()
Bluetooth: RFCOMM: free the skb when the DLC has no owner
====================
Link: https://patch.msgid.link/20261005194436.3594085-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
bc_transmit() sends the original skb through the last txable port.
When no port is txable, it returns false without consuming the skb, but
team_xmit() still returns NETDEV_TX_OK. AF_PACKET sends then retain the
skb and its socket write-memory charge.
A TX-enabled, link-down port keeps the broadcast transmit op installed.
Carrier can stay up through a second link-up, TX-disabled port or be
forced on by userspace. Free the original skb on the no-port path while
leaving the existing drop accounting intact.
Discovered by Sashiko during review of previous commit.
Cc: stable@vger.kernel.org
Fixes: 5fc889911a99 ("team: add broadcast mode")
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20261001182326.572045-3-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
A nonzero dev_queue_xmit() result does not return ownership of the skb.
When an override port's qdisc drops it, the queue override helper used to
try another port or the mode transmit op with the freed skb.
Return the handoff decision separately from transmit success. Select only
the first override port under RCU and stop after handing it the skb,
regardless of the lower transmit result. Keep the existing success and
drop accounting.
BUG: KASAN: slab-use-after-free in sk_skb_reason_drop
Write of size 4 by task poc/5040
sk_skb_reason_drop (net/core/skbuff.c:1240)
dev_kfree_skb_any_reason (net/core/dev.c:3470)
ab_transmit (drivers/net/team/team_mode_activebackup.c:49)
team_xmit (drivers/net/team/team_core.c:1849)
__dev_direct_xmit (net/core/dev.c:4934)
Freed by task 5040:
__tcf_kfree_skb_list (net/sched/sch_generic.c:43)
__dev_queue_xmit (net/core/dev.c:4831)
team_xmit (drivers/net/team/team_core.c:1846)
Kernel panic - not syncing: KASAN: panic_on_warn set ...
Cc: stable@vger.kernel.org
Fixes: 8ff5105a2b9d ("team: add support for queue override by setting queue_id for port")
Reported-by: <co+2cf741c3005abc20@bugs.sh>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20261001182326.572045-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Currently the base non-PTP-capable lan969x variants don't forward any
Ethernet frames:
$ ip link set eth10 up
$ ip addr add 10.0.0.54/24 dev eth10
$ ping 10.0.0.1
PING 10.0.0.1 (10.0.0.1): 56 data bytes
--- 10.0.0.1 ping statistics ---
7 packets transmitted, 0 packets received, 100% packet loss
$ cat /proc/interrupts | grep fdma
20: 0 GIC-0 120 Level sparx5-fdma
$ ip -s link show dev eth10
RX: bytes packets errors dropped missed mcast
20828 102 0 0 0 72
TX: bytes packets errors dropped carrier collsns
0 0 0 0 0 0
Testing showed that starting the domain 0 TOD counter gets networking
working again, but start all three for consistency with the PTP variants.
Fixes: 207966787b71 ("net: sparx5: add feature support")
Suggested-by: Daniel Machon <daniel.machon@microchip.com>
Signed-off-by: Quentin Freimanis <quentin@q-lab.dev>
Reviewed-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20261005052826.14251-1-quentin@q-lab.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Rather surprisingly, opening /dev/zero read-only then mmap()'ing it
MAP_SHARED gets you true anonymous memory (albeit in a VMA with non-NULL
vma->vm_file).
This is a by-product of MAP_PRIVATE-/dev/zero being how anonymous memory
was mapped in Linux's distant past.
It happens because mmap_zero_prepare() gates on VMA_SHARED_BIT and when
mapping a read-only file MAP_SHARED, do_mmap() clears VMA_SHARED_BIT and
VMA_MAYWRITE_BIT.
The gating is incorrect - the (poorly named) VMA_MAYSHARE_BIT flag exists
explicitly to tell you if something was originally mapped MAP_SHARED.
So the fix is simple - gate on this instead.
This isn't exactly a common use case, but it's unexpected behaviour which
now causes an assert if CONFIG_DEBUG_VM is set.
While this bug has existed since the dawn of time for linux (or at least
since 2.6.12), it hasn't caused issues in the past, so while it's
incorrect behaviour, it doesn't seem necessary to backport that far.
The mapping is now accounted at mmap time and can fail with -ENOMEM under
strict overcommit, and read faults allocate folios. However this is
normal behaviour for a read-only shmem mapping.
Commit 93c0c8dc87f6 ("mm/rmap: use anon pgoff to track MAP_PRIVATE
file-backed anon folios") is the first patch at which the debug assert
fires, so target that instead.
Link: https://lore.kernel.org/20260924-fix-dev-zero-readonly-shared-v1-1-153c2111e323@kernel.org
Fixes: 93c0c8dc87f6 ("mm/rmap: use anon pgoff to track MAP_PRIVATE file-backed anon folios")
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: <syzbot+c181d3198e98f8aef8b9@syzkaller.appspotmail.com>
Closes: https://lore.kernel.org/linux-mm/6ab4ae75.80e1c6cc.1e8e5f.000d.GAE@google.com/
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Jann Horn <jannh@google.com>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Lance Yang <lance.yang@linux.dev>
|
|
Commit 9e8db5913264 ("net: avoid false positives in untrusted gso
validation") added a '&& skb->network_header' check before flow-dissecting
GSO packets without VIRTIO_NET_HDR_F_NEEDS_CSUM in
__virtio_net_hdr_to_skb(), because some callers (such as tun_get_user(),
tun_xdp_one(), virtnet_receive_done(), and raw_verify_header()) called
virtio_net_hdr_*_to_skb() before initializing skb->network_header and
skb->dev.
Because __alloc_skb() and __build_skb_around() zero-initialize
skb->network_header to 0 (unlike mac_header and transport_header which
are initialized to ~0U), those four callers always had
skb->network_header == 0 and bypassed flow dissection in
__virtio_net_hdr_to_skb(). More generally, skb->network_header is an
offset from skb->head (where 0 is also a valid offset whenever
skb_headroom(skb) is 0), not a boolean flag.
Whenever the 'if (gso_type && skb->network_header)' branch was skipped,
the fallback 'else if (gso_type)' only pulled nh_min_len + thlen (40 bytes
for TCPv4) without dissecting the packet, without validating ip_proto or
n_proto, and without setting skb->transport_header.
If the packet has a malformed network header, it is not rejected and a
subsequent skb_probe_transport_header() also fails, leaving
skb->transport_header at ~0U (0xffff). Similarly, if an IPv4 packet
carries IP options (ihl > 5) or an IPv6 packet carries extension headers,
pulling only nh_min_len + thlen can leave the TCP header outside
skb->head. In both cases, tcp_hdrlen(skb) in skb_gso_transport_seglen()
reads out-of-bounds:
BUG: KASAN: slab-out-of-bounds in skb_gso_transport_seglen
Read of size 2 by task poc/133
skb_gso_transport_seglen (net/core/gso.c:155)
skb_gso_validate_mac_len (net/core/gso.c:270)
tbf_enqueue (net/sched/sch_tbf.c:260)
dev_qdisc_enqueue (net/core/dev.c:4227)
__dev_queue_xmit (net/core/dev.c:4884)
In addition, checking virtio_net_hdr_match_proto() only inside
'if (!skb->protocol)' before flow dissection both skipped validation when
skb->protocol was pre-set by the caller and rejected VLAN-tagged frames
whose outer L2 protocol is ETH_P_8021Q or ETH_P_8021AD.
Fix this by:
1. Initializing skb->dev and skb->network_header (plus skb->protocol for
IFF_TUN) before virtio_net_hdr_*_to_skb() in tun_get_user(),
tun_xdp_one(), virtnet_receive_done(), and raw_verify_header(). In
tun_get_user(), drop the redundant skb_reset_mac_header(skb) in the
IFF_TUN case since __virtio_net_hdr_to_skb() unconditionally resets
mac_header.
2. Removing '&& skb->network_header' and the unvalidated
'else if (gso_type)' fallback in __virtio_net_hdr_to_skb() so all GSO
packets without VIRTIO_NET_HDR_F_NEEDS_CSUM are flow-dissected, have
their transport header pulled into linear data, and have
skb->transport_header set.
3. Moving the virtio_net_hdr_match_proto() check to after
skb_flow_dissect_flow_keys_basic(), validating keys.basic.n_proto
against hdr_gso_type.
Fixes: 9e8db5913264 ("net: avoid false positives in untrusted gso validation")
Fixes: d5be7f632bad ("net: validate untrusted gso packets without csum offload")
Fixes: 924a9bc362a5 ("net: check if protocol extracted by virtio_net_hdr_set_proto is correct")
Reported-by: Weiming Shi <bestswngs@gmail.com>
Closes: https://lore.kernel.org/netdev/20260927163117.746432-2-bestswngs@gmail.com/
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Cc: Michael S. Tsirkin <mst@redhat.com>
Link: https://patch.msgid.link/20261001191140.2818991-3-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fix from Niklas Cassel:
- Set CHECK CONDITION for failed ATAPI commands.
Since commit 2e1d2e65e773 ("ata: libata-scsi: terminate deferred
commands on time out") failed ATAPI commands incorrectly stopped
having CHECK CONDITION set for commands that had a SCSI midlayer
byte set by scsi_check_sense(). SG_IO users therefore saw failed
ATAPI commands as successful (Hengyu)
* tag 'ata-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
ata: libata-scsi: do not lose CHECK CONDITION for failed ATAPI commands
|
|
ath12k_core_panic_handler() is invoked via atomic_notifier_call_chain(),
which runs inside an RCU read-side critical section. The current code calls
ath12k_pci_sw_reset() synchronously from this context, which eventually
reaches mhi_device_get_sync() and schedule_timeout(), triggering a voluntary
context switch within RCU.
Call trace:
rcu_note_context_switch+0x4c4/0x508 (P)
__schedule+0xbc/0x1204
schedule+0x34/0x110
schedule_timeout+0x84/0x11c
__mhi_device_get_sync+0x164/0x228 [mhi]
mhi_device_get_sync+0x1c/0x3c [mhi]
ath12k_wifi7_pci_bus_wake_up+0x20/0x2c [ath12k_wifi7]
ath12k_pci_read32+0x58/0x350 [ath12k]
ath12k_pci_clear_dbg_registers+0x28/0xb8 [ath12k]
ath12k_pci_panic_handler+0x20/0x44 [ath12k] ath12k_core_panic_handler+0x28/0x3c [ath12k]
notifier_call_chain+0x78/0x1c0
atomic_notifier_call_chain+0x3c/0x5c
Revert change "wifi: ath12k: add panic handler" to avoid this issue.
Tested-on: WCN7850 hw2.0 PCI WLAN.HMT.1.1.c7-00108-QCAHMTSWPL_V1.0_V2.0_SILICONZ_UPSTREAM-3
Signed-off-by: Yingying Tang <yingying.tang@oss.qualcomm.com>
Tested-by: Joonhoe Kim <26rote@gmail.com>
Reviewed-by: Rameshkumar Sundaram <rameshkumar.sundaram@oss.qualcomm.com>
Fixes: 809055628bce ("wifi: ath12k: add panic handler")
Link: https://patch.msgid.link/20260612032332.2278338-1-yingying.tang@oss.qualcomm.com
Signed-off-by: Jeff Johnson <jeff.johnson@oss.qualcomm.com>
|
|
The RCU guard is declared in the body of the per-TID loop, so it is still
held across the teardown pass. nxpwifi_del_rx_reorder_entry() flushes the
Rx workqueue and deletes the reorder timer synchronously, both of which
sleep, so a station deauthenticating from the AP splats under
CONFIG_DEBUG_ATOMIC_SLEEP.
Scope the guard to the collection walk. The teardown does not need RCU: it
serialises on priv->rx_reorder_tbl_lock[] and frees with kfree_rcu().
Fixes: 73b01e57ed3e ("wifi: nxp: add nxpwifi driver for IW61x")
Assisted-by: Claude:claude-opus-5
Signed-off-by: David Carlier <devnexen@gmail.com>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|
|
nxpwifi_is_tx_ba_stream_ptr_valid() returns true when the entry is still
linked, and every caller passes an entry that is on the list, so the early
return always fires and nothing is ever unlinked or freed. Entries leak on
every teardown and, since nxpwifi_space_avail_for_new_ba_stream() counts
them, Tx aggregation stops being negotiated once the stale count reaches
the maximum.
Changing the original dead && test to || to silence a NULL dereference
report inverted the validity test along with it.
Fixes: 00c786a7581e ("wifi: nxpwifi: fix multiple static analysis errors and warnings")
Assisted-by: Claude:claude-opus-5
Signed-off-by: David Carlier <devnexen@gmail.com>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|
|
nxpwifi_cfg80211_authenticate() allocates the buffer it builds the
management frame in and never checks the result:
mgmt = kzalloc(frame_len, GFP_KERNEL);
skb = dev_alloc_skb(...);
if (!skb) {
...
return -ENOMEM;
}
...
memcpy(mgmt->da, req->bss->bssid, ETH_ALEN);
On allocation failure the memcpy() a dozen lines later dereferences NULL.
The same sequence leaks mgmt when dev_alloc_skb() fails: that path
returns without freeing it, while the success path drops it after
nxpwifi_form_mgmt_frame().
Bail out when the allocation fails, and free it before returning on the
skb error path.
Fixes: 73b01e57ed3e ("wifi: nxp: add nxpwifi driver for IW61x")
Signed-off-by: Linmao Li <lilinmao@kylinos.cn>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|
|
nxpwifi_cfg80211_authenticate() builds a 3-address management frame in a
temporary buffer and hands it to nxpwifi_form_mgmt_frame(), which copies
it into the skb and splices in address4. The frame length is tracked in
a single u16 that conflates the two lengths, and it gets both wrong.
Truncation. req->ie_len and req->auth_data_len are both size_t. The IE
policy caps req->ie_len at IEEE80211_MAX_DATA_LEN, but
NL80211_ATTR_AUTH_DATA only has a four-byte minimum length policy. Since
nla_len is a u16, a single authentication data attribute can carry up to
65531 bytes, enough for the combined frame length to exceed U16_MAX
before it is assigned to pkt_len. With auth_data_len == 65510 and no IEs
the length wraps to 6, kzalloc(6) succeeds, and the memcpy() below then
writes 65506 user-provided bytes past a 6-byte heap object.
Geometry. NXPWIFI_MGMT_HEADER_LEN is 30, which is the 24-byte 3-address
header plus the address4 that the firmware expects. The temporary buffer
only holds a struct ieee80211_hdr_3addr, so six of the bytes accounted
for are never used, and passing that length to nxpwifi_form_mgmt_frame()
makes it append ETH_ALEN on top of a length that already included it.
The skb is sized without those six bytes, so the last skb_put_data()
overruns the tailroom by exactly ETH_ALEN. It usually goes unnoticed
because SKB_DATA_ALIGN() rounding leaves slack. The driver's other
caller of the helper, nxpwifi_cfg80211_mgmt_tx(), adds ETH_ALEN to the
skb size before allocating and passes the 3-address length, which is what
this path should do too.
Keep the two lengths apart. frame_len is the 3-address frame the driver
builds: it sizes the temporary buffer and is what the helper is told.
pkt_len is what the firmware sees, frame_len plus address4: it sizes the
skb and goes into tx_info. Computing frame_len in size_t and rejecting
anything that cannot be represented once ETH_ALEN is added removes the
wrap.
The value handed to the firmware is unchanged, so this is not a
behavioural change for frames that were already valid. A tighter 802.11
bound may make sense, but that depends on the firmware's frame format and
is a separate decision.
Fixes: 73b01e57ed3e ("wifi: nxp: add nxpwifi driver for IW61x")
Signed-off-by: Linmao Li <lilinmao@kylinos.cn>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|
|
adapter->chan_stats comes from vmalloc(), which does not clear the
memory, and nxpwifi_cfg80211_dump_survey() reads it as soon as user
space asks for survey data. For the entries no scan has filled in, a
non-zero cca_scan_dur passes the validity check and stale bytes are
reported to user space as noise, time and time_busy.
Use kcalloc() instead. The array holds two entries per supported
channel, small enough not to need vmalloc(), and the two callers change
to kfree() accordingly.
mwifiex fixed the same issue in commit 0e20450829ca ("wifi: mwifiex:
Initialize the chan_stats array to zero").
Fixes: 73b01e57ed3e ("wifi: nxp: add nxpwifi driver for IW61x")
Signed-off-by: Linmao Li <lilinmao@kylinos.cn>
Reviewed-by: Jeff Chen <jeff.chen_1@nxp.com>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|
|
nxpwifi_11n_aggregate_pkt() drops ra_list_spinlock while it copies each
subframe, so it rechecks the RA list after taking the lock again. The
check inside the aggregation loop returns without releasing the
skb_aggr it has been filling, leaking one tx_buf_size buffer along with
the subframes already aggregated into it.
Release skb_aggr there, the way the same check on the -EBUSY path
already does.
mwifiex fixed the same issue in commit 990a73dec3fd ("wifi: mwifiex: Fix
memory leak in mwifiex_11n_aggregate_pkt()").
Fixes: 73b01e57ed3e ("wifi: nxp: add nxpwifi driver for IW61x")
Signed-off-by: Linmao Li <lilinmao@kylinos.cn>
Reviewed-by: Jeff Chen <jeff.chen_1@nxp.com>
Signed-off-by: Jeff Chen <jeff.chen_1@nxp.com>
|