summaryrefslogtreecommitdiff
path: root/include/linux
AgeCommit message (Collapse)Author
2026-09-22packet: use ubuf_info completion for TX_RING packetsWillem de Bruijn
tpacket_snd sends skbs with frags pointing into its ring slots. Slots are released when skb->destructor is called. A call to skb_orphan calls skb->destructor before the skb is freed. This can cause the slot to be reused while still linked into the skb. Switch to standard zerocopy completion (ubuf_info) so the slot is only released once all references to the payload are freed or copied. Restore skb->destructor to standard sock_wfree. The ubuf_info completion callback can be called with a NULL skb, but only from net_zcopy_put and related API, used by zerocopy implementations that hold their own reference on the uarg, such as MSG_ZEROCOPY. This uarg is only ever completed from skb_zcopy_clear, so skb is always set. To prevent userspace from aliasing in-flight state on shared ring slots, allocate tpacket_uarg per packet, rather than per slot. This adds a small allocation to the transmit path. Use standard kmalloc to allow backporting to stable kernels. The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing the ring pages. Always allocate vec->deferred for tx_ring so page-backed rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs before calling tpacket_ubuf_complete. Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before accessing the slot. The sk_wmem_alloc reference now guarantees that the slot is valid. The test is also not sufficient by itself, as it reads pg_vec without pg_vec_lock, so it can race with packet_set_ring. As a result a slot is released when its payload is copied, which can be before transmission (e.g., in skb_orphan_frags_rx). If copied before skb_tx_timestamp() is called, no slot timestamp is recorded, similar to when skb_orphan() was called early in the datapath before this patch. Revert the now unused previous skb_zcopy_.._nouarg infra. Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING free until skbs finish"). Reported-by: Katherine Leaver <kleaver@janestreet.com> Reported-by: Bjoern Doebel <doebel@amazon.de> Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/ Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22bpf: Prevent variable arena/non-arena register contentsEmil Tsalapatis
The verifier marks ALU instructions that include at least one arena operand with needs_zext: These instructions are fixed up after verification to be ALU32 instructions to ensure that the result is a valid offset into an arena. However, different code paths may provide two non-arena 64-bit arguments to the same instruction. The result of the operation in that code path is wrong, since it is now unexpectedly truncated to 32 bits and zero-extended. Add logic to the verifier to ensure every instruction either always has at least one PTR_TO_ARENA argument, or never does. Since needs_zext already tracks the first scenario, add a prevent_zext field in bpf_insn_aux to track the latter. Reject instructions that use arena arguments and have prevent_zext set, or do not have arena arguments and have needs_zext set. Fixes: 6082b6c328b5 ("bpf: Recognize addr_space_cast instruction in the verifier.") Reported-by: Nicholas Carlini <nicholas@carlini.com> Suggested-by: Nicholas Carlini <nicholas@carlini.com> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260922172028.6269-8-emil@etsalapatis.com
2026-09-22bpf: Fix bounds check for skb-backed dynptrsEmil Tsalapatis
The skb_pointer_if_linear() function checks whether a memory region of length len starting at offset off into the skb is in the linear area, and returns a pointer to the region if so. The check currently subtracts between skb_headlen and offset of the check, and since skb_headlen is unsigned the subtraction can underflow. This causes the bounds check to spuriously pass and generate an arbitrary pointer of the form *(skb->data + off). The only user of this helper is currently skb-backed BPF dynptr code. Returning the wrong pointer leads to the dynptr erroneously being backed with invalid memory. Ensure the subtraction cannot underflow, and fail the check if it would. Use u64 arithmetic to also prevent overflow when calculating (skb_headlen(skb) - off) since off is unsigned. Fixes: 6f5a630d7c57 ("bpf, net: Introduce skb_pointer_if_linear().") Reported-by: Nicholas Carlini <nicholas@carlini.com> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev> Link: https://patch.msgid.link/20260922172028.6269-2-emil@etsalapatis.com
2026-09-22sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity ↵Davi Chaves Azevedo
underestimation bug The scheduler scales LLC capacity by the fraction of cache-sharing CPUs covered by a domain: llc_bytes = cache_size * span_weight / shared_weight During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The new domains therefore use the old sharing weight. The later call to sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has already been detached, and returns without correcting the surviving CPUs. On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC, offlining one SMT sibling left the remaining CPUs with: llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes The correct capacity is still 16777216 bytes. On systems with active cache-aware scheduling, an underestimated capacity can cause exceed_llc_capacity() to reject aggregation for a process whose footprint would fit. Unchanged cpuset partitions sharing the physical cache can also retain stale capacity when a CPU comes online in another partition. Pass the cache-sharing mask already retained by cacheinfo to the scheduler update. Refresh every surviving CPU using its own LLC domain so that each partition receives the correct share. This also preserves the correction needed as cache-sharing maps grow during boot. Keep the existing CPU-hotplug and scheduler-domain synchronization. The update remains on the hotplug path; no steady-state scheduling operation or persistent allocation is added. Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain") Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Chen Yu <yu.c.chen@intel.com> Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com> Tested-by: Chen Yu <yu.c.chen@intel.com> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Introduce task_struct->sched_cache_grp to fix UAFTim Chen
Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes the use-after-free when account_mm_sched() reaches the group through a task whose mm is being switched, as reported by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Keep the fork/exec/exit reference management out of the generic mm paths: add sched_cache_fork(), sched_cache_fork_cleanup(), sched_cache_exec_mmap() and sched_cache_exit_mm() in kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE), so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper instead of open-coding the refcounting under #ifdef. Also add sched_cache_group_get() and task_cache_group_get(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Decouple sched_cache_group from mm to fix UAFTim Chen
Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. This allows us in the next patch access sched_cache_group directly from task, and add a refcount on sched_cache_group when a task links to it. This prevents the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with a user defined grouping, or cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
2026-09-21ata: libata-core: Extend Samsung LPM quirk to AMD controllersNiklas Cassel
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA controller is reported to time out on STANDBY IMMEDIATE during system suspend with med_power_with_dipm enabled. The command completes when using max_performance instead. The existing Samsung LPM quirk only matches ATI controllers, leaving AMD controllers unaffected. Rename it to ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for the same Samsung SSD model patterns. Keep LPM behavior unchanged for other controller vendors, including Intel. Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD issue concerns LPM rather than NCQ. Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986 Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> --- Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-19vlan: require the MAC header to be present in __vlan_insert_inner_tag()Xiang Mei
__vlan_insert_inner_tag() only guarantees head room via skb_cow_head(), never that mac_len bytes of MAC header are present. Its ETH_HLEN wrappers - __vlan_insert_tag() under skb_vlan_push(), and vlan_insert_tag() under validate_xmit_vlan() on the generic transmit path - therefore rewrite the first 16 bytes at skb->data: a 12-byte memmove plus two 2-byte stores at +12 and +14. No caller supplies the bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull(). An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a hwaccel tag; the next - clsact "action vlan push" or bpf_skb_vlan_push() - enters the helper with skb->len still 1. The head comes from skbuff_small_head without __GFP_ZERO, so each push drags bytes from beyond skb->tail into the frame. After three the one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised slab: 0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81 `------------------------------' only 0x5a was sent; the rest is slab, here the top 56 bits of a linear-map address Require the MAC header the helper rewrites to be present, so such a frame is dropped rather than transmitted. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: co+0ea1ac045375cf05@bugs.sh Signed-off-by: Xiang Mei <xmei5@asu.edu> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-19bpf: Compare stack frames in regs_exact()Kumar Kartikeya Dwivedi
regs_exact() compares the register state up to id, followed by the ID mappings, but does not compare frameno. The PTR_TO_STACK case in regsafe() checks frameno separately, which is bypassed when exact comparison is requested. Consequently, infinite-loop detection can treat pointers to different stack frames as the same pointer and reject a finite loop. For example, initialize fp-8 to zero in the caller and to one in the callee, then pass the caller's fp-8 to the callee as r1: loop: r0 = *(u64 *)(r1 + 0); if r0 != 0 goto done; r1 = r10; r1 += -8; goto loop; done: exit; The loop terminates after reading the callee's slot on its second iteration. At the loop header, however, the only relevant difference is r1's frameno, so exact comparison incorrectly reports an infinite loop. The same problem occurs when the pointer is spilled to the stack. Move frameno into the type-specific metadata union, ahead of id, so the existing prefix comparison in regs_exact() covers it. Ordinary stack pointers do not use another union member. Iterator and IRQ stack-slot states use their dedicated union views and do not need a frame lookup. This also keeps bpf_reg_state at 80 bytes. Since frameno now shares storage with other pointer metadata, it is only meaningful for PTR_TO_STACK registers. Return NULL from bpf_func() for other register types. process_iter_arg(), get_constant_map_key() and is_dynptr_reg_valid_init() look up the frame before checking the register type and would otherwise index frame[] with a byte of the register's map or BTF pointer. They dereference the frame only after their type check. Move the states_maybe_looping() boundary from frameno to precise after the field relocation. Its prefix comparison continues to cover the complete value state and now includes frameno. Continue to ignore precise. Precision marks control whether pruning may ignore scalar ranges; they do not change the represented values, and exact comparison already compares those ranges unconditionally. Marks can also change through backtracking while an ancestor state is still being explored. Fixes: d5b892fd607a ("bpf: make infinite loop detection in is_state_visited() exact") Reported-by: Eduard Zingerman <eddyz87@gmail.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260919014213.1840880-2-memxor@gmail.com
2026-09-18net: ethtool: keep rtnl_lock for the ioctl self testAlexander Duyck
An offline self test that brings the interface down and back up with netif_close() / netif_open() requires rtnl_lock for both. Since the ethtool IOCTL path became rtnl-optional for ops-locked drivers, the ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an ops-locked driver the self test now tears the device down without rtnl_lock. With lockdep this reproduces deterministically on every offline self test on such a driver; note the sole lock held is the instance lock, not rtnl: WARNING: suspicious RCU usage net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage! 1 lock held by ethtool/107: #0: (&dev->lock){+.+.}, at: dev_ethtool Call Trace: netpoll_poll_disable __dev_close_many netif_close_many netif_close fbnic_self_test dev_ethtool_locked dev_ethtool dev_ioctl sock_ioctl __x64_sys_ioctl Without lockdep the same condition trips ASSERT_RTNL() in __dev_close_many() / __dev_open(); that check only samples the global rtnl state, so it can be masked by a concurrent rtnl holder, but the device is still being reconfigured without the lock it requires. The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST case is only needed on the ioctl path. Add an opt-in bit for drivers whose self test needs rtnl_lock and set it on the ops-locked drivers whose offline self test tears the interface down and up: - fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path uses netif_close() / netif_open(). - bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path goes through bnxt_close_nic() / bnxt_half_open_nic() / bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the device. Fixes: f994752b1127 ("net: ethtool: optionally skip rtnl_lock on IOCTL path") Signed-off-by: Alexander Duyck <alexanderduyck@fb.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18Merge tag 'drm-fixes-2026-09-19' of https://gitlab.freedesktop.org/drm/kernelLinus Torvalds
Pull drm fixes from Dave Airlie: "Things have picked back up a bit this week, mostly amdgpu, xe and msm this time. There are a bunch of scattered changes across the rest of drivers and core stuff, nouveau, i915. core: - fix vblank pending event leak ttm: - swapout fixes dma-buf: - scattergather fixes - enable dma-buf debug on debug kernels dma-fence: - fix signaling bit checks sched: - fix virtual runtime race msm: - DT: - Corrected indentation - Core: - Marked fbdev as system memory - GPU: - Fixed autosuspend cleanup on teardown - a750: fix timestamps - Increase GMU fw init timeout - Misc fixes/cleanups - DPU: - Fixed clock rounding, unbreaking newest platforms - Cleared pending flush state - DP: - Skip PUSH_IDLE when link was never enabled - Fixed bandwidth checks - HDMI: - Fixed runtime PM cleanup on probe failure xe: - shrinker related fixes - xe_mmio_gem fault handler and destroy fixes - xe disable i2c irq on unbind i915: - Revert a commit touching registers that don't necessarily exist - Check for negative numbers before passing to BIT() amdgpu: - SMU 14.x fix - DC IRQ fix - Runtime PM fix for P2P - RAS fix - PCIe reporting fix - DCN 6 fix - Device removal fix - DC MALL fix amdkfd: - GC 12.x fixes - Boundary checks - Mapping clear fix nouveau: - suspend/resume fixes gud: - out of bounds access fix - ignore damage clips in full update vc4: - use-after-free fix versilicon: - plane format fix longsoon: - blend mode property fix" * tag 'drm-fixes-2026-09-19' of https://gitlab.freedesktop.org/drm/kernel: (59 commits) drm/amd/display: fix MALL hysteresis timer underflow at high refresh rates drm/amdgpu: fix rmmio iounmap skipped on device removal drm/amdgpu: Skip KFD mapping clear before initialization drm/amd/display: Fix NULL dereference in dcn50/dcn60 init_hw drm/amdkfd: Avoid integer underflow in EOP ring size calculation. drm/amdkfd: Avoid integer underflow with ffs in EOP ring size calc drm/amdgpu: Fix GPU PCIe link capability reporting drm/amdgpu: check ras and obj before dereference drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments drm/amdkfd: implement restore_mqd callbacks for GFX12/12.1 drm/amd/display: Atomize IRQ register read/modify/write ops drm/amd/pm: report energy accumulator for smu 14.0.3 drm/loongson: Create blend mode property for cursor plane drm/xe/i2c: Disable IRQ on unbind Revert "drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable" drm/verisilicon: remove ARGB formats from primary plane drm/verisilicon: add primary modifier for format tables drm/verisilicon: set blend mode for the cursor plane drm/sched: Fix virtual runtime race drm/i915/display: check configuration index before shifting ...
2026-09-17bpf: Assign lock identity to callback map valuesKumar Kartikeya Dwivedi
A nested bpf_for_each_map_elem() callback can unlock a different element of the same map: static long inner(void *map, int *key, struct value *v, struct value **outer_value) { bpf_spin_lock(&v->lock); bpf_spin_unlock(&(*outer_value)->lock); return 0; } static long outer(void *map, int *key, struct value *v, void *ctx) { bpf_for_each_map_elem(map, inner, &v, 0); return 0; } Both callback values currently have ID zero and the same map_ptr. process_spin_lock() compares those two fields, so it accepts the unlock even though the two callbacks can receive different map elements. Assign a fresh ID to every callback map value in the for-each, timer/workqueue, and task-work constructors. Copies of one callback argument retain its ID, so locking and unlocking through that argument continues to work. Distinct callbacks also get distinct IDs for single-element arrays, including inner arrays sharing inner_map_meta. Preserve map_uid for every inner-map lookup and compare it through check_ids() during state pruning. This preserves relationships between maps, keys, and values while allowing equivalent states with different lookup IDs to match. It avoids field-specific rules for when an inner map needs an identity. Move map_uid out of the metadata union and next to the other IDs, so register comparisons can use the existing memcmp() ranges and remap the IDs separately. Clear it when resetting a register or converting a map lookup result to a socket pointer. Shrink frameno to u8, which is enough for MAX_CALL_FRAMES, to make room without growing bpf_reg_state. Fixes: d0d78c1df9b1 ("bpf: Allow locking bpf_spin_lock global variables") Reported-by: Nicholas Carlini <npc@anthropic.com> Suggested-by: Nicholas Carlini <npc@anthropic.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260917233222.2542500-9-memxor@gmail.com Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17bpf: Apply CO-RE relocations before subprogram validationKumar Kartikeya Dwivedi
check_subprogs() verifies that each subprogram ends in an exit or an unconditional jump before in-kernel CO-RE relocations are applied. An unresolved relocation can then replace that terminal instruction with an invalid helper call. The resulting fall-through into another subprogram breaks the CFG invariant used by postorder and stack liveness analysis, which can write past their per-subprogram arrays. Apply CO-RE relocations immediately after preparing the program BTF, before subprogram discovery and validation. Keep func_info and line_info validation after subprogram discovery because those records depend on the complete subprogram layout. Reject an ldimm64 first slot at the end of the instruction stream before CO-RE can inspect its missing second slot. check_subprogs() previously rejected this form before relocation processing because it is not a valid subprogram terminator. Moving CO-RE ahead of check_subprogs() removes that implicit protection, so perform an explicit check before applying relocations. Include core_relo_cnt when deciding whether to prepare program BTF. A load that supplied only CO-RE relocation metadata previously skipped both BTF setup and relocation processing. Fixes: fbd94c7afcf9 ("bpf: Pass a set of bpf_core_relo-s to prog_load command.") Suggested-by: Andrii Nakryiko <andrii@kernel.org> Suggested-by: Alexei Starovoitov <ast@kernel.org> Suggested-by: Eduard Zingerman <eddyz87@gmail.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260917233222.2542500-5-memxor@gmail.com Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17bpf: Verify global subprogs in each sleepability contextKumar Kartikeya Dwivedi
Global subprograms are verified independently with a fresh verifier root. do_check_common() currently seeds that root's in_sleepable state from the program, even though a global subprogram can also run from callbacks whose execution context differs from the program's main entry point. In particular, workqueue and task-work callbacks are sleepable even when the containing program is not. A global subprogram of that program is therefore verified as non-sleepable, making in_rcu_cs() true and allowing loads of RCU-protected kptrs to produce trusted MEM_RCU pointers. The same subprogram can then be called from a sleepable callback without a classic RCU reader. It can retain such a pointer while the object is freed and use it after free. The verifier's execution-context predicates are complementary. A state is sleepable only when in_sleepable is set and no RCU, preemption, IRQ, or lock region is active. Each condition which prevents sleeping also provides RCU protection, while in_rcu_cs() treats a non-sleepable state as implicitly protected. Use this relationship to represent a global subprogram caller with only the result of in_sleepable_context(). A protected sleepable caller is normalized to in_sleepable=false at the independent verification root. This both prevents sleepable operations and makes in_rcu_cs() true without copying caller-owned lock state. Track only the contexts in which each global subprogram is actually reached. Verify it once if all reachable calls use the same context, and twice only if both sleepable and non-sleepable calls reach it. Calls found while verifying globals or asynchronous callbacks mark further contexts for checking. Repeat the existing subprogram walk until all called contexts have been verified; unreachable global calls remain unchecked. Accumulate instruction counts over those verification passes. Preserve the total recorded before each pass, since path accounting has already added this pass's synchronous instructions and its root total must also include asynchronous subprograms. This makes an unprotected callback verify the global subprogram as sleepable, turning its RCU-protected kptr load into an untrusted pointer. Protected callers and global subprograms which do not depend on implicit RCU protection remain valid. Fixes: 81f1d7a583fa ("bpf: wq: add bpf_wq_set_callback_impl") Fixes: 38aa7003e369 ("bpf: task work scheduling kfuncs") Reported-by: Nicholas Carlini <npc@anthropic.com> Suggested-by: Nicholas Carlini <npc@anthropic.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Acked-by: Eduard Zingerman <eddyz87@gmail.com> Link: https://patch.msgid.link/20260914131923.2544250-2-memxor@gmail.com Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17Merge tag 'net-7.3-rc4' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Paolo Abeni: "Including fixes from Netfilter, Bluetooth, IPSec and WiFi. Previous releases - regressions: - netfilter: hold reference on ct until flow is released - bridge: - move switchdev call outside rcu - vlan: fix bugs caused by switchdev deletion errors - wifi: - mac80211: reset state when starting AP fails - cfg80211: don't free driver-owned scan requests - tcp: don't call skb_clone_and_charge_r() for close()d listener in tcp_v6_do_rcv() - mptcp: return sk_wait_data() errors from recvmsg() - xfrm: serialize state GC with device state flush - drop_monitor: synchronize tracepoint unregistration on error path - bluetooth: - eir: validate service data length before reading UUID - hci_sync: serialize local codec list cleanup - RFCOMM: avoid socket lock inversion in listener cleanup - eth: - lan743x: fix RX checksum use-after-free - mvpp2: prevent buffer overflow in page_pool allocation Previous releases - always broken: - core: lock the socket in sock_gettstamp() - neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS. - sched: codel: bound the dropping loop per dequeue call - wifi: mac80211: include TIM bitmap control for buffered S1G mcast traffic - psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb() - xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk() - bluetooth: hci_qca: do not write to the serial port after it is closed - dsa: mxl862xx: disable the stats poll on teardown - eth: - stmmac: fix TSO header length truncation - ip_tunnel: initialize `options_len` before referencing options" * tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (159 commits) mptcp: fix bad accounting in __mptcp_subflow_push_pending() mptcp: close race between scheduler and state change mptcp: avoid unneeded actions on subflow reset net: skbuff: do not leave stale header offsets after pskb_carve() selftests: net: packetdrill: test exclusion of old ACK from TCP fast path tcp: exclude old ACKs from tcp fast path dpll: reject a reference sync pin which is not on the pin's dpll net: mvpp2: prevent buffer overflow in page_pool allocation net: macb: fix ordering around PTP timestamp read selftests: drv-net: psp: test PSP and TCP ULP mutual exclusion net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb() net: stmmac: preserve real_num_tx_queues on mqprio setup failure net: stmmac: propagate FPE preemption-class mapping errors net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb() net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain net: ethernet: cortina: Ack RX overrun interrupt correctly net: lock the socket in sock_gettstamp() eth: fbnic: ring the doorbell if a burst ends in a drop net: netsec: fix device_node reference leak on phy_np ...
2026-09-16Merge tag 'wireless-2026-09-16' of ↵Jakub Kicinski
https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless Johannes Berg says: ==================== Many fixes: - mac80211: S1G TIM bitmap fix - ath12k: remove undocumented DT ABI implementation - various firmware API and over-the-air hardening changes - fixes for most cfg80211/mac80211 syzbot reports * tag 'wireless-2026-09-16' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (67 commits) wifi: brcmsmac: fix UAF in brcms_free_timer() wifi: brcmfmac: fix lost 802.1x TX completion wakeup wifi: ath11k: cleanup arsta in ath11k_mac_peer_cleanup_all() wifi: wcn36xx: Fix potential use-after-free in TX ack timer teardown wifi: ath12k: ahb: Revert undocumented ABI and dead code wifi: mac80211: refuse to make a monitor active when it has no queue wifi: libipw: reject TKIP frames without a full MIC wifi: virt_wifi: don't transfer operstate before register wifi: cfg80211: check if AP has been started or joined a mesh before adding new station wifi: cfg80211: move link_id validation earlier in nl80211_new_station() wifi: cfg80211: do not support direct add of station to AP_VLAN interfaces wifi: cfg80211: verify if AP_VLAN belongs to the correct AP wifi: mac80211: set up the TX info early to fix failure paths wifi: mac80211: mesh: release the channel if start fails wifi: mac80211: mesh: reset the CSA state when leaving wifi: mac80211: add HE 6 GHz capability in the scan elems len wifi: mac80211: don't access the TSF of a down interface wifi: mac80211: don't RCU-dereference the mesh CSA settings we just set wifi: mac80211: don't allow link changes when iface is down wifi: mac80211: require a peer station for TDLS setup confirm ... ==================== Link: https://patch.msgid.link/20260916083642.110609-3-johannes@sipsolutions.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16siphash: clean up kernel-doc commentsRandy Dunlap
Use the correct function parameter names and add function return value descriptions to avoid kernel-doc warnings: Warning: include/linux/siphash.h:82 function parameter 'len' not described in 'siphash' Warning: include/linux/siphash.h:82 No description found for return value of 'siphash' Warning: include/linux/siphash.h:132 function parameter 'len' not described in 'hsiphash' Warning: include/linux/siphash.h:132 No description found for return value of 'hsiphash' Signed-off-by: Randy Dunlap <rdunlap@infradead.org> Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
2026-09-16Merge fdo/drm/drm-fixes into drm-misc-fixesMaxime Ripard
Backmerging to get drm-misc-fixes up to v7.3-rc3. Signed-off-by: Maxime Ripard <mripard@kernel.org>
2026-09-15dma-buf/dma-fence: fix checking signaling bit for timeline and driver name v3Christian König
The patch "dma-buf: dma-fence: Fix potential NULL pointer dereference" changed the check to test for the ops pointer instead of the signaled bit to avoid a potential NULL dereference when the ops pointer has been cleared. The problem is now that the ops pointer is cleared only when neither the release nor the wait callback is implemented and this isn't true for a lot of dma_fence implementations yet. So those implementations lost the RCU protection after signaling of the returned string resulting in potential use after free. Add the signaling check additional to the ops pointer check so that we have both the protection against NULL dereference as well as the RCU protection after signaling for the returned string. v2: improve comments to note RCU protection and explain why we check both signaling state and ops pointer v3: some comment improvements suggested by Philip Signed-off-by: Christian König <christian.koenig@amd.com> Fixes: 035219a760ed ("dma-buf: dma-fence: Fix potential NULL pointer dereference") CC: stable@vger.kernel.org # 7.2+ Reported-by: Jonghyuk Kim(MalHyuk) <malhyuk97@gmail.com> Tested-by: Jonghyuk Kim(MalHyuk) <malhyuk97@gmail.com> Reviewed-by: Philipp Stanner <phasta@kernel.org> Link: https://lore.kernel.org/r/20260914182740.1587-1-christian.koenig@amd.com
2026-09-13Merge tag 'x86_urgent_for_7.3-rc4' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Dave Hansen: "The most notable fix is THP not silently losing user data and having been around for a couple of years. The main explanation I'd have for its longevity is that it requires a few different things to align at the same time: MADV_FREE, THP and heavy reclaim. - Fix user-space data loss with THP - Fix set_memory oopses - Fix addition of large constants in mul_u64_add_u64_div_u64() - Fix FineIBT hash offset in cfi_get_func_hash() - Fix PCI device reference counting in amd_smn_init()" * tag 'x86_urgent_for_7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/amd_node: Fix PCI device reference counting in amd_smn_init() x86/div64: Fix addition of large constants in mul_u64_add_u64_div_u64() x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash() x86/mm: Fix user-space data loss with MADV_FREE and THP x86/mm/pat: Allocate split page tables as kernel page tables x86/alternatives: Exclude text poking against change_page_attr() x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
2026-09-13Merge tag 'trace-v7.3-rc2' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull tracing fixes from Steven Rostedt: - Don't destroy user event fields when removal fails User event fields are destroyed before the event is removed from visibility. But that can fail leaving the still visible event with no fields. Move the destroying of the fields to after the event is successfully removed from visibility. - Initialize function graph state is fork before calling copy_exec_state() For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state. If that allocation fails, ftrace_graph_exit_task() will free the tasks ret_stack pointer. Since that pointer is still using the parent's ret_stack, it mistakenly frees the parent's pointer too. Call ftrace_graph_init() on the task first which will NULL out the new tasks's ret_stack and if the copy fails, it will not free anything. - Remove FGRAPH_MAX_INDEX The macro FGRAPH_MAX_INDEX was added but never used. Remove it. - Save ent_size in function graph printing of nested functions The function graph tracer needs to look at the next event to see if the next event is the return of the current function entry. If it is, it prints a single line: ktime_get(); Otherwise it prints it like a nested function: tick_nohz_irq_exit() { ktime_get(); kcpustat_irq_exit(); } In order to look at the next event, it must save the current event so that it has the information to print from it. It saves the event in the iterator descriptor called "ent". What it doesn't save is the ent_size of the event which is now used to know if the function graph arguments are to be printed. The peek doesn't save the size so the size used happens to be that of the size of the last event that was seen. Save the entry event size in the iterator descriptor so that the correct size is used. - Fix several errors with freeing data in the histogram code The histogram code had a lot of leaked or or incorrect accounting when failures happen. Correct them. - Fix histogram regression of .percent and .graph modifiers Up until 6.3 histogram values could have "percent" or "graph" modifiers that changed how they were printed. But a change that added restricting histograms values from being strings, stack traces and other modifiers inadvertently prevented them from using the percent and graph modifiers, which were legal use cases for values. Put back the percent and graph modifiers. - Fix various typos in the comments - Set the trace_clock before initializing a histogram with clock argument The histogram API allows the user to specific which trace clock to use via a "clock=" string. The histogram is set up first before the clock is checked. If the passed in clock is not valid, it exits without fully fixing up the histogram leaving it on the list and a use-after-free can trigger. Update the clock argument first and if it fails then exit gracefully before the histogram trigger is placed on any lists. - Restore :mod: trailer after parsing in ftrace_set_clr_event The function ftrace_set_clr_event() modifies the parse string and needs to put it back to what was passed in. It searches for ":mod:" via a strsep() but fails to put back the first ':' in the string. Add back the ':' in the passed in string. - Take trace_array reference when opening a tracer options file The options files are dynamically created and some tracers add their own options. When a tracer adds their own list of options, the trace_array holding them has an array to hold the list of options for each tracer. This array increases in size via a krealloc(), and the new entry gets a newly allocated array to hold the options of the new tracer being added. The element in each entry of the tracer's option array holds a pointer back to the trace_array, a pointer to the tracer it is associated to, a pointer to the flags of the option. The issue is that these arrays are freed when the trace_array is freed when its instance it represents is removed from the instances directory. There's a race that an open of one of these options files can happen when the instance is being removed. Add a new helper function to be called by the open function of the options file to iterate all existing trace_arrays under a lock and find the one that has the given option element in one of it's tracer arrays. If found, then update the associated trace_array's reference counter to keep it from being freed. If not found, have the open call return -ENODEV. - Disable interrupts when acquiring the lock in rb_wake_up_waiters() The function rb_wake_up_waiters() assumes it will be called in interrupt context and does not disable irqs when taking cpu_buffer->reader_lock, which can be called in hard interrupt context. The issue is in PREEMPT_RT, this function is called in thread context leaving this lock open to a deadlock. Take the lock with interrupts disabled. - Use rcu_assign_pointer() for tmp_ops filter hash The tmp_ops used in update_ftrace_direct_mod() assigns its filter_hash field directly, but that field is annotated as __rcu and sparse complains. Assign it with rcu_assign_pointer() - Fix use-after-free in enable_trigger_private_data_free() The trace_event_call is accessed through the event_trigger_data's trace_event_file pointer to put the trace_event_call on freeing. The issue is that the trace_event_file data may have been freed already causing a use-after-free. Add a field to the event_trigger_data that points directly to the trace_event_call so that it can decrement its reference directly without needing to go through the trace_event_file. - Fix accounting of buffer data remote headers trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the number of pages is needed for the asked for size as it doesn't take into account the meta data on each page. Add a helper function to do the calculation properly and use that in these functions. - Catch nr_page_va overflow in ring_buffer_desc sizing The number of pages per remote ring buffer is capped by ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to overflow that field would silently allocate a descriptor smaller than what was asked for. - Do not resize the subbuf order if any per_cpu buffer is disabled The mmapping of ring buffers disables resizing the subbuffers, but it is done per-cpu whereas the subbuf size change is done for all the per_cpu buffers under the buffer->mutex. It could change the size of some while the mapping is happening on others. Have the resize of the subbuf order check all the per_cpu buffers under the lock to see if any of them is disabled before starting and causing an inconsistency between buffers that are being mapped. * tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: (25 commits) ring-buffer: Check resize_disabled before publishing the new subbuf order tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing tracing/remotes: Account for ring buffer page header in size calculation tracing: Don't dereference trace_event_file in deferred trigger free ftrace: Use rcu_assign_pointer() for tmp_ops filter hash ring-buffer: Acquire the lock with irqsave in rb_wake_up_waiters() tracing: Take trace_array reference when opening a tracer options file tracing: Fix ring_buffer_read_page_size() kernel-doc tracing: Restore :mod: trailer after parsing in ftrace_set_clr_event() tracing: Fix memory corruption from a "STACKTRACE" histogram key tracing: Fix memory corruption from the histogram stacktrace modifier tracing: Undo the registration when enabling the histogram trigger fails tracing: Take the reference before publishing the named histogram trigger tracing: Set the trace clock before registering the histogram trigger tracing: Fix typo "preceeded" in comment tracing: Fix typo "availabe" in comment tracing: Let histogram values keep the percent and graph modifiers tracing: Keep the entry count when the histogram stats allocation fails tracing: Free histogram the field rejected for a bad modifier tracing: Free histogram the var ref when its initialization fails ...
2026-09-13tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizingVincent Donnefort
The number of pages per remote ring buffer is capped by ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to overflow that field would silently allocate a descriptor smaller than what was asked for. Return SIZE_MAX from trace_buffer_desc_size() on nr_page_va overflow. Link: https://patch.msgid.link/20260911193937.602202-3-vdonnefort@google.com Fixes: 2e67fabd8b77 ("ring-buffer: Introduce ring-buffer remotes") Signed-off-by: Vincent Donnefort <vdonnefort@google.com> Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-13tracing/remotes: Account for ring buffer page header in size calculationVincent Donnefort
trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the required pages because every ring buffer page contains a header (BUF_PAGE_HDR_SIZE). Account for that header to ensure allocated remote ring buffers aren't smaller than requested by the user. The newly introduced helper __calc_nr_pages_ring_buffer_desc() can return a value that overflows the descriptor nr_pages field (32 bits). Link: https://patch.msgid.link/20260911193937.602202-2-vdonnefort@google.com Fixes: 2e67fabd8b77 ("ring-buffer: Introduce ring-buffer remotes") Signed-off-by: Vincent Donnefort <vdonnefort@google.com> Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-13Merge tag 'sched-urgent-2026-09-13' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull scheduler fixes from Ingo Molnar: - Fix EEVDF se->max_slice value on enqueueing (Vincent Guittot) - Fix EEVDF augmented rb-trees re-balancing with multiple fields (Vincent Guittot) - In proxy scheduling, account cgroup CPU time to the execution context, not the scheduling context (Hui Su) - Likewise, call wq_worker_tick() for the execution context, not the scheduling context (Hui Su) * tag 'sched-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: sched/core: Call wq_worker_tick() for the execution context sched: Account cgroup CPU time to the execution context sched/eevdf: Fix rb augmented with multi fields sched/eevdf: Fix augmented max_slice
2026-09-13Merge tag 'core-urgent-2026-09-13' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull entry code fix from Ingo Molnar: - Fix generic entry code cross-build failure on !CONFIG_AUDITSYSCALL kernels using older RISCV64 and S390 cross-compilers (Thomas Gleixner) * tag 'core-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: entry: Guard syscall_enter_audit() invocation with CONFIG_AUDITSYSCALL
2026-09-11Merge tag 'riscv-for-linus-7.3-rc3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/riscv/linux Pull RISC-V fixes from Paul Walmsley: "From a RISC-V point of view, there's one notable fix here, reverting an earlier bogus fix to the pointer masking code. Fortunately the practical impact appears to be small. - Revert a bad fix, likely LLM-generated, in the pointer masking code that confused the RISC-V hardware pointer masking implementation with the Linux kernel tagged address feature - Fix unexpected faults caused by kprobe instruction slot writes when !CONFIG_STRICT_MODULE_RWX - Fix unexpected faults on minimal configurations during runtime code patching on !CONFIG_STRICT_MODULE_RWX systems - Fix a misplaced variable clear causing incorrect reuse of previous values in the RISC-V hardware feature probing code - Fix two bugs in the PMU SBI perf code on rv32: use BIT_ULL rather than BIT on 64-bit masks; and use a bitmap rather than an unsigned long on a quantity that can exceed 32 bits And a few miscellaneous cleanups: - Avoid a potential dereference-before-NULL-pointer-check bug in the PMU SBI perf driver - Use CONFIG_GENERIC_BUG_RELATIVE_POINTERS to simplify the rv32 bug table code (like x86 and PPC) - Report the RISC-V standard ISA extensions Z[v]fhmin when support is claimed for the superset RISC-V standard ISA extensions Z[v]fh; and simplify our FPU test code to only check for the presence of the D extension - Use an existing kernel string helper in place of some open-coded code in kernel/usercfi.c - Fix some yamllint issues in the RISC-V DT bindings for CPUs - Convert one use of __ASSEMBLY__ to __ASSEMBLER__ that snuck into the RISC-V CFI selftest code - Update the translation for the simplified Chinese translation of the RISC-V kernel patch acceptance policy" * tag 'riscv-for-linus-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/riscv/linux: riscv: skip software algning code for HAVE_EFFICIENT_UNALIGNED_ACCESS kselftest/riscv: Replace __ASSEMBLY__ with __ASSEMBLER__ docs/zh_CN: Update arch/riscv/patch-acceptance.rst translation dt-bindings: riscv: cpus: Fix yamllint style issues riscv: hwprobe: simplify has_fpu() to check D extension only perf: RISC-V: check cpu_hw_evt before dereference in overflow IRQ riscv: report Zfhmin/Zvfhmin when Zfh/Zvfh are present perf: RISC-V: store available counter mask as bitmap perf: RISC-V: use BIT_ULL for u64 overflow masks riscv: bug: Make RV32 use GENERIC_BUG_RELATIVE_POINTERS riscv: hwprobe: initialize pair->value in hwprobe_one_pair() riscv: use string helper in setup_global_riscv_enable() Revert "riscv: Reset pmm when PR_TAGGED_ADDR_ENABLE is not set" riscv: patch: skip fixmap mapping when kernel text is already writable riscv: mm: make EXECMEM_KPROBES writable without CONFIG_STRICT_MODULE_RWX
2026-09-11KVM: Never clear KVM_REQ_VM_DEAD from a vCPU's requestsSean Christopherson
Use kvm_test_request() instead of kvm_check_request() when querying KVM_REQ_VM_DEAD, i.e. don't clear KVM_REQ_VM_DEAD, as the entire purpose of KVM_REQ_VM_DEAD is to prevent the vCPU from enterring the guest ever again, even if userspace insists on redoing KVM_RUN. Ensuring KVM_REQ_VM_DEAD is never cleared will allow relaxing KVM's rule that ioctls can't be invoked on dead VMs, to only disallow ioctls if the VM is bugged, i.e. if KVM hit a KVM_BUG_ON(). Opportunistically add compile-time assertions to guard against clearing KVM_REQ_VM_DEAD through the standard APIs. Reviewed-by: Kai Huang <kai.huang@intel.com> Acked-by: Marc Zyngier <maz@kernel.org> Link: https://patch.msgid.link/20260806214618.82180-1-seanjc@google.com Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-09-11Merge tag 'usb-serial-7.3-rc3' of ↵Greg Kroah-Hartman
ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial into usb-linus Johan writes: USB serial fixes for 7.3-rc3 Here is a fix for a long-standing ioctl-hangup race and a couple of fixes for port lifetime issues that can lead to NULL-pointer dereferences when disconnecting devices or deregistering drivers. Included are also a fix for a related dynamic id leak and some new modem device ids. All have been in linux-next with no reported issues. * tag 'usb-serial-7.3-rc3' of ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial: USB: serial: fix ioctl hangup race USB: serial: use iterator for driver deregistration USB: serial: fix driver deregistration order USB: serial: fix dynamic id driver deregistration race USB: serial: fix port tear down use-after-free USB: serial: option: add Compal EXC-T1 support USB: serial: option: add Quectel RG660QB USB: serial: option: add Compal EXM-G1x support USB: serial: option: add support for SIMCom SIM8260C USB: serial: option: add Quectel EG060W
2026-09-10bpf: Add KF_PERFMON kfunc flagDaniel Borkmann
Tracing related BPF helpers e.g. under bpf_base_func_proto() are gated behind CAP_PERFMON. However, the same is currently not true for kfuncs and they are accessible via plain CAP_BPF. Add a new KF_PERFMON flag which can be used such that check_kfunc_call() ensures env->allow_ptr_leaks is permitted. This follows similar pattern to existing KF_DESTRUCTIVE flag. The rejection returns -EPERM to match the other CAP_PERFMON gates in the verifier, that is, check_ptr_to_btf_access() and check_ptr_to_map_access(), which report the very same policy to user space. Signed-off-by: Daniel Borkmann <daniel@iogearbox.net> Link: https://lore.kernel.org/r/20260910213510.49358-1-daniel@iogearbox.net Signed-off-by: Alexei Starovoitov <ast@kernel.org>
2026-09-10Merge tag 'net-7.3-rc3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Nothing too exciting, usual stream of fixes. Including fixes from Netfilter, Bluetooth and WPAN. Current release - new code bugs: - Bluetooth: hci_sync: fix not setting CE length properly - eth: enic: match mailbox replies to request numbers Previous releases - regressions: - tunnels: drop stale dst when building an ICMP error for PMTUD - ipv6: null-check fib6_node before accessing in __ip6_del_rt_siblings() (bug in the rtnl_lock -> RCU conversion) - eth: bnxt_en: - fix crashes on Thor2 due to OOB coalescing buffer accesses - prevent queue stop with deferred completions Previous releases - always broken: - eth: - ice: don't dereference pointers from TP_printk() - fix OOB writes on ethtool flow rule dump in 3 drivers - mlx5: fix FEC configuration with RS_544_514_INTERLEAVED_QUAD - dsa: tag_brcm: legacy FCS: request needed tailroom Misc: - net: cap tx_queue_len at S16_MAX to prevent oversized ring alloc - ipv6: flowlabel: cap duplicate leases per socket" * tag 'net-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (164 commits) selftests: tc-testing: test action batch failure cleanup net/sched: act_api: release all action references on NEWACTION failure openvswitch: fix wrong flag value in get_ipv6_ext_hdrs() ipmr: account multicast table and route memory net: phy: dp83td510: handle the active-high LED polarity mode net: macb: initialize PTP state before registering clock net: hsr: enable promiscuous mode on interlink port with fwd offload ipv6: fix fib6 walker UAF on seq stop net: stmmac: fix TX descriptor availability check for TSO traffic net/rds: fix tcp stream corruption with large pages net: mana: restore the XDP program pointer when pre-allocation fails net: phy: dp83867: handle the active-high LED polarity mode octeontx2-af: fix PF/CGX debugfs PCI bus lookup net: net_failover: Fix the deadlock in net_failover_slave_name_change() net: phy: mediatek-ge: disable EEE on the MT7530 PHY tcp: reject non zerocopy devmem tx net: ethernet: mtk_eth_soc: populate lpi_interfaces to fix EEE support net: dsa: mt7530: populate lpi_interfaces to fix EEE support net: hinic: fix mailbox segment buffer overflow net: sun4i-emac: fix missing of_node_put() for phy_node ...
2026-09-10sched/eevdf: Fix rb augmented with multi fieldsVincent Guittot
The eevdf rb tree maintains 3 augmented fields but only one is currently copied when balancing the tree. Add a more generic define that can be used when there are several augmented fields. In this case, we provide a function that takes care of copying all fields. Fixes: aef6987d8954 ("sched/eevdf: Propagate min_slice up the cgroup hierarchy") Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260909150522.858312-1-vincent.guittot@linaro.org
2026-09-09Merge tag 'vfs-7.3-rc3.fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: - netfs: - Fix an uninitialized return value in netfs_unbuffered_write() when preparing the first subrequest fails - For partial unbuffered/DIO writes return the amount transferred rather than an error - Update i_size with the amount actually written when a partial transfer ends in an error - Fix a subrequest reference leak when the io_iter ends up empty - Handle netfs_alloc_subrequest() failure during unbuffered writes - Load all readahead folios into the rolling buffer upfront and drop the readahead references once the first subrequest is dispatched - Mark folios for copy-to-cache while issuing subrequests - Fix read progress reporting - afs: - Add the missing kunmap in the error path of afs_dir_search_bucket() - Fix a double kunmap in afs_edit_dir_remove() - Don't free an existing server's endpoint state when cleaning up a candidate server in afs_lookup_server() - Unbind peers removed from a server's address list - ufs: - Load the cylinder group metadata before creating the root dentry - Validate the cylinder group index and rotor positions before caching them - Treat an unreadable directory block as not empty - exec: - Close the close-on-exec files before taking exec_update_lock Closing a file can block on the filesystem, so a hung filesystem blocked everything that takes exec_update_lock and a FUSE server inspecting the calling process could deadlock - Drop the bprm loader before closing bprm->file in free_bprm() - exit: Hold a reference to thread_pid across proc_flush_pid() - reboot: Fix a use-after-free on cad_pid - nsfs: Keep the namespace tree fields out of the rcu_head used by kfree_rcu() - nstree: Check listing permission before taking a namespace reference in listns() - super: Return 0 when a nested thaw drops its hold while other freezers remain - ext4: Don't set I_METADATA_WRITEBACK during fastcommit replay - adfs: Free s_fs_info in ->kill_sb() - autofs: Free the inode info allocated in autofs_fill_super() when the root inode allocation fails - ovl: Return EINVAL instead of EIO on a user namespace mismatch now that it's a plain refusal and not an internal error - cachefiles: Don't cast the variable-length coherency data to a __be64 in the coherency tracepoint * tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (28 commits) nstree: check listing permission before taking a namespace reference exec: do_close_on_exec() before taking exec_update_lock exit: hold a reference to thread_pid across proc_flush_pid fs: autofs: fix memory leak in autofs_fill_super() exec: Drop bprm loader before closing bprm->file afs: Clear stale peer app data after address list changes afs: Fix incorrect free in candidate cleanup in afs_lookup_server() afs: Fix double-unmap of directory block afs: Fix missing kunmap in afs_dir_search_bucket() ovl: return EINVAL instead of EIO in case of mismatched user_ns reboot: fix cad_pid use-after-free race cachefiles: Fix potential UAF/KASAN warning netfs: Fix read progress reporting netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs netfs: Fix readahead synchronisation issues by loading all folios upfront netfs: break unbuffered write when netfs_alloc_subrequest() fails netfs: Fix subreq ref leak netfs: Fix i_size update for partial transfer netfs: Fix error vs transferred passed to ->ki_complete() netfs: Fix unbuffered/DIO write partial transfer error return ...
2026-09-09Merge tag 'printk-for-7.3-rc3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux Pull printk fixes from Petr Mladek: - Use lazy irq_work for waking printk kthreads - Flush pending irq_work before destroying printk kthreads - Remove redundant WARN() when a printk kthread can't be created - Typo fix * tag 'printk-for-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/printk/linux: printk/nbcon: Change nbcon_irq_work to IRQ_WORK_LAZY printk/nbcon: Flush nbcon_irq_work in nbcon_free() console: fix /dev/kmsg reference in flags kernel doc printk: Don't WARN on kthread_run failure.
2026-09-09Merge tag 'thunderbolt-for-v7.3-rc3' of ↵Greg Kroah-Hartman
ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/westeri/thunderbolt into usb-linus Mika writes: thunderbolt: Fixes for v7.3-rc3 This includes following USB4/Thunderbolt fixes: - Fix various issues around asynchronous DisplayPort tunnel activation when there is no graphics driver doing doing the capability exchange. - Fix potential NULL pointer dereference when XDomain connection is removed. - Fix use-after-free when control channel request is canceled. - Fix lockdep false positive. - Revert a commit that causes XDomain properties ping-pong. All these have been in linux-next with no reported issues. * tag 'thunderbolt-for-v7.3-rc3' of ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/westeri/thunderbolt: Revert "thunderbolt: xdomain: Notify peers after enumeration" thunderbolt: Use separate lock class for each ring thunderbolt: Fix KASAN reported use-after-free when request is canceled thunderbolt: Fix NULL dereference in tb_remove_work() thunderbolt: Tear down inactive DP tunnels when the domain is stopped thunderbolt: Mark discovered tunnels as active thunderbolt: Don't access a DP tunnel after its DPRX read was canceled thunderbolt: Fix domain reference leak when DPRX read is canceled thunderbolt: Make the DP tunnel activation callback mandatory thunderbolt: Hold a router reference for each allocated HopID
2026-09-09x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAFLorenzo Stoakes (ARM)
x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and BPF when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. In addition, concurrent CPA collapse operations are possible which can also cause races. Resolve the issue by acquiring the mmap write lock on init_mm across the whole operation. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org> Reviewed-by: Kiryl Shutsemau (Meta) <kas@kernel.org> Reviewed-by: David Hildenbrand (Arm) <david@kernel.org> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Will Deacon <will@kernel.org> Reviewed-by: David Carlier <devnexen@gmail.com> Tested-by: Atish Patra <atishp@meta.com> Tested-by: Nikunj A Dadhania <nikunj@amd.com> Cc:stable@vger.kernel.org Link: https://patch.msgid.link/20260813-cpa-fixes-v2-1-39b4ff90f91d@kernel.org
2026-09-08Merge tag 'nf-26-09-07' of ↵Jakub Kicinski
git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf Pablo Neira Ayuso says: ==================== Netfilter/IPVS fixes for net The following patchset contains Netfilter/IPVS fixes for net: 1) Reject malformed messages in IPVS sync, from Kyle Zeng. 2) Fix possible stale infoleak in IPVS sync, also from Kyle Zeng. 3) Out-of-bound read in the SIP conntrack helper, from Joas Antonio dos Santos. 4) UaF on cttimeout module removal, from Chengfeng Ye. 5) Unregister nf_loggers before netns teardown to fix UaF, also from Chengfeng Ye. 6) Fix race in nfnetlink_log due to concurrent instance destruction, from Florian Westphal. 7) Remove arp_table 32bit compat interface, this is already off in many distributions, from Florian Westphal. 8) Set IP6T_F_PROTO flag is e->ipv6.proto is set on to deal with insufficient validation of xtables extensions when used from legacy ip6tables, from Florian. 9) Set on the NLM_F_DUMP_FILTERED flag when all is filtering out in ctnetlink, from Ilya Maximets. * tag 'nf-26-09-07' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf: netfilter: report NLM_F_DUMP_FILTERED when all is filtered out netfilter: ip6_tables: set F_PROTO when proto value is nonzero netfilter: arp_tables: remove the 32bit compat interface netfilter: nfnetlink_log: cope with concurrent instance destruction netfilter: nf_log: unregister loggers before per-net teardown netfilter: cttimeout: prevent UAF during module unload netfilter: nf_conntrack_sip: fix OOB read in sip_skip_whitespace() ipvs: fix reversed sequence option serialization ipvs: reject invalid states in connection template sync records ==================== Link: https://patch.msgid.link/20260907171732.1407739-1-pablo@netfilter.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-08USB: serial: fix driver deregistration orderJohan Hovold
USB serial driver modules register one driver for the USB bus and one or more drivers for the ports on the USB serial bus. When unloading a driver module, the USB driver must be deregistered before the USB serial bus drivers so that I/O is stopped before unbinding the ports to avoid use-after-free in completion handlers accessing port data. Note that the "new_id" attributes must first be removed to prevent new ids from being added and triggering a probe of the USB driver after it has been deregistered. Fixes: 765e0ba62613 ("usb-serial: new API for driver registration") Cc: stable@vger.kernel.org # 3.4 Cc: Alan Stern <stern@rowland.harvard.edu> Signed-off-by: Johan Hovold <johan@kernel.org>
2026-09-07netfilter: arp_tables: remove the 32bit compat interfaceFlorian Westphal
This feature is required to use 32bit arptables binary on 64bit kernels. It's already off in many distributions including Debian and Fedora for many years. Zap arptables first, it's the most esoteric of the 4 flavors. Signed-off-by: Florian Westphal <fw@strlen.de> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-07entry: Guard syscall_enter_audit() invocation with CONFIG_AUDITSYSCALLThomas Gleixner
A bunch of older cross compilers notably RISCV64 and S390 fail to eliminate the dead code when CONFIG_AUDITSYSCALL=n. The code in question is: if (unlikely(audit_context()) syscall_enter_audit(regs); and in case of CONFIG_AUDITSYSCALL=n: static inline struct audit_context *audit_context(void) { return NULL; } which should make the compiler eliminate the syscall_enter_audit() call. But a RISV64 GCC12 cross compiler translates that into: if (unlikely(audit_context())) 1c34: 00000097 auipc ra,0x0 1c38: 000080e7 jalr ra # 1c34 <.L785> 1c3c: c511 beqz a0,1c48 <.L787> syscall_enter_audit(regs); 1c3e: 8526 mv a0,s1 1c40: 00000097 auipc ra,0x0 1c44: 000080e7 jalr ra # 1c40 <.L785+0xc> and then claims in the failing link: include/asm-generic/preempt.h:54:(.noinstr.text+0x1a20): undefined reference to 'syscall_enter_audit' which is obviously hallucination. Add an explicit IS_ENABLED(CONFIG_AUDITSYSCALL) check into the condition to cure this compiler madness. Fixes: 6f25517010dd ("entry: Rework syscall_audit_enter()") Reported-by: kernel test robot <lkp@intel.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/87tso45bqq.ffs@fw13 Closes: https://lore.kernel.org/oe-kbuild-all/202609031938.ZvZZaRQy-lkp@intel.com/
2026-09-06Merge tag 'trace-v7.3-rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull tracing fixes from Steven Rostedt: - Fix several tracefs files that did not take the trace_array reference A trace instance can be created and destroyed in the tracefs "instances" directory via mkdir and rmdir respectively. The instance is represented by a trace_array descriptor. Most tracefs files pass the trace_array as the private data of the inode to the open/read/write functions. Since there is no locking between the time a task opens a file and the deletion of the instance (and the freeing of the trace_array), each open needs to get a reference to the trace_array and each close must remove it. An instance can't be removed if there's any reference taken on its trace_array. The open function uses trace_array_get() that takes a lock (preventing removal of instances) and iterates the list of all existing trace_arrays and if it finds a match, it takes the reference and releases the lock. If it doesn't find a match, it causes the open to return -ENODEV. There were some added files that did not take the trace_array reference on open that needed to be fixed. Sashiko also correctly pointed out that there were some files that took an address of an field or element of the trace_array which had a pointer back to the trace_array to take its reference on open. But this leaves a slight race between referencing this element to get the trace_array as the element itself could be freed. To solve this, some helper functions were created to look for trace_arrays with this field or element in the search so that the element did not have to be dereferenced before the trace_array's reference was taken. - Add a lock around ftrace_ops initialization When a ftrace_ops is first used by ftrace, some internal initialization is performed on the ops. But if multiple tasks were calling functions that did this initialization, it could race and perform doing the initialization more than once, corrupting the internal data. Add a lock in the initialization code to prevent this from happening. - Fix splice reads on mmapped buffers The logic in the ring buffer splice code for mmapped buffers is supposed to do a copy of the memory as the mapped buffers can't be given to splice. But there was an if statement within the copy code that would return a -1 if a request for a full page was done and it wasn't a partial read. This is because this logic was written before mmapped buffers existed and this case didn't make sense at the time. For mmapped buffers it makes perfect sense and by returning early can drop a lot of pages unnecessarily. - Have the persistent ring buffer validation check nr_subbufs Sashiko reported that the validation code was relying on the saved nr_subbufs to match the calculated nr_pages + 1 and if they were off, that the code could cause corruption. Sashiko is correct, and the saved nr_subbufs should be validated before assuming it is correct. - Do not allow more than one instance with the same name on cmdline If an admin were to add more than one trace instances with the same name they all would be created, but only the first one would be accessible via tracefs. This used to not be allowed but some restructuring of code has since made it possible. - Fix the race between subbuf resize and trace_pipe_raw readers If a task was reading trace_pipe_raw while another task was changing the ring buffer subbuf size, it could crash the reader. The trace_pipe_raw readers do get their own copy of the page from the buffer, but the code needs some restructuring to not have the resize of the subbuffers cause issues. - Cap the size of the mapped (static) ring buffer nr_pages The meta data used for ring buffer mapped buffers is 32 bit in size. A normal ring buffer could (in theory) have more than 4 billion pages. But this is not allowed by mapped buffers, so enforce it. * tag 'trace-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: ring-buffer: Use a macro for static buffer bits tracing: Fix comment in tracing_buffers_splice_read() ring-buffer: Prevent truncation of nr_pages / nr_subbufs ring-buffer: Cap static ring buffer nr_pages tracing: Fix subbuf resize races with trace_pipe_raw readers tracing: Fix to avoid creating trace instances with duplicate names ring-buffer: Add checking nr_subbufs to persistent ring buffer validation ring-buffer: Allow splice reads on static buffers tracing: Take trace_array reference when opening options file ftrace: Synchronize the initialization of ftrace_ops ftrace: Take trace_array reference before accessing its ftrace_ops tracing: Have show_event_filters/triggers files take trace array ref
2026-09-06Merge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: "This mainly contains verifier fixes that address bugs reported by Nicholas Carlini. - Fix incorrect non-NULL inference in pointer comparisons: pointer types that may be NULL at runtime, pointers with unbounded offsets, JMP32 comparisons with zero, and imprecise zero registers (Eduard Zingerman) - Fix precision tracking for half-dead zero spills, ld_abs/ld_ind implicit subprog exit, bpf_loop() callbacks, linked scalar ids and NULL call arguments (Eduard Zingerman) - Reject BPF_PSEUDO_FUNC reference to the main program, fix zero extension of arena 32-bit cmpxchg, don't rewrite bpf_fastcall patterns entered by a jump (Eduard Zingerman) - Fix percpu map update and BPF_F_CPU validation with sparse CPU IDs (Hui Su) - Fix NULL-ptr-derefs in bpf_snprintf_btf() for void and VAR types, and reject key-less BTF for hash maps (Jiayuan Chen) - Various fixes (Kumar Kartikeya Dwivedi): - Fix out-of-bounds access in disassembler on invalid LDSX instruction - mark siginfo of signal tracepoints as scalar and sched_process_wait argument as nullable - mark faultable stack helpers as sleepable - reject tail calls and legacy packet loads from callbacks - enforce rbtree callback lock restrictions for resilient locks - require MEM_PERCPU for percpu kptr stores - clear NON_OWN_REF after RCU protection ends - mark NULL kptr stores precise - preserve inner map identity in callback frames - reject non-scalar bpf_loop() iteration counts - Fix trampoline allocation slowdown on x86 by using EXECMEM_MODULE_DATA (Mike Rapoport) - Keep bpf_refcount_acquire() nullable for borrowed RCU kptrs and reject untrusted allocated-object pointers (Ning Ding) - Fix special fields handling in recycled rhtab elements (Nuoqi Gui, Yuan Chen)" * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (86 commits) bpf, riscv: Make arena support depend on ZACAS selftests/bpf: Test pointer bpf_loop iteration count rejection bpf: Reject non-scalar bpf_loop iteration counts bpf: use mark_arg_precision() in check_mem_size_reg() bpf: propagate mark_chain_precision() errors out of loop_flag_is_zero() selftests/bpf: precision of a NULL global subprogram BTF_ID argument bpf: mark a NULL BTF_ID argument of a global subprogram precise selftests/bpf: precision of a NULL kfunc argument bpf: mark a NULL kfunc argument precise selftests/bpf: precision of a NULL global subprogram memory argument bpf: mark a NULL memory argument of a call precise selftests/bpf: precision of a NULL helper argument bpf: mark a NULL call argument precise selftests/bpf: Test inner map identities in callbacks bpf: Preserve inner map identity in callback frames selftests/bpf: Test imprecise scalar kptr stores bpf: Mark NULL kptr stores precise selftests/bpf: Test rhtab kptr cancellation semantics bpf: Cancel special fields when recycling rhtab elements selftests/bpf: Test timer field on recycled rhtab element ...
2026-09-06Merge tag 'locking-urgent-2026-09-06' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull locking fixes from Ingo Molnar: - Fix a softirq processing delay bug in local_interrupt_disable(), which should mostly only affect the Rust runtime (Boqun Feng) - Remove the hardirq_disable_count() function which caused the previous bug and is now unused & unnecessary (Boqun Feng) - lockdep: Invalidate stale class_cache entries for zapped classes (Eric Dumazet) - Fix rt_mutex specific futex scheduling helpers (Sebastian Andrzej Siewior) - Fix rcuwait use-after-free race during futex requeue PI (Yao Kai) * tag 'locking-urgent-2026-09-06' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: futex: Prevent rcuwait use-after-free during requeue PI futex: Provide rt_mutex_.*_schedule() equivalents for futex scheduling locking/lockdep: Invalidate stale class_cache entries for zapped classes preempt: Remove hardirq_disable_count() interrupt: Disable interrupt before modifying hardirq_disable counter
2026-09-06Merge tag 'irq-urgent-2026-09-06' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull IRQ subsystem fixes from Ingo Molnar: - Revert a commit to the mbigen irqchip driver that caused a regression on two-port Hi1616 chips (Caina) - Fix a too-long-preemption-off bug in the stm32mp-exti irqchip driver, caused by a time unit ambiguity & mismatch (Ju Nan) - Remove the now completely unused irq_domain_add_linear() inline function (Jiri Slaby) * tag 'irq-urgent-2026-09-06' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: irqchip/stm32mp-exti: Fix the unit of the hwspinlock timeout Revert "irqchip/mbigen: Fix mbigen node address layout" irqdomain: Delete irq_domain_add_linear()
2026-09-05bpf: Reject non-scalar bpf_loop iteration countsKumar Kartikeya Dwivedi
bpf_loop() declares its nr_loops argument as ARG_ANYTHING. Privileged programs may pass pointer values to such arguments, so check_func_arg() lets a pointer-valued R1 reach the helper-specific checks. Since commit bb124da69c47 ("bpf: keep track of max number of bpf_loop callback iterations"), the verifier marks R1 precise and reads its upper bound to limit callback simulation. Precision backtracking only accepts scalar registers, so passing a pointer instead triggers the "backtracking misuse" verifier warning. Kernels with panic_on_warn enabled subsequently panic. Introduce ARG_SCALAR for helper arguments that only accept scalar values and use it for bpf_loop() nr_loops. Generic helper argument validation then rejects pointers before loop inlining and precision processing. Fixes: bb124da69c47 ("bpf: keep track of max number of bpf_loop callback iterations") Reported-by: syzbot+7b47f87674e9a1569110@syzkaller.appspotmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> Link: https://patch.msgid.link/20260905014735.1452988-2-memxor@gmail.com Closes: https://lore.kernel.org/bpf/6a9ad24c.b5d4176b.238c3e.0001.GAE@google.com/ Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-05Merge tag 'block-7.3-20260905' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux Pull block fixes from Jens Axboe: - NVMe fixes via Keith: - nvme-tcp fixes for an out-of-bounds write on an over-long PDU - nvmet-tcp, nvmet-rdma and nvme-rdma leak and cleanup-ordering fixes - FDP placement id array racy access fix - nvme-fc double free of fabrics options on nvme_add_ctrl() failure, and a secret leak failure - Fault injection opcode filtering - stale namespace removal during scan - Various other smaller fixes and cleanups - Flag zoned disks with GENHD_FL_NO_PART - Save the page offset gaps in a cloned bio - Fix dma_alignment for large or unreported limits in loop and zloop - Clear VM_MAYWRITE on a read-only ublk char device mmap * tag 'block-7.3-20260905' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux: (25 commits) nvme-tcp.h: drop kernel-doc comments, fix a few descriptions nvme-fc: fix double free of fabrics options when nvme_add_ctrl() fails nvmet: reject namespace enable without device path nvmet-auth: Synchronize timeout work during SQ teardown MAINTAINERS: update nvme entry nvmet-tcp: reject unsolicited H2CData PDUs nvme-tcp: defer TLS inline send to io_work nvmet-tcp: fix out-of-bounds write when receiving an over-long PDU nvme-tcp: return -EPROTO for a C2HData on a write nvmet: print namespace IDs as unsigned 32bit value nvme: print namespace IDs as unsigned 32bit value nvme: remove stale namespaces by NSID range during scan nvme: add missing SRCU grace period in error path nvme-fabrics: fix DHCHAP secret leak on parse failure ublk: clear VM_MAYWRITE on read-only ublk char device mmap loop, zloop: fix dma_alignment for large or unreported limits block: save page offset gaps in cloned bio block: flag zoned disks with GENHD_FL_NO_PART nvmet-rdma: fix queue leak when connect backlog is exceeded nvme: add opcode filtering for fault injection ...
2026-09-04ethtool: document that GRXCLSRLALL rule_cnt is a caller-provided limitJakub Kicinski
Three drivers have shipped a get_rxnfc() which dumps its entire rule table into rule_locs, reading rule_cnt as "how many rules do I have" rather than "how many entries did the caller allocate". Nothing in the callback's documentation contradicted that reading. The distinction only matters because the ioctl lets an unprivileged caller pick rule_cnt directly, so getting it wrong is a heap overflow rather than a truncated dump. Reviewed-by: Joe Damato <joe@dama.to> Link: https://patch.msgid.link/20260903032611.3000029-6-kuba@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04Merge tag 'drm-fixes-2026-09-05' of https://gitlab.freedesktop.org/drm/kernelLinus Torvalds
Pull drm fixes from Dave Airlie: "Lots of scattered fixes: nouveau has a bunch of display fixes for blackwell GPUs that should mean we light up monitors properly and fix some desktop rendering problems, amdgpu and intel display changes as usual. There also changes to the core pagemap, then the usual amouny of AI inspired validation fixes. core: - Fix drm_crtc_commit leak when PAGE_FLIP_EVENT is used dma-buf: - Publish the dma-buf only after copy_to_user succeeds - fix some kernel-doc warnings atomic-state-helpers: - set pixel_blend_mode to prop default on reset sysfb: - Fix integer overflow - fix constant comparison bug pagemap: - Prevent double migration of device pages - Reset migration page count on eviction retry - dma-unmap pages before handling migration errors - use after free fixes prime: - fix prime exports tracing amdgpu: - Fix for drm_amdgpu_info_device with mixed 64 bit kernel and 32 bit userspace - plane blend mode fixes - SR-IOV fix - GFX8 fix - MES queue reset fix - GPUVM fixes - DCN 6 warning fix - DCN 3.5/3.6 fix - DML fix - Backlight fix - Colorop fix - DC get_estimated_bw() fix - devcoredump fix - Userq fixes - APU PSP fix - Cursor fix amdkfd: - MES queue eviction fix - MQD debugfs fix xe: - oa uapi error handling fix - drm info message to report FLAT_CSS base misalignment i915: - Drop an accidentally duplicated panel fitter call in DP MST - Fix DDI clock programming for Cx0 and LT PHY - Fix PTL CDCLK handling at probe, causing a glitch - Fix dg2_power_well_count() return type - Fix a NULL pointer deref at forced probe - Fix selective fetch disable amdxdna: - out-of-bounds access fix - reject commands chains with no commands - handle chained mapping BO failures - refuse to flush an imported BO ethosu: - handle mmio mapping failures - handle storage modes only on hardware that supports it - fix job completion fence cleanup fastrpc: - Publish the dma-buf only after copy_to_user succeeds gud: - Improve TV modes and rotation handling nouveau: - use-after-free fixes - add missing scanline position support - HDMI and DP fixes - null pointer dereference fix - dmem accounting fixes for large folios - use write-combined maps for coherent qaic: - out-of-bounds access fix tegra: - Add blend mode properties virtio: - exit path and error handling fixes * tag 'drm-fixes-2026-09-05' of https://gitlab.freedesktop.org/drm/kernel: (83 commits) drm/xe/vram: report FLAT_CCS base misalignment MAINTAINERS, mailmap: use Aditya Garg's linux.dev account drm/amd/display: use plane color_mgmt_changed to track colorop changes drm/amdgpu/userq: fix struct drm_amdgpu_info_device padding for 32bit compile drm/amd/display: Fix cursor disable with horizontally split planes drm/amdgpu/userq: dont overwrite the error of subsequent map call drm/amdgpu: Skip accessing psp rum time db for APUs drm/amdgpu: update the fw version for gfx12 userqueues drm/amdgpu: update the fw version for gfx11 userqueues drm/amdgpu: fix byte/dword unit mismatch in coredump IB dump drm/amdkfd: fix scope of mqd_mgr dereference in pqm_debugfs_mqds drm/amd/display: fix division by zero in get_estimated_bw() drm/amd/display: use halving distribution for all encode-to-linear curves drm/amd/display: Fix backlight control for luminance-capable OLED drm/amd/display: Remove const Qualifier From Non-Pointer Fields drm/amd/display: Set gpuvm min page size to 4K on dcn35/36 drm/amd/display: Fix DCN5/6 DML2 compilation warnings drm/amdgpu: fix Idle BOs list in VM debugfs status info drm/amdgpu: use AMDGPU_GPU_PAGE_SHIFT instead of PAGE_SHIFT drm/amdgpu: Update queue reset support version ...
2026-09-04tracing: Fix subbuf resize races with trace_pipe_raw readersVincent Donnefort
Concurrent subbuffer resizes may crash trace_pipe_raw readers or leak uninitialized memory to userspace due to stale size values. Modify ring_buffer_alloc_read_page() to handle the resizing of an existing buffer_data_read_page if necessary and add a new ring_buffer_read_page_size(). This new function enables ring-buffer buffer_data_read_page users to not call the racy ring_buffer_subbuf_size_get(). This makes the spare_size member of ftrace_buffer_info redundant. Finally, handle buffer_data_read_page/reader_page order discrepancy in ring_buffer_read_page(). On a mismatch simply copy manually the data to the buffer_data_read_page. Link: https://lore.kernel.org/all/20260817140812.2C7D41F00A3A@smtp.kernel.org/ Link: https://patch.msgid.link/20260904164450.1345852-3-vdonnefort@google.com Fixes: bce761d75745 ("ring-buffer: Read and write to ring buffers with custom sub buffer size") Signed-off-by: Vincent Donnefort <vdonnefort@google.com> Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-04mtd: spinand: fix NULL pointer dereference with no ECC engineNuno Sá
When "nand-no-ecc-engine" is set in DT, nanddev_get_ecc_engine() takes the NAND_ECC_ENGINE_TYPE_NONE path and returns success while leaving nand->ecc.engine NULL. The SPI-NAND code nevertheless dereferences it unconditionally to test for a pipelined engine, so probing such a device oopses immediately. Rather than open-coding the test three times, add a nand_ecc_is_pipelined() helper to the NAND core that folds the NULL check into the integration comparison, and use it everywhere. Future callers then cannot reintroduce the problem. Fixes: f9d7c7265bcf ("mtd: spinand: Create direct mapping descriptors for ECC operations") Cc: stable@vger.kernel.org Signed-off-by: Nuno Sá <nuno.sa@analog.com> Signed-off-by: Miquel Raynal <miquel.raynal@bootlin.com>
2026-09-04Merge tag 'probes-fixes-v7.3-rc1' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull probes fixes from Masami Hiramatsu: - Protect kprobe_blacklist with RCU RCU-protect kprobe_blacklist and use kfree_rcu() to prevent UAF races during module unloading and enable safe atomic lookups. - Fix multi-probe field use-after-free Duplicate field and type strings on trace_probe_event to prevent UAF when freeing primary probe - Fix probe BTF member lookup: Check the containing inner struct/union kflag when resolving anonymous members to ensure correct bitfield offset calculation Prevent unnamed bitfields from being pushed to anon_stack in btf_find_struct_member(), avoiding false lookup errors Fix code block indentation in get_bitoffset_of_field() - uprobes error pointer safety Guard free_trace_uprobe() with IS_ERR_OR_NULL() to avoid crashing during automatic cleanup when an error pointer is returned * tag 'probes-fixes-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: kprobes: Protect kprobe_blacklist with RCU tracing/probes: Fix use-after-free on field name/type of events with multiple probes tracing/probes: Fix code indent in get_bitoffset_of_field() tracing/probes: Fix BTF kflag check for anonymous struct member access tracing/probes: Fix anon_stack check for unnamed bitfields in btf_find_struct_member uprobes: guard trace cleanup against error pointers