summaryrefslogtreecommitdiff
path: root/tools
AgeCommit message (Collapse)Author
108 min.Merge tag 'vfs-7.3-rc7.fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs Pull vfs fixes from Christian Brauner: "This contains fixes for the current development cycle. All of them came out of a review of the mount code that started with a bug report. The review modeled the corner cases of mount propagation, unmounting and mount reference counting and turned up a lot of bugs. Most of them years old. Most fixes come with a selftest. - Rework connected mounts. A mount that is unmounted together with its parent can stay attached to the parent to keep its mountpoint covered. That happens when the mountpoint is removed with rmdir(), unlink() or rename(), when a detached tree is dissolved, and for locked mounts in any umount that isn't synchronous, including the teardown of their mount namespace. The parent then owns the child and drops it on its own final mntput(). So any reference from the child's superblock back to one of its ancestors becomes a cycle that is never freed. A loop device backed by an image on a tmpfs and mounted on that same tmpfs is enough. Remove the directory the tmpfs is mounted on from the host, let the container's mount namespace exit, and the loop device, the tmpfs and the filesystem on the loop device are leaked for good. The same works with autofs, zram, ecryptfs, binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md, and the selftests have reproducers for them. This has been possible since v4.1. It is also why "put_mnt_ns(): leave mounts connected" was reverted in -rc5. Keeping every mount of a dying mount namespace connected made these cycles trivial to create. Every unmounted mount is now detached from its parent. Where the mountpoint has to stay covered the mount leaves a cover on the parent instead, allocated together with the mount. A lookup on the unmounted parent that hits a cover finds an empty immutable directory or file on the private nullfs instance. Nothing leads from a cover to another mount, so no unmounted mount owns another one and no cycle can form. This is visible to userspace. A formerly connected mount can no longer be reached through its unmounted parent and ".." inside it leads nowhere, as for every other lazily unmounted mount. With that the private nullfs instance becomes reachable from userspace, so it now refuses mounts on top, is mounted read-only, and refuses fsnotify marks and file locks. Its inodes are shared by every holder and a watch or a lock would otherwise reach across users. may_decode_fh() now also decides its subtree check under a single mount_lock hold, as a racing umount could otherwise let it decode into what a locked child covered. - umount: * Don't silently unmount busy mounts. Since v4.13 propagate_umount() takes down propagated copies of the victim with children as long as each child is an overmount or another copy of the victim, but propagate_mount_busy() only ever checked copies without children or with just an overmount. A container that moved a tree beneath its copy of a host mount lost that tree from under its open file descriptor to a plain umount() on the host. propagate_mount_busy() now applies the same rules, walking each chain of copies once. * Don't let a migrating task hide its reference from umount(). mnt_get_count() sums the per-cpu counters under mount_lock but the mntget() and mntput() fast paths don't take it. A task that takes a reference on a cpu the sum has already passed and drops it after migrating to one the sum hasn't reached yet hides the reference it held to begin with, and umount() succeeds with the file still open. Gets and puts now live in separate per-cpu counters and all puts are summed before all gets with a full barrier in between, the way srcu_readers_active_idx_check() does it. mntget() is unchanged and mntput() gains an smp_wmb(). * Check each submount for references right before unmounting it. shrink_submounts() and mark_mounts_for_expiry() checked all their victims up front. Unmounting the first could move a busy overmount to where the next victim's propagated copy is looked up and it was then unmounted without a check. * Never expire a locked mount. A shrinkable mount moved beneath a locked mount with MOVE_MOUNT_BENEATH takes over the lock, and umount() of an unlocked ancestor expired it and revealed what it covered. That umount() now fails with EBUSY as it does for any other locked child. A lazy umount still takes the whole tree. - Overmounts and locked mounts: * Unhash a dentry before detaching the mounts on it. unlink(), rmdir() and rename() detach the mounts on the victim but only d_delete() it once its inode is unlocked, a window that includes an expedited RCU grace period. In between, a lookup from a mount namespace in which the dentry is a mountpoint found it hashed, positive and uncovered. Drop the dentry first, as d_invalidate() already does. * Don't reveal overmounted entries in refwalk. A refwalk that had grabbed the dentry before the unlink never rechecked it the way rcuwalk does with d_seq and mount_lock. Without any artificial widening three walkers read the covered file 27 times in a minute. step_into() now fails an unhashed dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is retried. * Keep covered mounts covered in OPEN_TREE_NAMESPACE. Creating such a mount namespace only takes a user namespace and the copy followed bind mount rules: no children without AT_RECURSIVE and no unbindable mounts with it. An unprivileged user could see what mounts covered in the source, such as the parts of /proc and /sys that container runtimes mask. If the caller doesn't own the source mount namespace a non-recursive copy of a mount with something mounted below the requested directory is now refused and a recursive copy includes unbindable mounts, the way unshare() copies. * Keep the lock on a mount that a propagated copy is moved beneath. MNT_LOCKED moved to any mount that ended up beneath a locked mount, propagated copies included. A host mount and umount on a directory covered by a locked mount in a less privileged mount namespace left that cover unlocked for the namespace's owner to remove. Only mounts the caller places beneath take over the lock now. * Handle mount locking for automounts correctly. Which copies to lock was decided by the mount namespace of the task that triggered the automount. A task in a user namespace that triggered one on a host mount through a file descriptor got the host's own automount locked while its own copy stayed unlocked and could have nosuid, nodev and noexec cleared. Use the owner of the mount namespace the mount lands in. - Use-after-free and crashes: * Refuse an automount below a mount that is in no namespace. The private clones overlayfs uses for its layers have the MNT_NS_INTERNAL error pointer as their namespace, which finish_automount() let through and count_mounts() dereferenced. A fanotify filesystem mark on an overlayfs lower layer hands out file descriptors on such a clone. With debugfs as the lower layer opening "tracing" oopses with namespace_sem held for writing and every mount operation on the system blocks from then on. * Reset the old parent's ->overmount in mnt_change_mountpoint(). When propagate_umount() moved an overmount off a mount that a file descriptor kept alive, MOVE_MOUNT_BENEATH through that descriptor later followed the stale pointer into the freed overmount. * statmount() with STATMOUNT_BY_FD and pivot_root() read the parent of a mount that may be unmounted and only held by a file descriptor, while the parent's final mntput() can free it. statmount() now reads it under mount_lock and pivot_root() first checks that both mounts are in the caller's mount namespace. * Queue a mount only once for mount notifications. A mount reparented by one umount_tree() and taken down by the next under the same namespace_sem hold, as in shrink_submounts(), was queued twice. That cut the mounts queued in between out of notify_list while it still pointed at them, and once they were freed every later mount operation walked freed memory. * Don't let a pseudo dentry become the root of a mount. A bind mount of a bpf token file did that with a DCACHE_NORCU dentry, which is freed without an RCU grace period while lockless path walks may still look at it. Refuse to clone such a mount. * Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily unmounted nsfs or pidfs mount through its file descriptor started out flagged as unmounted. Among other things __detach_mounts() then dropped the namespace's reference on it, the mount outlived its namespace and mount_setattr() through the descriptor read the freed namespace. A recursive bind mount of such a mount also copied the unmounted stack still attached to it. That now fails with EINVAL, copying the mount itself still works. * Remove the fsnotify marks of a mount namespace in free_mnt_ns() instead of the RCU callback that frees the namespace, where taking the group mutexes meant sleeping in softirq context. - Propagation and copies: * Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't copy the unbindable flag, so every mount namespace created with CLONE_NEWNS had bindable copies of all unbindable mounts. This had been fixed once before. * Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the mount an unbindable slave, a state nothing else can produce, or silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE after restoring sharing and isn't affected. * Check a recursive bind mount for mount namespace loops. Recursively bind mounting a tree from another mount namespace could put a mount of a namespace's file inside that same namespace, which then pins itself and all its mounts. Repeating it leaks without limit, the reproducer took Shmem from 380 kB to 65916 kB. The copy is now checked with check_for_nsfs_mounts() before it is grafted, as move_mount() does. * Look at the topmost mount for a mount namespace file. attach_recursive_mnt() never looked at the topmost mount of the source's chain of overmounts. If that was the chain's only mount namespace file an existing mount at a propagated destination got buried below the root of the nsfs file where no path walk reaches it. * Don't put a mountpoint on a dentry that's being removed. attach_recursive_mnt() makes a mountpoint of the source's root without its inode lock, so a racing rmdir() of that directory could leave a mount on it that nothing ever detaches. d_set_mounted() now checks cant_mount() as well. - nullfs: * Take no inode lock for readdir of an immutable directory. The root of every empty mount namespace is the same nullfs directory and iterate_dir() held its i_rwsem across ->iterate_shared(). A reader whose buffer faults on a FUSE mount of its own holds it for as long as its server wants, and with an exclusive locker queued behind it every lookup that misses the dcache, every create and every mount in that directory waits. One user of an empty mount namespace stalls all others. Directories with the new FOP_IMMUTABLE flag skip the lock. * Refuse to reconfigure internal superblocks through fspick(), MS_REMOUNT or the read-only remount that a synchronous umount() of the root does. For nullfs only root in the initial user namespace could do it, but the superblock is shared by every mount namespace and the flags showed up in statfs() for all of them. * Don't update the access time on nullfs and refuse F_SET_RW_HINT on an immutable inode. - unshare: Free an nsproxy that was never installed with nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped active references that were never taken, which triggered a warning and hid the caller's own namespaces from listns(). - Smaller changes: mount_setattr() checks the target before it walks the tree to allocate peer group ids, unshare() puts the old fs_struct before the old namespaces, dissolve_on_fput() drops the file's reference to the tree itself, disconnect_mount() is simplified and the documentation of the propagated unmount rule is brought up to date" * tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits) namespace: simplify disconnect_mount() selftests/filesystems: test covered mounts namespace: rework connected mounts nullfs: add an empty immutable regular file selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody selftests/filesystems: add a helper that holds a readdir in a page fault readdir: take no inode lock on an immutable directory nullfs: refuse file locks fsnotify: let a filesystem refuse marks on its objects namespace: nothing is mounted on or written through knullfs namespace: keep the private nullfs instance in knullfs fhandle: decide the subtree check under mount_lock selftests/filesystems: check that an automount below an overlay layer is refused selftests/filesystems: check the atime of the empty mount namespace root selftests/filesystems: check that a lock lands on the right mount and stays namespace: keep the lock on a mount that a propagated copy is moved beneath namespace: never expire a locked mount nullfs: don't update the access time namespace: handle mount locking for automounts correctly namespace: refuse an automount below a mount that is in no namespace ...
19 hoursMerge tag 'net-7.3-rc7' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Jakub Kicinski: "Including fixes from wireless, wireguard, CAN and Bluetooth. We have one known regression to wrap up in VLAN handling. Current release - regressions: - Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex Previous releases - regressions: - can: fix regression in handling RPS after migrating metadata to skb_ext - eth: - iavf: fix regressions in reconfig impacting bonding - mana: fix packet forwarding performance regression - stmmac: remove buggy VLAN acceleration support Previous releases - always broken: - a few high prio fixes for tun, and af_packet - amt: fix a UaF on tunnel teardown - eth: - bnxt: fix PCIe AER recovery and FLR handling issues - macb: don't modify Tx skbs before taking ownership - axienet: don't leak Tx skbs on interface stop - wifi: - nxpwifi: number of LLM-ish fixes - assorted mt76 fixes" * tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits) net: macb: copy shared skbs before appending the FCS net: macb: check TX ring before modifying skb vsock: Fix memory leak in vmci_transport_recv_dgram_cb() wireguard: noise: reject response consumption after intermediate initiation wireguard: queueing: preserve tstamp_type when encapsulating packet net: openvswitch: validate transport header presence in set_ipv6_addr net/smc: protect clcsock lifetime in smc_getname ipv6: do not warn on route notification size race ipv4: do not warn on route notification size race ipv4: validate checksum_start before completing checksum ptp: ocp: fix PCIe delay estimation calculation xen/netfront: don't leak the skb when xennet_fill_frags() fails net/packet: call packet_parse_headers after virtio_net_hdr_to_skb xen/netfront: drop RX packets with a short Ethernet header net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits() net: sparx5: free the matchall entry on destroy selftests: mlxsw: Test port range occupancy on template create mlxsw: spectrum_flower: Fix port range register leak in tmplt_create() net: dsa: microchip: fix KSZ8765 fiber detection net/mlx5e: Order ICOSQ cc update after CQ doorbell ...
30 hoursselftests: mlxsw: Test port range occupancy on template createPetr Machata
Add a test that creates a tc chain template matching on both source and destination port ranges and verifies via devlink-resource occupancy that this does not leak port range registers, neither while the template exists nor after it is deleted. Reviewed-by: Ido Schimmel <idosch@nvidia.com> Signed-off-by: Petr Machata <petrm@nvidia.com> Reviewed-by: Jacob Keller <jacob.e.keller@intel.com> Link: https://patch.msgid.link/e7b37b80adb7ac8b0ef20a22d7654a6656c1c545.1791294384.git.petrm@nvidia.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
30 hoursselftests: net: veth: test peer ndo-xmit after GRO toggle while downTianyi Gao
Test that toggling GRO on a veth device while it is down updates the peer's ndo-xmit XDP feature once the device comes up, for both GRO on and GRO off. xdp-features is only visible through netlink, so read it with the ynl CLI, as double_udp_encap.sh does. Skip the checks if the CLI is not found, as in an installed kselftest tree. Signed-off-by: Tianyi Gao <tianyi@cloudflare.com> Link: https://patch.msgid.link/20261006173241.65945-3-tianyi@cloudflare.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
37 hoursMerge tag 'mm-hotfixes-stable-2026-10-07-21-48' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull MM fixes from Andrew Morton: - Update .mailmap entries for Andy Yan and John Garry - Fix read-only MAP_SHARED /dev/zero mappings so they retain shared-file semantics instead of being treated as anonymous memory, also avoiding a CONFIG_DEBUG_VM assertion - Fix 32-bit build warnings in the hugetlb-mmap selftest caused by using the wrong printf format for size_t values - Fix a boot-time crash when early function tracing causes CPA to free kernel page tables before the workqueues used for deferred freeing are available - Fix two MREMAP_DONTUNMAP locked_vm accounting leaks: one caused by an mlock-on-fault VMA self-merging, and one caused by partially remapping a locked VMA * tag 'mm-hotfixes-stable-2026-10-07-21-48' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: mailmap: update entry for Andy Yan drivers/char/mem: mmap readonly MAP_SHARED-/dev/zero correctly selftests/mm: cleanup -Wformat issues in hugetlb-mmap mm: don't schedule deferred kernel page table freeing while booting mailmap: update addresses for John Garry mm/mremap: fix locked_vm leak by splitting VMA for MREMAP_DONTUNMAP mm/mremap: fix locked_vm leak from MREMAP_DONTUNMAP self-merge
3 daysselftests: tc-testing: add cls_route no-routing-attribute change testsVictor Nogueira
Add a boundary pair for route4 change without a routing attribute: - a classid-only change of a wildcard filter (handle 0xffff8000) is a metadata-only update and must succeed; - the same change on a non-wildcard handle (0x10001) would rekey the filter to the wildcard bucket, leaving the handle userspace stored stale, and must be rejected with -EINVAL. The change of a wildcard filter is the flow iproute2 uses by default ("tc filter add ... route classid X:Y" gives handle 0xffff8000). Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com> Signed-off-by: Victor Nogueira <victor@mojatatu.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/QDISC-0JB3.v1.20260923100921@mojatatu.com.2 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
3 daysMerge tag 'sched_ext-for-7.3-rc6-fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext Pull sched_ext fixes from Tejun Heo: - Taking a CPU offline could hang, or stall until the watchdog ejected the BPF scheduler, when tasks on the dying CPU were still held by the scheduler or sitting on a user dispatch queue. Re-enqueue them onto the local queue when the runqueue goes offline so that the CPU pushes them off like the other sched classes. - The sequence number guarding against stale dispatches was per runqueue, so a task re-enqueued on another CPU could get the same number and a dispatch meant for its earlier instance was applied to the new one. Use a per-task counter. - A task dispatched to another CPU's local queue got its ops.dequeue() only when picked to run and flagged as a core-sched pick. Call it at insertion like for same-CPU dispatches. * tag 'sched_ext-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: sched_ext: Generate qseq from a per-task counter selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ sched_ext: Fix CPU hotplug hang when a dying CPU's tasks sit in the BPF scheduler
3 daysselftests/mm: cleanup -Wformat issues in hugetlb-mmapCarlos Llamas
Commit ae571cd6015c ("selftests/mm: hugetlb-mmap: add setup of HugeTLB pages") and commit 9c5a65f374f8 ("selftests/mm: merge map_hugetlb into hugepage-mmap") added logs of 'hugepage_size' which has a size_t type. However, the incorrect format specifier '%lu' was used which triggers -Wformat warnings when building for 32-bit: hugetlb-mmap.c:125:55: warning: format specifies type 'unsigned long' but the argument has type 'size_t' (aka 'unsigned int') [-Wformat] 125 | ksft_print_msg("Default size hugepages (%lu kB)\n", hugepage_size >> 10); | ~~~ ^~~~~~~~~~~~~~~~~~~ | %zu hugetlb-mmap.c:134:47: warning: format specifies type 'unsigned long' but the argument has type 'size_t' (aka 'unsigned int') [-Wformat] 134 | ksft_exit_skip("Not enough %lu Kb pages\n", hugepage_size >> 10); | ~~~ ^~~~~~~~~~~~~~~~~~~ | %zu Fix this by switching to the expected '%zu' format specifier. Link: https://lore.kernel.org/20260927162419.820609-1-cmllamas@google.com Fixes: ae571cd6015c ("selftests/mm: hugetlb-mmap: add setup of HugeTLB pages") Fixes: 9c5a65f374f8 ("selftests/mm: merge map_hugetlb into hugepage-mmap") Signed-off-by: Carlos Llamas <cmllamas@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com> Reviewed-by: SJ Park <sj@kernel.org> Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Cc: Mike Rapoport <rppt@kernel.org> Cc: David Hildenbrand <david@kernel.org> Cc: Liam R. Howlett <liam@infradead.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Michal Hocko <mhocko@suse.com> Cc: Shuah Khan <shuah@kernel.org> Cc: <stable@vger.kernel.org>
3 daysmm/mremap: fix locked_vm leak from MREMAP_DONTUNMAP self-mergeLorenzo Stoakes (ARM)
Patch series "mm/mremap: fix two issues with MREMAP_DONTUNMAP", v2. The MREMAP_DONTUNMAP feature is highly unusual in that it permits mremap() operations that keep the original VMA in place. Historically this has led to a lot of bugs where non-obvious interactions occur between existing mremap() operations and the original VMA. Commit 397432cab17b ("mm/mremap: account mm->locked_vm correctly for MREMAP_DONTUNMAP") fixed an accidentally introduced bug around mm->locked_vm accounting, but this wasn't the only issue. And thus history repeats itself, as it turns out that mm->locked_vm accounting is broken by MREMAP_DONTUNMAP yet again by two further cases, and has been broken ever since the feature was introduced. Both relate to the fact that VMA_LOCKED_BIT is cleared on the source VMA (it has to be as all page tables are moved): 1. If an unfaulted VMA_LOCKONFAULT_BIT anonymous VMA self-merges it clears the VMA_LOCKED_BIT flag and permanently leaks mm->locked_vm pages. 2. If a partial mremap() is performed on a locked VMA there is a leak equal to the number of pages not copied. (Both for MREMAP_DONTUNMAP operations only) Both issues can be fixed by treating the source range as distinct from the destination range, which is the definition of what MREMAP_DONTUNMAP does so is appropriate. In case 1, simply disallow the self-merge, keeping adjacent source and destination VMAs distinct. In case 2, split the source range ahead of time if the VMA is mlock()'d, so accounting is always correct. Both changes were tested locally and confirmed to fix the issues. For the purposes of a backport, the fixes are kept distinct, a follow-up series can add self-tests. This patch (of 2): The MREMAP_DONTUNMAP feature is highly unusual in that it permits mremap() operations that keep the original VMA in place. Historically this has led to a lot of bugs where non-obvious interactions occur between existing mremap() operations and the original VMA. Fix another of these - self-merge. Self-merge occurs when a VMA is moved in front of or behind itself and the attributes of the VMA permit such a merge. Practically this can only happen for unfaulted anonymous VMAs due to the page offset equality requirement for merge: |------------| | | | v |...........||-----------||...........| | || unfaulted || | |...........||-----------||...........| ^ | | | |------------| This becomes problematic if the VMA is configured by the user to mlock-on-fault, i.e. the VMA_LOCKED_BIT, VMA_LOCKONFAULT_BIT VMA flags are set. MREMAP_DONTUNMAP clears mlock flags for the source VMA and maintains them for the destination VMA. Self-merge makes this impossible (there is only one VMA) and incorrectly clears the destination VMA's mlock flags. This causes a leak in mm->locked_vm as clearing this flag does not decrement the counter and the VMA no longer has VMA_LOCKED_BIT set so it is not decremented on unmap. Resolve this by simply disallowing a self-merge in this case - the source and destination VMAs are kept distinct and then are able to have distinct mlock() flags. Update dontunmap_complete() to make the now-redundant self-merge check a VM_WARN_ON_ONCE() instead to guard against future regressions. Also update the VMA userland tests to reflect the change. Link: https://lore.kernel.org/20260930-fix-dontunmap-partial-self-merge-v2-0-f388985a0f0a@kernel.org Link: https://lore.kernel.org/20260930-fix-dontunmap-partial-self-merge-v2-1-f388985a0f0a@kernel.org Fixes: e346b3813067 ("mm/mremap: add MREMAP_DONTUNMAP to mremap()") Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Reviewed-by: Pedro Falcato <pfalcato@suse.de> Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org> Reviewed-by: Jose A. Perez de Azpillaga <azpijr@gmail.com> Tested-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com> Cc: Liam R. Howlett <liam@infradead.org> Cc: Vlastimil Babka <vbabka@kernel.org> Cc: Jann Horn <jannh@google.com> Cc: Brian Geffon <bgeffon@google.com> Cc: Minchan Kim <minchan@kernel.org> Cc: <stable@vger.kernel.org>
3 daysselftests/filesystems: test covered mountsChristian Brauner
Test that mount cycles are resolved and test that mount covers behave as expected. Link: https://patch.msgid.link/20261002-work-mount-cover-v1-3-232a8f52b43c@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
3 daysselftests: net: tun: add test for VLAN-tagged GSO without NEEDS_CSUMEric Dumazet
Add selftests in tun.c verifying that a VLAN-tagged (802.1Q) TCPv4 GSO packet without VIRTIO_NET_HDR_F_NEEDS_CSUM (both flags = 0 and flags = VIRTIO_NET_HDR_F_DATA_VALID) is accepted when written to a TAP device (/dev/net/tun with IFF_TAP | IFF_NO_PI | IFF_VNET_HDR). Also verify that the following are rejected with -EINVAL: - a mismatched GSO type (VIRTIO_NET_HDR_GSO_TCPV6 on a VLAN-tagged IPv4 packet). - a frame whose TCP header is truncated after 10 bytes. Pulling only sizeof(struct iphdr) + sizeof(struct tcphdr) bytes would accept it, so this requires the transport offset found by flow dissection. Finally, verify that a 65540-byte frame is accepted. Its skb->len is above U16_MAX while it is flow-dissected, before eth_type_trans() pulls the Ethernet header. Based on a reproducer by Michael S. Tsirkin <mst@redhat.com>. Signed-off-by: Eric Dumazet <edumazet@kernel.org> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20261001191140.2818991-4-edumazet@kernel.org Signed-off-by: Jakub Kicinski <kuba@kernel.org>
6 daysselftests/filesystems: check that reading the root of an empty mount ↵Christian Brauner
namespace stalls nobody A readdir of the root of an empty mount namespace stuck in the page fault of its buffer must stall neither a create nor a lookup in that directory. The test needs userfaultfd and skips without it. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-21-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
6 daysselftests/filesystems: add a helper that holds a readdir in a page faultChristian Brauner
A getdents64() whose buffer is a page registered with userfaultfd sits in handle_userfault() with whatever iterate_dir() took before it copied the entries. Add readdir_hold.h for tests that want to know what that blocks: it opens the userfaultfd before the test enters a user namespace, the fault happens in the kernel and needs CAP_SYS_PTRACE in the initial one, starts the readdir in a thread, waits until that thread is stuck, queues a create behind it and waits for it to settle, then probes a lookup with a watchdog and says whether it came back. The page is released afterwards so that everything drains on a kernel that still holds the lock across the copy. The two users follow. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-20-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
6 daysselftests/filesystems: check that an automount below an overlay layer is refusedChristian Brauner
Make sure that we don't automount on top of internal things. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-12-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
6 daysselftests/filesystems: check the atime of the empty mount namespace rootChristian Brauner
No access time updates for immutable nullfs. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-11-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
6 daysselftests/filesystems: check that a lock lands on the right mount and staysChristian Brauner
Check MNT_LOCKED: - A child in a user namespace of its own triggers the tracefs automount through a directory descriptor on the host's debugfs mount. The copy in the child's namespace has to be locked and the host's mount unlocked. - A child moves a bind of that automount beneath a locked covering mount, unmounts the covering mount and asks for the umount of an unlocked ancestor. This must not expire the mount which holds the lock now. - The host mounts and unmounts on a directory that a locked mount covers in a user namespace further down the propagation chain. That cover has to be locked afterwards as it was before. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-10-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
6 daysselftests/filesystems: check that an immutable inode takes no write hintChristian Brauner
Check that an immutable inode takes no write hint. Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-4-dd44b89d44ce@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
7 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: - Fix overflow of backward jump offset in constant blinding (Alexei Starovoitov) - Fix packet range of packet pointers sharing an id when var_off tightens umax of one pointer and not the other (Alexei Starovoitov) - Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc (Alexei Starovoitov) - Fix use-after-free of progs detached from busy trampolines: wait for an RCU tasks grace period before freeing trampoline progs, and patch detached progs out of trampoline images that are still in use (Florent Revest) - Hold map BTF for the memory allocator destructor record to fix UAF in deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi) - Fix missing migration protection in resizable hashtab lookup_and_delete batch operation (Ömer Mete Kaya) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch() selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace bpf: Fix objects stuck in free_by_rcu_ttrace bpf: Factor out __do_call_rcu_ttrace() selftests/bpf: Test packet range of pointers sharing an id bpf: Fix packet range of pointers sharing an id selftests/bpf: Detach a trampoline prog while a task sleeps before it bpf: Skip detached progs in trampoline images that are still in use bpf: Wait for an RCU tasks grace period before freeing trampoline progs bpf: Hold map BTF for the memory allocator destructor record bpf: Fix overflow of jump offset in constant blinding
7 daysMerge tag 'block-7.3-20261002' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux Pull block fixes from Jens Axboe: - NVMe fixes via Keith: - Fix an out-of-bounds write in nvmet_auth_challenge(), where sizeof() on a void pointer undercounted the challenge header and let a short AUTH_RECEIVE buffer pass the check - nvme-multipath fixes for an ANA log bounds check underflow, the command effects log lifetime for multipath heads, and only setting BLK_FEAT_ZONED after the zone info is known. - nvmet fixes for ns->enabled teardown ordering, rejecting I/O after the percpu ns reference is killed, device path preservation on allocation failure, and too-short SGL segments in pci-epf - nvme-tcp: revert the per-socket dynamic lockdep keys, and delay the socket reclassification - A DMA pool alignment quirk for the Micron 4100AT - Controller state/reset race fixes, and -Wformat-security workarounds - blk-mq: set RQF_USE_SCHED when the operation is known, and allow cached requests to be used for flush operations - Reject polled dio with user integrity metadata - Save the IRQ state in blkg_tryget_closest() - Set the zone write granularity in virtio_blk - ublk selftest fixes * tag 'block-7.3-20261002' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux: (23 commits) virtio_blk: set the zone write granularity nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known nvme: fix command effects log lifetime for multipath heads nvmet: don't allow I/O admission after percpu ns reference is killed nvmet: defer setting ns->enabled to false in nvmet_ns_disable() nvmet: copy the hostid into the ctrl before creating PR pc_refs nvmet-auth: fix out-of-bounds write in nvmet_auth_challenge() nvmet: pci-epf: reject too-short SGL segments nvme-multipath: fix underflow in ANA log bounds checks nvme: work around all -Wformat-security warnings nvme: work around -Wformat-security warning nvme: do not reset controllers in NVME_CTRL_NEW state nvme-tcp: delay nvme_tcp_reclassify_socket() Revert "nvme-tcp: lockdep: use dynamic lockdep keys per socket instance" drbd: remove unused drbd_nl_mcgrps[] array blk-mq: allow cached requests to be used for flush operations blk-mq: set RQF_USE_SCHED when the operation is known block: reject polled dio with user integrity metadata selftests: ublk: fix unused_result error blk-cgroup: save IRQ state in blkg_tryget_closest() ...
7 daysMerge tag 'hid-for-linus-2026100201' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid Pull HID fixes from Benjamin Tissoires: - Revert of the Bolt integration into hid-logitech-dj (Benjamin Tissoires) - A couple of buffer overflow in Intel-thc-hid (Even Xu) - A couple of Sashiko findings fixes in hid-multitouch and HID-BPF (Aldo Ariel Panzardo and Benjamin Tissoires) * tag 'hid-for-linus-2026100201' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid: selftest/hid: add test for negative return codes for hid_bpf_hw_request HID: bpf: cast size to ssize_t when checking hid_bpf_hw_request HID: Intel-thc-hid: Intel-quickspi: Fix buffer overflow HID: Intel-thc-hid: Intel-quicki2c: Fix buffer overflow HID: universal-pidff: Add support for Turtle Beach VelocityOne Race HID: multitouch: stop the release timer from being rearmed on remove Revert "HID: logitech: add Bolt receiver support for Logitech HID++ devices"
8 daysselftests: net: amt: check that the relay's queries bypass the amt deviceOmar Ramadan
The relay used to hand its General Queries to dev_queue_xmit() on the amt device, where a query could wait in a qdisc and outlive the tunnel it pointed to. The previous patch sends them directly from the receive path instead. Count the IGMP and MLD queries that leave the relay through amtr with tc flower filters on its egress, installed before the gateway comes up, and check that there are none. The forwarding tests before it already show that the gateway received its queries, since it cannot join without one. Without the previous patch the new test fails (one run counted 7 IGMP and 6 MLD queries); with it, all of amt.sh passes. Signed-off-by: Omar Ramadan <omar@blockcast.net> Link: https://patch.msgid.link/20260928202312.74574-3-omar@blockcast.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
8 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "The most intrusive change is reverting a commit from 7.3-rc1 that made struct kvm a bit too large, and fixing the same issue otherwise. There are again a lot of selftests lines; the sheer number of commits is not small but I don't expect much more for 7.3 due to people travelling to Plumbers next week. ARM: - Take a reference on the last IRQ loaded into an LR to prevent it from being freed while running the guest (Marc Zyngier) - Ensure that the ITS MOVALL command only affects LPIs that were previously affined to the source redistributor (Marc Zyngier) - Fix + test for honoring the host's trap configuration when running non-protected VMs while KVM is in protected mode (Fuad Tabba) - Use the host stage-1 mapping granularity for VM_PFNMAP mappings at stage-2 (Mostafa Saleh) x86: Various bugfixes where the guest could do stupid things on purpose to cause problems in the host: - Failed VMRUNs can cause pending TLB flushes to be dropped, and in general some actions done through VMCB control fields have to be redone if VMRUN fails - Toggling MSR interceptions or eVMCS execution controls can cause the host to use a stale MSR permission bitmap - Bad page tables can cause a WARN. Also fix issues in last week's pull request (my fault, for changing email workflow and thus missing feedback sent to kvm@ but not LKML). Generic: - Take kvm_lock when creating vCPUs. For almost two decades everybody thought it was not done for some unspecified performance reasons, but in reality it was only done because kvm_lock was originally a spinlock. This is a better fix than 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible", from the 7.3 merge window), and does not waste 2K per VM, hence its inclusion here" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits) KVM: arm64: Use stage-1 leaf size for VM_PFNMAP KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt KVM: selftests: Verify that L0's TPR doesn't get clobbered KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded ...
8 daysMerge tag 'kvmarm-fixes-7.3-2' of ↵Paolo Bonzini
https://git.kernel.org/pub/scm/linux/kernel/git/kvmarm/kvmarm into HEAD KVM/arm64 fixes for 7.3, round #2 - Take a reference on the last IRQ loaded into an LR to prevent it from being freed while running the guest (Marc Zyngier) - Ensure that the ITS MOVALL command only affects LPIs that were previously affined to the source redistributor (Marc Zyngier) - Fix + test for honoring the host's trap configuration when running non-protected VMs while KVM is in protected mode (Fuad Tabba) - Use the host stage-1 mapping granularity for VM_PFNMAP mappings at stage-2 (Mostafa Saleh)
8 daysselftests/bpf: Add a test for objects stuck in free_by_rcu_ttraceAlexei Starovoitov
Delete all elements of BPF_F_NO_PREALLOC hash map in one batch. The first free_bulk() starts RCU tasks trace GP and the rest of the elements are freed while it's in flight. Wait for call_rcu_ttrace_in_progress to clear in bpf_mem_cache of every cpu and check that free_by_rcu_ttrace and waiting_for_gp_ttrace lists are empty. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20260930095920.601738-4-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysselftests/bpf: Test packet range of pointers sharing an idAlexei Starovoitov
Add tests where two packet pointers share an id and tightening one pointer's umax from its var_off would put it less than their constant distance from the other's umax: with an index & 0x38 capped at 50, the base pointer keeps umax 50, so the pointer 8 bytes further on must keep umax 58, even though its known bits allow at most 56. These refused a valid program or accepted an out-of-bounds access before the fix: - check the advanced copy, load through the base: valid, was refused; - check the base, load the byte at base + 1 through a copy advanced by 8: was accepted; - check base + 4, load 4 bytes at base + 2 through base + 8: reads two bytes past the checked range, was accepted; - the same as the second with data_meta pointers checked against data: was accepted. These pass with and without the fix and cover nearby paths: - subtract an unknown scalar from a checked pointer and load below it (the range is kept across a new id); - reach a load through two paths whose checks cover 8 and 7 bytes after the loaded pointer; the second path must not be pruned by the first; - spill a copy of a pointer, check the pointer, fill the copy and load one byte past the checked range: the load is refused, and the copy has the range of the check. Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261001145255.855630-2-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysbpf: Fix packet range of pointers sharing an idAlexei Starovoitov
Since commit 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers"), find_good_pkt_pointers() sets the range of all packet pointers sharing an id from the umax of the compared pointer, and check_packet_access() requires umax + off + size <= range. That assumes the umax of two such pointers differ by exactly their constant distance. reg_bounds_sync() breaks it when var_off tightens one umax and not the other: r4 &= 0x38 if r4 > 50 goto exit ; umax 50, var_off (0x0; 0x38) r5 = pkt + r4 ; umax 50 r6 = r5 r6 += 8 ; umax 56, not 58 Comparing r6 with pkt_end sets the range to 56, and the valid 8-byte load at r5 is rejected (50 + 8 > 56). Comparing r5 sets it to 50, and the out-of-bounds 1-byte load at r6 - 7, i.e. r5 + 1, is accepted (56 - 7 + 1 <= 50). Don't call reg_bounds_sync() on a packet pointer that keeps its id (a constant was added or subtracted) or its range (an unknown non-negative value was subtracted), so that var_off cannot tighten its umax. Only update the 32-bit bounds from var_off: reg_bounds_sanity_check() wants them constant when the lower half of var_off is, e.g. for pkt + 8. This relies on nothing else changing the 64-bit bounds of a packet pointer, which holds today. var_off of such a pointer is no longer narrowed by its bounds. Adjust three verifier_align expectations; the low bits, which the alignment checks use, don't change. veristat on the selftests shows no verdict changes and +0.8% insns in test_cls_redirect_subprogs. Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers") Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://lore.kernel.org/bpf/20261001145255.855630-1-alexei.starovoitov@gmail.com Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
8 daysMerge tag 'net-7.3-rc6' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net Pull networking fixes from Paolo Abeni: "Including fixes from Bluetooth, WiFi and netfilter. We are actively retargeting several non-urgent fixes towards next, but the traffic on the ML looks ever-increasing, and propagating the push-back towards subsystems is not immediate. No known outstanding regressions. Current release - regressions: - netfilter: nft_set_rbtree: skip transaction elements during GC Previous releases - regressions: - sched: cls_api: reclaim an empty proto on the error path - core: - fix checksum offsets in skb_splice_from_iter() - cap skb->queue_mapping when the tx queue is picked - page_pool: fix use-after-free in page_pool_recycle_ring_bulk() - wifi: - mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue() - mac80211: drop oversized fragments to avoid extra_len overflow - netfilter: - flowtable: restore ieee80211 forward path - bluetooth: hci_conn: Lock parent access during enhanced SCO setup - eth: - bcmgenet: allocate RX buffers as page fragments - stmmac: fix rx Scatter-Gather support - octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled - gve: DQO: accept TSO packets with non-protocol gso_type bits - r8169: disable EEE on RTL8168h/8111h Previous releases - always broken: - tcp: refresh TS.Recent for accepted old ACKs - wifi: - ath11k: reset ar->num_stations on hardware start - cfg80211: fix RTS threshold setting for single-radio PHY - bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame() - eth: bcmgenet: fix NULL dereference in set_coalesce before first open Misc: - Eric is retiring from google and updating his contact info" * tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits) net: phy: aquantia: fix system interface type not updated in forced mode net: usb: qmi_wwan: add Rolling Wireless RN947R net: mvneta: clear XDP pfmemalloc flag between frames ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel() net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference r8169: disable EEE on RTL8168h/8111h octeontx2-pf: Fix RSS indirection table size sctp: check RCV_SHUTDOWN after the sendmsg connect wait net: sparx5: make ports inherit the switch base mac address type net: microchip: vcap: stop scanning after deleting key field netfilter: flowtable: restore ieee80211 forward path netfilter: flowtable: generalize pending status bit netfilter: bpf: reject invalid NAT manipulation types netfilter: nft_set_rbtree: skip transaction elements during GC ipvs: filter some flags received in the backup server ipvs: do not create invisible templates ipvs: bound LBLCR and LBLC cache growth ipvs: fix missing counter decrement in lblc netfilter: nft_flow_offload: drop flowtable reference on init error path selftests: net: check timestamp echo after an old ACK ...
8 daysselftests/filesystems: check that the nullfs root can't be reconfiguredChristian Brauner
Add a test for the root of an empty mount namespace: - fspick() of the root fails with EINVAL - mount(MS_REMOUNT) of the root fails with EINVAL - umount() of the root fails and doesn't remount it read-only Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-12-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysselftests/filesystems: check that a mount namespace file on top doesn't bury ↵Christian Brauner
a mount Add a test for moving a mount with a mount namespace file on top of it: - S, a bind mount of a file, with a newer mount namespace's file on top - A shared with a slave B that has Q on B/file - S moved onto A/file, its copy lands on B/file below Q B/file keeps reading Q and once Q is unmounted it reads the copy. Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-9-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysselftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts coveredChristian Brauner
Add a test for open_tree(OPEN_TREE_NAMESPACE) from a user namespace that doesn't own the mount namespace it copies from: - a file covered by a private or an unbindable tmpfs mount - the non-recursive copy of the parent mount fails with EINVAL - the recursive copy keeps the file covered in the new mount namespace - the owner of the mount namespace keeps the bind mount semantics Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-7-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysselftests/filesystems: check that a recursive bind mount can't pin the ↵Christian Brauner
caller's mount namespace Add a test for a recursive bind mount of a namespace file in another mount namespace: - a bind mount of the network namespace file, held through a descriptor - the caller's own mount namespace file stacked on top of it there - the recursive bind mount through the descriptor fails with EINVAL - a plain bind mount of the file still works The copy would pin the namespace it is put in otherwise. Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-5-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
8 daysselftests/filesystems: check that a busy submount survives a synchronous umountChristian Brauner
Add a test for the shrinkable submounts of a synchronous umount: - P shared, P1 a slave that is shared in turn, P2 its peer - B on P/options with copies on P1 and P2, R on top of the copy in P2 - P with B moved below P1, R is the working directory - umount(P1) slides R to where the copy of B is looked up The umount fails with EBUSY and R stays mounted. Needs the tracefs automount below debugfs for the shrinkable mounts. Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-3-be34c83956ae@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
9 daysselftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ movesKuba Piecuch
Add a dequeue_remote test that makes moves of tasks in the BPF scheduler's custody to another CPU's local DSQ the common case, both via SCX_DSQ_LOCAL_ON dispatch and via scx_bpf_dsq_move_to_local(). The BPF scheduler tracks each task's custody state and triggers scx_bpf_error() if a custody period doesn't end with exactly one ops.dequeue() before the task runs, or if it ends with an SCX_DEQ_CORE_SCHED_EXEC dequeue of a task without a core cookie. Without the previous patch, the test fails with: sched_ext: dequeue_remote: dequeue_remote.bpf.c:141: 15 (rcu_preempt): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=3 cpu=2 seq=1) ... ops_dequeue+0x114/0x170 set_next_task_scx+0x104/0x1e0 __pick_next_task+0xc7/0x180 __schedule+0x154/0x1870 Reviewed-by: Andrea Righi <arighi@nvidia.com> Assisted-by: Claude:claude-opus-5.5 Signed-off-by: Kuba Piecuch <jpiecuch@google.com> Signed-off-by: Tejun Heo <tj@kernel.org>
10 daysselftests: net: check timestamp echo after an old ACKJeff Jo
Add a regression test for a gap-filling packet whose acknowledgment has become old. Linux first sends data to the peer. Deliver two peer data packets out of order: the later packet acknowledges Linux's data, while the delayed packet still carries the earlier acknowledgment. Require the ACK that closes the receive gap to echo the delayed packet's timestamp, 301000. Without the fix, Linux accepts the data but still echoes the previously saved timestamp, 1000. Also check that the application can read all 34 bytes. Use a large jump in peer timestamps to represent the idle interval, with no real wait, loss or retransmission. The test directly checks the outgoing timestamp echo. It requires packetdrill's merged TSecr verification fix (linked below); older tools incorrectly pass on an unfixed kernel. Use the existing packetdrill selftest runner for IPv4, IPv6 and IPv4-mapped IPv6. Link: https://github.com/google/packetdrill/commit/83f72d3f9085d0e26eb4d206fe4d7cfab5b6d872 Assisted-by: LLM Signed-off-by: Jeff Jo <jeffjo@openai.com> Reviewed-by: Eric Dumazet <edumazet@google.com> Signed-off-by: David S. Miller <davem@davemloft.net>
10 daysselftests: net: cover UDP splice checksum fragment alignmentAlireza Asgari
Exercise software UDP checksums when splice() transfers multiple pipe fragments into a corked datagram. Use separate source pipes so short writes cannot merge, and use odd fragment lengths to expose checksum-position errors between fragments. Cover IPv4, IPv6 and IPv4-mapped destinations, connected and unconnected sockets, resident prefixes, even and odd fragment lengths, and a fragment starting near a page boundary. Ordinary sends provide a control. Check the complete payload and reject extra datagrams after a successful send. Keep receive waits bounded and skip unavailable address families or pipe capacities. The test uses only loopback sockets with ephemeral ports and has no dependency on the application that exposed the regression. Signed-off-by: Alireza Asgari <alireza@asgari.net> Link: https://patch.msgid.link/20260924-fix-udp-splice-checksum-v1-2-fe61d65447a8@asgari.net Signed-off-by: Jakub Kicinski <kuba@kernel.org>
10 daysKVM: arm64: selftests: Check a feature hidden in an ID register is UNDEFFuad Tabba
Userspace can hide a feature from a guest by clearing its field in a writable ID register, and KVM then makes the feature's instructions UNDEFINED in the guest by trapping or disabling them. No selftest checks that. Add a test that runs the instruction of each of TLBI OS, MOPS, TCR2_EL1 and FPMR once with its field as advertised and once with it cleared, and expects an UNDEF only when cleared. A feature the vCPU doesn't advertise is skipped, as is hidden TLBI OS on a CPU with neither FGT nor FEAT_EVT2, where KVM can't trap it. Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev> Link: https://patch.msgid.link/20260929090031.3829185-5-fuad.tabba@linux.dev Signed-off-by: Oliver Upton <oupton@kernel.org>
11 daysselftests/bpf: Detach a trampoline prog while a task sleeps before itFlorent Revest (Anthropic)
Add a test for the use-after-free fixed by the previous commit. A sleepable prog on bpf_fentry_test1() blocks a task on a userfaultfd page, like bpf_mod_race does, while the two progs that run after it in the same trampoline are detached and freed one after the other, which the test knows from their .bss maps going away. After the first detach the task is in an image that isn't the trampoline's current one anymore. The task is then released and must not call into the freed progs. This is done with fentry progs, with fexit progs, where the task is already past the original function, and with a task sleeping in a fentry prog while fexit progs are detached, where the original function must still be called. bpf_mod_race's userfaultfd helper moves to testing_helpers.c so that both tests can use it. Without the fix, on a kernel with KASAN: BUG: KASAN: vmalloc-out-of-bounds in __bpf_prog_enter_recur+0xed/0x1e0 Read of size 8 at addr ffa0000000144040 by task test_progs/171 CPU: 6 UID: 0 PID: 171 Comm: test_progs Tainted: G OE 7.2.0+ #1 PREEMPT(full) Call Trace: <TASK> __bpf_prog_enter_recur+0xed/0x1e0 bpf_trampoline_6442545468+0x72/0xe3 bpf_fentry_test1+0x9/0x20 bpf_prog_test_run_tracing+0x183/0x3e0 __sys_bpf+0xd3f/0x38f0 ... Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260926135605.1217928-4-florent.revest@linux.dev
11 daysKVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12Sean Christopherson
Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260813223610.2043560-7-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
11 daysKVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virtSean Christopherson
Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260813223610.2043560-6-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
12 daysMerge tag 'cgroup-for-7.3-rc4-fixes-2' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup Pull cgroup fix from Tejun Heo: - A cpuset partition could claim CPUs an ancestor partition already held exclusively. Restore the rejection an earlier change had turned into a warning. * tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup: cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
12 daysMerge tag 'sched_ext-for-7.3-rc4-fixes-2' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext Pull sched_ext fix from Tejun Heo: - The CPU topology helper for BPF schedulers took no buffer size, so its structure couldn't grow without breaking schedulers built against the older layout. Add a size argument. * tag 'sched_ext-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: sched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo can grow
13 dayssched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo ↵Tejun Heo
can grow scx_bpf_cid_topo() copies struct scx_cid_topo into a buffer the BPF program sized from its own vmlinux.h while the verifier sizes the write from the running kernel's BTF. The struct may grow and each growth then breaks every scheduler built against the older layout, rejected at load or written past its buffer. This is the usual hole for a struct handed to BPF, closed elsewhere with a size argument, and it was missed here. Take the buffer size, copy the smaller of it and the kernel's struct and set the rest to -1. Accesses to the copy are CO-RE relocated, so the struct can grow by appending fields, which its comment now states. The kfunc changes in place: the cid interface is still being finalized and no released scheduler uses the current form. Fixes: e9b55af47edf ("sched_ext: Add topological CPU IDs (cids)") Cc: stable@vger.kernel.org # v7.2+ Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
13 dayscgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on ↵Hui Peng
subpartitions_cpus conflict When a remote partition is created underneath an existing local partition via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees parent->partition_root_state == PRS_MEMBER and calls remote_partition_enable(). Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus, subpartitions_cpus) error check in remote_partition_enable() with WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning and proceeds to enable the remote partition on CPUs that are already owned by the ancestor local partition in subpartitions_cpus. This can be reproduced on Linux 7.3.0-rc3 with: mkdir -p /tmp/cg1 mount -t cgroup2 none /tmp/cg1 echo "+cpuset" > /tmp/cg1/cgroup.subtree_control mkdir /tmp/cg1/A echo 1 > /tmp/cg1/A/cpuset.cpus echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive echo root > /tmp/cg1/A/cpuset.cpus.partition echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control mkdir /tmp/cg1/A/B echo 1 > /tmp/cg1/A/B/cpuset.cpus echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control mkdir /tmp/cg1/A/B/D echo 1 > /tmp/cg1/A/B/D/cpuset.cpus echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition which triggers: WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300 and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions claiming exclusive CPU 1. Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects subpartitions_cpus in remote_partition_enable(), matching the error code used by remote_cpus_update() for the same subpartitions_cpus conflict, and add a regression test case to tools/testing/selftests/cgroup/test_cpuset_prs.sh. Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and tools/testing/selftests/cgroup/test_cpuset_prs.sh. Fixes: 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs") Suggested-by: Guopeng Zhang <guopeng.zhang@linux.dev> Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Hui Peng <benquike@gmail.com> Reviewed-by: Waiman Long <longman@redhat.com> Signed-off-by: Tejun Heo <tj@kernel.org>
13 daysMerge tag 'probes-fixes-v7.3-rc4' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull probe fixes from Masami Hiramatsu: - kprobes: Fix permanent hang when flushing the kprobe optimizer Fix a deadlock when disabling kprobe optimization via sysctl or debugfs where flushers hung waiting for optimizer_completion. Replaced the completion with an optimizer_passes counter and wait_var_event_mutex() under kprobe_mutex so concurrent flushers can wait and wake up safely. - fprobe: Terminate the fgraph_data list when the reservation is not filled Fix an issue where unused shadow stack data left uninitialized by fprobe_fgraph_entry() was misparsed as stale fprobe headers on return. Explicitly write a zero word to terminate the list and update read_fprobe_header() to handle the zeroed slot properly. - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc Fix false test failures in kprobe_non_uniq_symbol.tc on architectures like s390 where a symbol exists once in core kernel but also in modules. Anchor the /proc/kallsyms search regex to the end of the line so that module symbols are not incorrectly counted. * tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: kprobes: Fix permanent hang when flushing the kprobe optimizer fprobe: Terminate the fgraph_data list when the reservation is not filled selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
13 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "Arm: - Invalidate the ITS translation cache when the guest changes the base address of the ITS tables (Fuad Tabba) - Skip saving ITS devices with device IDs that are out-of-bounds rather than failing the entire ITS save ioctl (Fuad Tabba) - Close race between VM teardown and invalidations of nested MMUs when handling MMU operations that are allowed to block (Lorenzo Stoakes) - Various fixes for the handling of the host's untrusted SVE configuration in pKVM (Fuad Tabba) - Make sure that empty SMCCC ranges based at 0 are rejected by the kvm_smccc_set_filter() (Karl Mehltretter) - Revoke the host mapping for pKVM's private stack pages, along with a new sanity check that all mappings in the hyp's private VA range have been correctly marked as hyp-owned (Fuad Tabba) - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that concurrent vCPU initialization cannot relocate in-use MMUs. Defer the freeing of shadow stage-2 MMUs to the point that no other users (e.g. MMU notifier) could reference them (Marc Zyngier) - Drop useless WARN when rejecting an unsupported ioctl for pKVM (Fuad Tabba) - Fix the steal_time selftest to install correctly-sized mappings for non-4K hosts (Sebastian Ott) - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark Brown) - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit guests (Karl Mehltretter) RISC-V: - Synchronize hrtimer during VCPU teardown - Fix the conversion between vsip and hvip values - Serialize IMSIC attributes with vCPU migration - Release unused page after MMU invalidation - Propagate interrupted G-stage faults to KVM user-space as EINTR - Fix nested acceleration hfence entry update order - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem - Preserve firmware counter value across PMU counter stop/start - Report PMU snapshot write failure to the guest - Fix perf-backed counter accounting across PMU stop and read - Correctly propagate error of a hart status SBI call s390: - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty the pages that contain indicator and summary bits - Fix compile warning for kvm_s390_update_cmma_dirty() - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to userspace - Move s390_kvm_mmu_commit_memory_region() into s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN - Add missing srcu in kvm_s390_set_irq_state() - Fix potential races in storage functions - Fix race in _destroy_pages_crste() - Fix issues in the handling of KVM interrupt and page resources, when a queue that is assigned to a mediated device (mdev) is removed from the host's AP configuration - Fix loop condition in uv_find_secrets - Prevent potential out-of-bounds read x86: - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl) - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely) - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior) - Fix a memory leak and a cache maintenance issue related to doing intra-host migration on an SEV guest" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits) KVM: SEV: Do cache maintenance on the source VM during intra-host migration KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg KVM: Don't treat reserved xarray entries as having memory attributes KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails KVM: arm64: Fix AArch32 DBGBXVR<n> handling KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP KVM: selftests: fix steal_time for arm64 with host page size > 4K KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction KVM: arm64: nv: Fix life cycle of the nested_mmus array KVM: arm64: Check every private mapping is hyp-owned at pKVM init KVM: arm64: Move the private VA allocation cursor to __io_map_next KVM: arm64: Match hyp text by physical address in fix_host_ownership() KVM: arm64: Transfer the hyp stack pages out of the host stage-2 KVM: arm64: selftests: Test empty SMCCC filter range at base 0 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 ...
13 daysselftests/filesystems: check that a dead mount's overmount isn't followedChristian Brauner
Add overmount_reparent_test: - a mount that a file descriptor keeps alive after propagate_umount() reparented its overmount and unmounted it must not be walked into the freed overmount by move_mount(MOVE_MOUNT_BENEATH) Link: https://patch.msgid.link/20260926-work-mount-fixes-2-v1-10-f357abf3d17b@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
13 daysselftests/filesystems: check that an unmounted mount tree isn't walkedChristian Brauner
Add unmounted_tree_test: - a detached nsfs or pidfs bind mount that a file descriptor keeps alive can't be copied recursively, a copy of the mount itself and a copy of a mounted one still work - a detached tree with a child that was vacated behind it can't be copied recursively either - a copy of a detached bind mount is an ordinary mount: unlinking its mountpoint from another mount namespace unmounts it - mount_setattr() refuses to change the propagation of a detached tree and still allows it for the root of an anonymous mount namespace Every case runs in a user and mount namespace of its own. Link: https://patch.msgid.link/20260926-work-mount-fixes-2-v1-9-f357abf3d17b@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
14 daysKVM: selftests: Verify that L0's TPR doesn't get clobberedSean Christopherson
Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260813223610.2043560-5-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: selftests: Run the nested x2APIC with and without APICv being inhibited ↵Sean Christopherson
in L2 Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260813223610.2043560-4-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: selftests: Add x2APIC MSR test for inhibiting APICv while nestedSean Christopherson
Add a selftest to verify that KVM intercepts x2APIC MSR accesses for L1 after APICv is inhibited while L2 is active. This is a regression test for an AVIC bug where KVM would skip updating x2APIC MSR intercepts while L2 is active, thus giving L1 access to a wide swath of L0's x2APIC surface. Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260710162052.2188574-3-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>