| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
"This contains fixes for the current development cycle.
All of them came out of a review of the mount code that started with a
bug report. The review modeled the corner cases of mount propagation,
unmounting and mount reference counting and turned up a lot of bugs.
Most of them years old. Most fixes come with a selftest.
- Rework connected mounts.
A mount that is unmounted together with its parent can stay
attached to the parent to keep its mountpoint covered. That happens
when the mountpoint is removed with rmdir(), unlink() or rename(),
when a detached tree is dissolved, and for locked mounts in any
umount that isn't synchronous, including the teardown of their
mount namespace. The parent then owns the child and drops it on its
own final mntput(). So any reference from the child's superblock
back to one of its ancestors becomes a cycle that is never freed.
A loop device backed by an image on a tmpfs and mounted on that
same tmpfs is enough. Remove the directory the tmpfs is mounted on
from the host, let the container's mount namespace exit, and the
loop device, the tmpfs and the filesystem on the loop device are
leaked for good. The same works with autofs, zram, ecryptfs,
binfmt_misc, fuse passthrough, zloop, a mass storage gadget and md,
and the selftests have reproducers for them. This has been possible
since v4.1. It is also why "put_mnt_ns(): leave mounts connected"
was reverted in -rc5. Keeping every mount of a dying mount
namespace connected made these cycles trivial to create.
Every unmounted mount is now detached from its parent. Where the
mountpoint has to stay covered the mount leaves a cover on the
parent instead, allocated together with the mount. A lookup on the
unmounted parent that hits a cover finds an empty immutable
directory or file on the private nullfs instance. Nothing leads
from a cover to another mount, so no unmounted mount owns another
one and no cycle can form.
This is visible to userspace. A formerly connected mount can no
longer be reached through its unmounted parent and ".." inside it
leads nowhere, as for every other lazily unmounted mount.
With that the private nullfs instance becomes reachable from
userspace, so it now refuses mounts on top, is mounted read-only,
and refuses fsnotify marks and file locks. Its inodes are shared by
every holder and a watch or a lock would otherwise reach across
users. may_decode_fh() now also decides its subtree check under a
single mount_lock hold, as a racing umount could otherwise let it
decode into what a locked child covered.
- umount:
* Don't silently unmount busy mounts.
Since v4.13 propagate_umount() takes down propagated copies of
the victim with children as long as each child is an overmount or
another copy of the victim, but propagate_mount_busy() only ever
checked copies without children or with just an overmount. A
container that moved a tree beneath its copy of a host mount lost
that tree from under its open file descriptor to a plain umount()
on the host. propagate_mount_busy() now applies the same rules,
walking each chain of copies once.
* Don't let a migrating task hide its reference from umount().
mnt_get_count() sums the per-cpu counters under mount_lock but
the mntget() and mntput() fast paths don't take it. A task that
takes a reference on a cpu the sum has already passed and drops
it after migrating to one the sum hasn't reached yet hides the
reference it held to begin with, and umount() succeeds with the
file still open. Gets and puts now live in separate per-cpu
counters and all puts are summed before all gets with a full
barrier in between, the way srcu_readers_active_idx_check() does
it. mntget() is unchanged and mntput() gains an smp_wmb().
* Check each submount for references right before unmounting it.
shrink_submounts() and mark_mounts_for_expiry() checked all their
victims up front. Unmounting the first could move a busy
overmount to where the next victim's propagated copy is looked up
and it was then unmounted without a check.
* Never expire a locked mount.
A shrinkable mount moved beneath a locked mount with
MOVE_MOUNT_BENEATH takes over the lock, and umount() of an
unlocked ancestor expired it and revealed what it covered. That
umount() now fails with EBUSY as it does for any other locked
child. A lazy umount still takes the whole tree.
- Overmounts and locked mounts:
* Unhash a dentry before detaching the mounts on it.
unlink(), rmdir() and rename() detach the mounts on the victim
but only d_delete() it once its inode is unlocked, a window that
includes an expedited RCU grace period. In between, a lookup from
a mount namespace in which the dentry is a mountpoint found it
hashed, positive and uncovered. Drop the dentry first, as
d_invalidate() already does.
* Don't reveal overmounted entries in refwalk.
A refwalk that had grabbed the dentry before the unlink never
rechecked it the way rcuwalk does with d_seq and mount_lock.
Without any artificial widening three walkers read the covered
file 27 times in a minute. step_into() now fails an unhashed
dentry marked DCACHE_CANT_MOUNT with -ESTALE and the walk is
retried.
* Keep covered mounts covered in OPEN_TREE_NAMESPACE.
Creating such a mount namespace only takes a user namespace and
the copy followed bind mount rules: no children without
AT_RECURSIVE and no unbindable mounts with it. An unprivileged
user could see what mounts covered in the source, such as the
parts of /proc and /sys that container runtimes mask. If the
caller doesn't own the source mount namespace a non-recursive
copy of a mount with something mounted below the requested
directory is now refused and a recursive copy includes unbindable
mounts, the way unshare() copies.
* Keep the lock on a mount that a propagated copy is moved beneath.
MNT_LOCKED moved to any mount that ended up beneath a locked
mount, propagated copies included. A host mount and umount on a
directory covered by a locked mount in a less privileged mount
namespace left that cover unlocked for the namespace's owner to
remove. Only mounts the caller places beneath take over the lock
now.
* Handle mount locking for automounts correctly. Which copies to
lock was decided by the mount namespace of the task that
triggered the automount. A task in a user namespace that
triggered one on a host mount through a file descriptor got the
host's own automount locked while its own copy stayed unlocked
and could have nosuid, nodev and noexec cleared. Use the owner of
the mount namespace the mount lands in.
- Use-after-free and crashes:
* Refuse an automount below a mount that is in no namespace.
The private clones overlayfs uses for its layers have the
MNT_NS_INTERNAL error pointer as their namespace, which
finish_automount() let through and count_mounts() dereferenced. A
fanotify filesystem mark on an overlayfs lower layer hands out
file descriptors on such a clone. With debugfs as the lower layer
opening "tracing" oopses with namespace_sem held for writing and
every mount operation on the system blocks from then on.
* Reset the old parent's ->overmount in mnt_change_mountpoint().
When propagate_umount() moved an overmount off a mount that a
file descriptor kept alive, MOVE_MOUNT_BENEATH through that
descriptor later followed the stale pointer into the freed
overmount.
* statmount() with STATMOUNT_BY_FD and pivot_root() read the parent
of a mount that may be unmounted and only held by a file
descriptor, while the parent's final mntput() can free it.
statmount() now reads it under mount_lock and pivot_root() first
checks that both mounts are in the caller's mount namespace.
* Queue a mount only once for mount notifications. A mount
reparented by one umount_tree() and taken down by the next under
the same namespace_sem hold, as in shrink_submounts(), was queued
twice. That cut the mounts queued in between out of notify_list
while it still pointed at them, and once they were freed every
later mount operation walked freed memory.
* Don't let a pseudo dentry become the root of a mount. A bind
mount of a bpf token file did that with a DCACHE_NORCU dentry,
which is freed without an RCU grace period while lockless path
walks may still look at it. Refuse to clone such a mount.
* Don't inherit MNT_UMOUNT in clone_mnt(). A bind mount of a lazily
unmounted nsfs or pidfs mount through its file descriptor started
out flagged as unmounted. Among other things __detach_mounts()
then dropped the namespace's reference on it, the mount outlived
its namespace and mount_setattr() through the descriptor read the
freed namespace. A recursive bind mount of such a mount also
copied the unmounted stack still attached to it. That now fails
with EINVAL, copying the mount itself still works.
* Remove the fsnotify marks of a mount namespace in free_mnt_ns()
instead of the RCU callback that frees the namespace, where
taking the group mutexes meant sleeping in softirq context.
- Propagation and copies:
* Keep a copied mount unbindable. Since v6.17 clone_mnt() didn't
copy the unbindable flag, so every mount namespace created with
CLONE_NEWNS had bindable copies of all unbindable mounts. This
had been fixed once before.
* Refuse MOVE_MOUNT_SET_GROUP on an unbindable mount. It made the
mount an unbindable slave, a state nothing else can produce, or
silently dropped the unbindable flag. CRIU applies MS_UNBINDABLE
after restoring sharing and isn't affected.
* Check a recursive bind mount for mount namespace loops.
Recursively bind mounting a tree from another mount namespace
could put a mount of a namespace's file inside that same
namespace, which then pins itself and all its mounts. Repeating
it leaks without limit, the reproducer took Shmem from 380 kB to
65916 kB. The copy is now checked with check_for_nsfs_mounts()
before it is grafted, as move_mount() does.
* Look at the topmost mount for a mount namespace file.
attach_recursive_mnt() never looked at the topmost mount of the
source's chain of overmounts. If that was the chain's only mount
namespace file an existing mount at a propagated destination got
buried below the root of the nsfs file where no path walk reaches
it.
* Don't put a mountpoint on a dentry that's being removed.
attach_recursive_mnt() makes a mountpoint of the source's root
without its inode lock, so a racing rmdir() of that directory
could leave a mount on it that nothing ever detaches.
d_set_mounted() now checks cant_mount() as well.
- nullfs:
* Take no inode lock for readdir of an immutable directory.
The root of every empty mount namespace is the same nullfs
directory and iterate_dir() held its i_rwsem across
->iterate_shared(). A reader whose buffer faults on a FUSE mount
of its own holds it for as long as its server wants, and with an
exclusive locker queued behind it every lookup that misses the
dcache, every create and every mount in that directory waits. One
user of an empty mount namespace stalls all others. Directories
with the new FOP_IMMUTABLE flag skip the lock.
* Refuse to reconfigure internal superblocks through fspick(),
MS_REMOUNT or the read-only remount that a synchronous umount()
of the root does. For nullfs only root in the initial user
namespace could do it, but the superblock is shared by every
mount namespace and the flags showed up in statfs() for all of
them.
* Don't update the access time on nullfs and refuse F_SET_RW_HINT
on an immutable inode.
- unshare: Free an nsproxy that was never installed with
nsproxy_free() when set_cred_ucounts() fails. put_nsproxy() dropped
active references that were never taken, which triggered a warning
and hid the caller's own namespaces from listns().
- Smaller changes: mount_setattr() checks the target before it walks
the tree to allocate peer group ids, unshare() puts the old
fs_struct before the old namespaces, dissolve_on_fput() drops the
file's reference to the tree itself, disconnect_mount() is
simplified and the documentation of the propagated unmount rule is
brought up to date"
* tag 'vfs-7.3-rc7.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (58 commits)
namespace: simplify disconnect_mount()
selftests/filesystems: test covered mounts
namespace: rework connected mounts
nullfs: add an empty immutable regular file
selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody
selftests/filesystems: add a helper that holds a readdir in a page fault
readdir: take no inode lock on an immutable directory
nullfs: refuse file locks
fsnotify: let a filesystem refuse marks on its objects
namespace: nothing is mounted on or written through knullfs
namespace: keep the private nullfs instance in knullfs
fhandle: decide the subtree check under mount_lock
selftests/filesystems: check that an automount below an overlay layer is refused
selftests/filesystems: check the atime of the empty mount namespace root
selftests/filesystems: check that a lock lands on the right mount and stays
namespace: keep the lock on a mount that a propagated copy is moved beneath
namespace: never expire a locked mount
nullfs: don't update the access time
namespace: handle mount locking for automounts correctly
namespace: refuse an automount below a mount that is in no namespace
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from wireless, wireguard, CAN and Bluetooth.
We have one known regression to wrap up in VLAN handling.
Current release - regressions:
- Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex
Previous releases - regressions:
- can: fix regression in handling RPS after migrating metadata to skb_ext
- eth:
- iavf: fix regressions in reconfig impacting bonding
- mana: fix packet forwarding performance regression
- stmmac: remove buggy VLAN acceleration support
Previous releases - always broken:
- a few high prio fixes for tun, and af_packet
- amt: fix a UaF on tunnel teardown
- eth:
- bnxt: fix PCIe AER recovery and FLR handling issues
- macb: don't modify Tx skbs before taking ownership
- axienet: don't leak Tx skbs on interface stop
- wifi:
- nxpwifi: number of LLM-ish fixes
- assorted mt76 fixes"
* tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits)
net: macb: copy shared skbs before appending the FCS
net: macb: check TX ring before modifying skb
vsock: Fix memory leak in vmci_transport_recv_dgram_cb()
wireguard: noise: reject response consumption after intermediate initiation
wireguard: queueing: preserve tstamp_type when encapsulating packet
net: openvswitch: validate transport header presence in set_ipv6_addr
net/smc: protect clcsock lifetime in smc_getname
ipv6: do not warn on route notification size race
ipv4: do not warn on route notification size race
ipv4: validate checksum_start before completing checksum
ptp: ocp: fix PCIe delay estimation calculation
xen/netfront: don't leak the skb when xennet_fill_frags() fails
net/packet: call packet_parse_headers after virtio_net_hdr_to_skb
xen/netfront: drop RX packets with a short Ethernet header
net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()
net: sparx5: free the matchall entry on destroy
selftests: mlxsw: Test port range occupancy on template create
mlxsw: spectrum_flower: Fix port range register leak in tmplt_create()
net: dsa: microchip: fix KSZ8765 fiber detection
net/mlx5e: Order ICOSQ cc update after CQ doorbell
...
|
|
Add a test that creates a tc chain template matching on both source and
destination port ranges and verifies via devlink-resource occupancy that
this does not leak port range registers, neither while the template
exists nor after it is deleted.
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Petr Machata <petrm@nvidia.com>
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Link: https://patch.msgid.link/e7b37b80adb7ac8b0ef20a22d7654a6656c1c545.1791294384.git.petrm@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Test that toggling GRO on a veth device while it is down updates the
peer's ndo-xmit XDP feature once the device comes up, for both GRO on
and GRO off.
xdp-features is only visible through netlink, so read it with the ynl
CLI, as double_udp_encap.sh does. Skip the checks if the CLI is not
found, as in an installed kselftest tree.
Signed-off-by: Tianyi Gao <tianyi@cloudflare.com>
Link: https://patch.msgid.link/20261006173241.65945-3-tianyi@cloudflare.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
- Update .mailmap entries for Andy Yan and John Garry
- Fix read-only MAP_SHARED /dev/zero mappings so they retain
shared-file semantics instead of being treated as anonymous memory,
also avoiding a CONFIG_DEBUG_VM assertion
- Fix 32-bit build warnings in the hugetlb-mmap selftest caused by
using the wrong printf format for size_t values
- Fix a boot-time crash when early function tracing causes CPA to free
kernel page tables before the workqueues used for deferred freeing
are available
- Fix two MREMAP_DONTUNMAP locked_vm accounting leaks: one caused by
an mlock-on-fault VMA self-merging, and one caused by partially
remapping a locked VMA
* tag 'mm-hotfixes-stable-2026-10-07-21-48' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
mailmap: update entry for Andy Yan
drivers/char/mem: mmap readonly MAP_SHARED-/dev/zero correctly
selftests/mm: cleanup -Wformat issues in hugetlb-mmap
mm: don't schedule deferred kernel page table freeing while booting
mailmap: update addresses for John Garry
mm/mremap: fix locked_vm leak by splitting VMA for MREMAP_DONTUNMAP
mm/mremap: fix locked_vm leak from MREMAP_DONTUNMAP self-merge
|
|
Add a boundary pair for route4 change without a routing attribute:
- a classid-only change of a wildcard filter (handle 0xffff8000) is a
metadata-only update and must succeed;
- the same change on a non-wildcard handle (0x10001) would rekey the
filter to the wildcard bucket, leaving the handle userspace stored stale,
and must be rejected with -EINVAL.
The change of a wildcard filter is the flow iproute2 uses by default
("tc filter add ... route classid X:Y" gives handle 0xffff8000).
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-0JB3.v1.20260923100921@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fixes from Tejun Heo:
- Taking a CPU offline could hang, or stall until the watchdog ejected
the BPF scheduler, when tasks on the dying CPU were still held by the
scheduler or sitting on a user dispatch queue. Re-enqueue them onto
the local queue when the runqueue goes offline so that the CPU pushes
them off like the other sched classes.
- The sequence number guarding against stale dispatches was per
runqueue, so a task re-enqueued on another CPU could get the same
number and a dispatch meant for its earlier instance was applied to
the new one. Use a per-task counter.
- A task dispatched to another CPU's local queue got its ops.dequeue()
only when picked to run and flagged as a core-sched pick. Call it at
insertion like for same-CPU dispatches.
* tag 'sched_ext-for-7.3-rc6-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Generate qseq from a per-task counter
selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves
sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ
sched_ext: Fix CPU hotplug hang when a dying CPU's tasks sit in the BPF scheduler
|
|
Commit ae571cd6015c ("selftests/mm: hugetlb-mmap: add setup of HugeTLB
pages") and commit 9c5a65f374f8 ("selftests/mm: merge map_hugetlb into
hugepage-mmap") added logs of 'hugepage_size' which has a size_t type.
However, the incorrect format specifier '%lu' was used which triggers
-Wformat warnings when building for 32-bit:
hugetlb-mmap.c:125:55: warning: format specifies type 'unsigned long'
but the argument has type 'size_t' (aka 'unsigned int') [-Wformat]
125 | ksft_print_msg("Default size hugepages (%lu kB)\n", hugepage_size >> 10);
| ~~~ ^~~~~~~~~~~~~~~~~~~
| %zu
hugetlb-mmap.c:134:47: warning: format specifies type 'unsigned long'
but the argument has type 'size_t' (aka 'unsigned int') [-Wformat]
134 | ksft_exit_skip("Not enough %lu Kb pages\n", hugepage_size >> 10);
| ~~~ ^~~~~~~~~~~~~~~~~~~
| %zu
Fix this by switching to the expected '%zu' format specifier.
Link: https://lore.kernel.org/20260927162419.820609-1-cmllamas@google.com
Fixes: ae571cd6015c ("selftests/mm: hugetlb-mmap: add setup of HugeTLB pages")
Fixes: 9c5a65f374f8 ("selftests/mm: merge map_hugetlb into hugepage-mmap")
Signed-off-by: Carlos Llamas <cmllamas@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com>
Reviewed-by: SJ Park <sj@kernel.org>
Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Shuah Khan <shuah@kernel.org>
Cc: <stable@vger.kernel.org>
|
|
Patch series "mm/mremap: fix two issues with MREMAP_DONTUNMAP", v2.
The MREMAP_DONTUNMAP feature is highly unusual in that it permits mremap()
operations that keep the original VMA in place.
Historically this has led to a lot of bugs where non-obvious interactions
occur between existing mremap() operations and the original VMA.
Commit 397432cab17b ("mm/mremap: account mm->locked_vm correctly for
MREMAP_DONTUNMAP") fixed an accidentally introduced bug around
mm->locked_vm accounting, but this wasn't the only issue.
And thus history repeats itself, as it turns out that mm->locked_vm
accounting is broken by MREMAP_DONTUNMAP yet again by two further cases,
and has been broken ever since the feature was introduced.
Both relate to the fact that VMA_LOCKED_BIT is cleared on the source VMA
(it has to be as all page tables are moved):
1. If an unfaulted VMA_LOCKONFAULT_BIT anonymous VMA self-merges it
clears the VMA_LOCKED_BIT flag and permanently leaks mm->locked_vm
pages.
2. If a partial mremap() is performed on a locked VMA there is a leak equal
to the number of pages not copied.
(Both for MREMAP_DONTUNMAP operations only)
Both issues can be fixed by treating the source range as distinct from the
destination range, which is the definition of what MREMAP_DONTUNMAP does
so is appropriate.
In case 1, simply disallow the self-merge, keeping adjacent source and
destination VMAs distinct.
In case 2, split the source range ahead of time if the VMA is mlock()'d,
so accounting is always correct.
Both changes were tested locally and confirmed to fix the issues.
For the purposes of a backport, the fixes are kept distinct, a follow-up
series can add self-tests.
This patch (of 2):
The MREMAP_DONTUNMAP feature is highly unusual in that it permits mremap()
operations that keep the original VMA in place.
Historically this has led to a lot of bugs where non-obvious interactions
occur between existing mremap() operations and the original VMA.
Fix another of these - self-merge.
Self-merge occurs when a VMA is moved in front of or behind itself and the
attributes of the VMA permit such a merge.
Practically this can only happen for unfaulted anonymous VMAs due to the
page offset equality requirement for merge:
|------------|
| |
| v
|...........||-----------||...........|
| || unfaulted || |
|...........||-----------||...........|
^ |
| |
|------------|
This becomes problematic if the VMA is configured by the user to
mlock-on-fault, i.e. the VMA_LOCKED_BIT, VMA_LOCKONFAULT_BIT VMA flags are
set.
MREMAP_DONTUNMAP clears mlock flags for the source VMA and maintains them
for the destination VMA.
Self-merge makes this impossible (there is only one VMA) and incorrectly
clears the destination VMA's mlock flags.
This causes a leak in mm->locked_vm as clearing this flag does not
decrement the counter and the VMA no longer has VMA_LOCKED_BIT set so it
is not decremented on unmap.
Resolve this by simply disallowing a self-merge in this case - the source
and destination VMAs are kept distinct and then are able to have distinct
mlock() flags.
Update dontunmap_complete() to make the now-redundant self-merge check a
VM_WARN_ON_ONCE() instead to guard against future regressions.
Also update the VMA userland tests to reflect the change.
Link: https://lore.kernel.org/20260930-fix-dontunmap-partial-self-merge-v2-0-f388985a0f0a@kernel.org
Link: https://lore.kernel.org/20260930-fix-dontunmap-partial-self-merge-v2-1-f388985a0f0a@kernel.org
Fixes: e346b3813067 ("mm/mremap: add MREMAP_DONTUNMAP to mremap()")
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Jose A. Perez de Azpillaga <azpijr@gmail.com>
Tested-by: Anirudh Srinivasan <asrinivasan@oss.tenstorrent.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Jann Horn <jannh@google.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Minchan Kim <minchan@kernel.org>
Cc: <stable@vger.kernel.org>
|
|
Test that mount cycles are resolved and test that mount covers behave as
expected.
Link: https://patch.msgid.link/20261002-work-mount-cover-v1-3-232a8f52b43c@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add selftests in tun.c verifying that a VLAN-tagged (802.1Q) TCPv4 GSO
packet without VIRTIO_NET_HDR_F_NEEDS_CSUM (both flags = 0 and
flags = VIRTIO_NET_HDR_F_DATA_VALID) is accepted when written to a TAP
device (/dev/net/tun with IFF_TAP | IFF_NO_PI | IFF_VNET_HDR).
Also verify that the following are rejected with -EINVAL:
- a mismatched GSO type (VIRTIO_NET_HDR_GSO_TCPV6 on a VLAN-tagged IPv4
packet).
- a frame whose TCP header is truncated after 10 bytes. Pulling only
sizeof(struct iphdr) + sizeof(struct tcphdr) bytes would accept it,
so this requires the transport offset found by flow dissection.
Finally, verify that a 65540-byte frame is accepted. Its skb->len is
above U16_MAX while it is flow-dissected, before eth_type_trans() pulls
the Ethernet header.
Based on a reproducer by Michael S. Tsirkin <mst@redhat.com>.
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20261001191140.2818991-4-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
namespace stalls nobody
A readdir of the root of an empty mount namespace stuck in the page
fault of its buffer must stall neither a create nor a lookup in that
directory. The test needs userfaultfd and skips without it.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-21-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
A getdents64() whose buffer is a page registered with userfaultfd sits
in handle_userfault() with whatever iterate_dir() took before it copied
the entries. Add readdir_hold.h for tests that want to know what that
blocks: it opens the userfaultfd before the test enters a user
namespace, the fault happens in the kernel and needs CAP_SYS_PTRACE in
the initial one, starts the readdir in a thread, waits until that
thread is stuck, queues a create behind it and waits for it to settle,
then probes a lookup with a watchdog and says whether it came back.
The page is released afterwards so that everything drains on a kernel
that still holds the lock across the copy.
The two users follow.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-20-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Make sure that we don't automount on top of internal things.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-12-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
No access time updates for immutable nullfs.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-11-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Check MNT_LOCKED:
- A child in a user namespace of its own triggers the tracefs automount
through a directory descriptor on the host's debugfs mount. The copy
in the child's namespace has to be locked and the host's mount unlocked.
- A child moves a bind of that automount beneath a locked covering mount,
unmounts the covering mount and asks for the umount of an unlocked
ancestor. This must not expire the mount which holds the lock now.
- The host mounts and unmounts on a directory that a locked mount covers
in a user namespace further down the propagation chain. That cover has
to be locked afterwards as it was before.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-10-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Check that an immutable inode takes no write hint.
Link: https://patch.msgid.link/20261002-work-mount-fixes-4-v1-4-dd44b89d44ce@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix overflow of backward jump offset in constant blinding
(Alexei Starovoitov)
- Fix packet range of packet pointers sharing an id when var_off
tightens umax of one pointer and not the other (Alexei Starovoitov)
- Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc
(Alexei Starovoitov)
- Fix use-after-free of progs detached from busy trampolines: wait for
an RCU tasks grace period before freeing trampoline progs, and patch
detached progs out of trampoline images that are still in use
(Florent Revest)
- Hold map BTF for the memory allocator destructor record to fix UAF in
deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi)
- Fix missing migration protection in resizable hashtab
lookup_and_delete batch operation (Ömer Mete Kaya)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf:
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
bpf: Fix objects stuck in free_by_rcu_ttrace
bpf: Factor out __do_call_rcu_ttrace()
selftests/bpf: Test packet range of pointers sharing an id
bpf: Fix packet range of pointers sharing an id
selftests/bpf: Detach a trampoline prog while a task sleeps before it
bpf: Skip detached progs in trampoline images that are still in use
bpf: Wait for an RCU tasks grace period before freeing trampoline progs
bpf: Hold map BTF for the memory allocator destructor record
bpf: Fix overflow of jump offset in constant blinding
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux
Pull block fixes from Jens Axboe:
- NVMe fixes via Keith:
- Fix an out-of-bounds write in nvmet_auth_challenge(), where
sizeof() on a void pointer undercounted the challenge header and
let a short AUTH_RECEIVE buffer pass the check
- nvme-multipath fixes for an ANA log bounds check underflow, the
command effects log lifetime for multipath heads, and only
setting BLK_FEAT_ZONED after the zone info is known.
- nvmet fixes for ns->enabled teardown ordering, rejecting I/O
after the percpu ns reference is killed, device path preservation
on allocation failure, and too-short SGL segments in pci-epf
- nvme-tcp: revert the per-socket dynamic lockdep keys, and delay
the socket reclassification
- A DMA pool alignment quirk for the Micron 4100AT
- Controller state/reset race fixes, and -Wformat-security
workarounds
- blk-mq: set RQF_USE_SCHED when the operation is known, and allow
cached requests to be used for flush operations
- Reject polled dio with user integrity metadata
- Save the IRQ state in blkg_tryget_closest()
- Set the zone write granularity in virtio_blk
- ublk selftest fixes
* tag 'block-7.3-20261002' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux: (23 commits)
virtio_blk: set the zone write granularity
nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
nvme: fix command effects log lifetime for multipath heads
nvmet: don't allow I/O admission after percpu ns reference is killed
nvmet: defer setting ns->enabled to false in nvmet_ns_disable()
nvmet: copy the hostid into the ctrl before creating PR pc_refs
nvmet-auth: fix out-of-bounds write in nvmet_auth_challenge()
nvmet: pci-epf: reject too-short SGL segments
nvme-multipath: fix underflow in ANA log bounds checks
nvme: work around all -Wformat-security warnings
nvme: work around -Wformat-security warning
nvme: do not reset controllers in NVME_CTRL_NEW state
nvme-tcp: delay nvme_tcp_reclassify_socket()
Revert "nvme-tcp: lockdep: use dynamic lockdep keys per socket instance"
drbd: remove unused drbd_nl_mcgrps[] array
blk-mq: allow cached requests to be used for flush operations
blk-mq: set RQF_USE_SCHED when the operation is known
block: reject polled dio with user integrity metadata
selftests: ublk: fix unused_result error
blk-cgroup: save IRQ state in blkg_tryget_closest()
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid
Pull HID fixes from Benjamin Tissoires:
- Revert of the Bolt integration into hid-logitech-dj (Benjamin
Tissoires)
- A couple of buffer overflow in Intel-thc-hid (Even Xu)
- A couple of Sashiko findings fixes in hid-multitouch and HID-BPF
(Aldo Ariel Panzardo and Benjamin Tissoires)
* tag 'hid-for-linus-2026100201' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid:
selftest/hid: add test for negative return codes for hid_bpf_hw_request
HID: bpf: cast size to ssize_t when checking hid_bpf_hw_request
HID: Intel-thc-hid: Intel-quickspi: Fix buffer overflow
HID: Intel-thc-hid: Intel-quicki2c: Fix buffer overflow
HID: universal-pidff: Add support for Turtle Beach VelocityOne Race
HID: multitouch: stop the release timer from being rearmed on remove
Revert "HID: logitech: add Bolt receiver support for Logitech HID++ devices"
|
|
The relay used to hand its General Queries to dev_queue_xmit() on the
amt device, where a query could wait in a qdisc and outlive the tunnel
it pointed to. The previous patch sends them directly from the receive
path instead.
Count the IGMP and MLD queries that leave the relay through amtr with
tc flower filters on its egress, installed before the gateway comes
up, and check that there are none. The forwarding tests before it
already show that the gateway received its queries, since it cannot
join without one.
Without the previous patch the new test fails (one run counted 7 IGMP
and 6 MLD queries); with it, all of amt.sh passes.
Signed-off-by: Omar Ramadan <omar@blockcast.net>
Link: https://patch.msgid.link/20260928202312.74574-3-omar@blockcast.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Pull kvm fixes from Paolo Bonzini:
"The most intrusive change is reverting a commit from 7.3-rc1 that made
struct kvm a bit too large, and fixing the same issue otherwise.
There are again a lot of selftests lines; the sheer number of commits
is not small but I don't expect much more for 7.3 due to people
travelling to Plumbers next week.
ARM:
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
x86:
Various bugfixes where the guest could do stupid things on purpose to
cause problems in the host:
- Failed VMRUNs can cause pending TLB flushes to be dropped, and in
general some actions done through VMCB control fields have to be
redone if VMRUN fails
- Toggling MSR interceptions or eVMCS execution controls can cause
the host to use a stale MSR permission bitmap
- Bad page tables can cause a WARN.
Also fix issues in last week's pull request (my fault, for changing
email workflow and thus missing feedback sent to kvm@ but not LKML).
Generic:
- Take kvm_lock when creating vCPUs. For almost two decades everybody
thought it was not done for some unspecified performance reasons,
but in reality it was only done because kvm_lock was originally a
spinlock.
This is a better fix than 97d65b544f48 ("KVM: Check for duplicate
vcpu_id as early as possible", from the 7.3 merge window), and does
not waste 2K per VM, hence its inclusion here"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits)
KVM: arm64: Use stage-1 leaf size for VM_PFNMAP
KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF
KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM
KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs
KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT
KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor
KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq
KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing
KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state
KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it
KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12
KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt
KVM: selftests: Verify that L0's TPR doesn't get clobbered
KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2
KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified
KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted
KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN
KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded
...
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kvmarm/kvmarm into HEAD
KVM/arm64 fixes for 7.3, round #2
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
|
|
Delete all elements of BPF_F_NO_PREALLOC hash map in one batch. The first
free_bulk() starts RCU tasks trace GP and the rest of the elements are
freed while it's in flight. Wait for call_rcu_ttrace_in_progress to clear
in bpf_mem_cache of every cpu and check that free_by_rcu_ttrace and
waiting_for_gp_ttrace lists are empty.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20260930095920.601738-4-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Add tests where two packet pointers share an id and tightening one
pointer's umax from its var_off would put it less than their constant
distance from the other's umax: with an index & 0x38 capped at 50, the
base pointer keeps umax 50, so the pointer 8 bytes further on must keep
umax 58, even though its known bits allow at most 56.
These refused a valid program or accepted an out-of-bounds access before
the fix:
- check the advanced copy, load through the base: valid, was refused;
- check the base, load the byte at base + 1 through a copy advanced by
8: was accepted;
- check base + 4, load 4 bytes at base + 2 through base + 8: reads two
bytes past the checked range, was accepted;
- the same as the second with data_meta pointers checked against data:
was accepted.
These pass with and without the fix and cover nearby paths:
- subtract an unknown scalar from a checked pointer and load below it
(the range is kept across a new id);
- reach a load through two paths whose checks cover 8 and 7 bytes after
the loaded pointer; the second path must not be pruned by the first;
- spill a copy of a pointer, check the pointer, fill the copy and load
one byte past the checked range: the load is refused, and the copy
has the range of the check.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261001145255.855630-2-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
Since commit 022ac0750883 ("bpf: use reg->var_off instead of reg->off
for pointers"), find_good_pkt_pointers() sets the range of all packet
pointers sharing an id from the umax of the compared pointer, and
check_packet_access() requires umax + off + size <= range. That assumes
the umax of two such pointers differ by exactly their constant distance.
reg_bounds_sync() breaks it when var_off tightens one umax and not the
other:
r4 &= 0x38
if r4 > 50 goto exit ; umax 50, var_off (0x0; 0x38)
r5 = pkt + r4 ; umax 50
r6 = r5
r6 += 8 ; umax 56, not 58
Comparing r6 with pkt_end sets the range to 56, and the valid 8-byte
load at r5 is rejected (50 + 8 > 56). Comparing r5 sets it to 50, and
the out-of-bounds 1-byte load at r6 - 7, i.e. r5 + 1, is accepted
(56 - 7 + 1 <= 50).
Don't call reg_bounds_sync() on a packet pointer that keeps its id (a
constant was added or subtracted) or its range (an unknown non-negative
value was subtracted), so that var_off cannot tighten its umax. Only
update the 32-bit bounds from var_off: reg_bounds_sanity_check() wants
them constant when the lower half of var_off is, e.g. for pkt + 8.
This relies on nothing else changing the 64-bit bounds of a packet
pointer, which holds today.
var_off of such a pointer is no longer narrowed by its bounds. Adjust
three verifier_align expectations; the low bits, which the alignment
checks use, don't change. veristat on the selftests shows no verdict
changes and +0.8% insns in test_cls_redirect_subprogs.
Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers")
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261001145255.855630-1-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Bluetooth, WiFi and netfilter.
We are actively retargeting several non-urgent fixes towards next,
but the traffic on the ML looks ever-increasing, and propagating the
push-back towards subsystems is not immediate.
No known outstanding regressions.
Current release - regressions:
- netfilter: nft_set_rbtree: skip transaction elements during GC
Previous releases - regressions:
- sched: cls_api: reclaim an empty proto on the error path
- core:
- fix checksum offsets in skb_splice_from_iter()
- cap skb->queue_mapping when the tx queue is picked
- page_pool: fix use-after-free in page_pool_recycle_ring_bulk()
- wifi:
- mac80211: fix slab-out-of-bounds read in ieee80211_monitor_select_queue()
- mac80211: drop oversized fragments to avoid extra_len overflow
- netfilter:
- flowtable: restore ieee80211 forward path
- bluetooth: hci_conn: Lock parent access during enhanced SCO setup
- eth:
- bcmgenet: allocate RX buffers as page fragments
- stmmac: fix rx Scatter-Gather support
- octeontx2-pf: fix aura BPID assignment when CONFIG_DCB is enabled
- gve: DQO: accept TSO packets with non-protocol gso_type bits
- r8169: disable EEE on RTL8168h/8111h
Previous releases - always broken:
- tcp: refresh TS.Recent for accepted old ACKs
- wifi:
- ath11k: reset ar->num_stations on hardware start
- cfg80211: fix RTS threshold setting for single-radio PHY
- bluetooth: btintel_pcie: fix plen overflow in btintel_pcie_recv_frame()
- eth: bcmgenet: fix NULL dereference in set_coalesce before first open
Misc:
- Eric is retiring from google and updating his contact info"
* tag 'net-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (96 commits)
net: phy: aquantia: fix system interface type not updated in forced mode
net: usb: qmi_wwan: add Rolling Wireless RN947R
net: mvneta: clear XDP pfmemalloc flag between frames
ipv6: sr: use skb_get_hash_net() in seg6_make_flowlabel()
net/mlx5e: Fix AF_XDP TX timestamp teardown NULL dereference
r8169: disable EEE on RTL8168h/8111h
octeontx2-pf: Fix RSS indirection table size
sctp: check RCV_SHUTDOWN after the sendmsg connect wait
net: sparx5: make ports inherit the switch base mac address type
net: microchip: vcap: stop scanning after deleting key field
netfilter: flowtable: restore ieee80211 forward path
netfilter: flowtable: generalize pending status bit
netfilter: bpf: reject invalid NAT manipulation types
netfilter: nft_set_rbtree: skip transaction elements during GC
ipvs: filter some flags received in the backup server
ipvs: do not create invisible templates
ipvs: bound LBLCR and LBLC cache growth
ipvs: fix missing counter decrement in lblc
netfilter: nft_flow_offload: drop flowtable reference on init error path
selftests: net: check timestamp echo after an old ACK
...
|
|
Add a test for the root of an empty mount namespace:
- fspick() of the root fails with EINVAL
- mount(MS_REMOUNT) of the root fails with EINVAL
- umount() of the root fails and doesn't remount it read-only
Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-12-be34c83956ae@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
a mount
Add a test for moving a mount with a mount namespace file on top of it:
- S, a bind mount of a file, with a newer mount namespace's file on top
- A shared with a slave B that has Q on B/file
- S moved onto A/file, its copy lands on B/file below Q
B/file keeps reading Q and once Q is unmounted it reads the copy.
Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-9-be34c83956ae@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add a test for open_tree(OPEN_TREE_NAMESPACE) from a user namespace that
doesn't own the mount namespace it copies from:
- a file covered by a private or an unbindable tmpfs mount
- the non-recursive copy of the parent mount fails with EINVAL
- the recursive copy keeps the file covered in the new mount namespace
- the owner of the mount namespace keeps the bind mount semantics
Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-7-be34c83956ae@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
caller's mount namespace
Add a test for a recursive bind mount of a namespace file in another mount
namespace:
- a bind mount of the network namespace file, held through a descriptor
- the caller's own mount namespace file stacked on top of it there
- the recursive bind mount through the descriptor fails with EINVAL
- a plain bind mount of the file still works
The copy would pin the namespace it is put in otherwise.
Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-5-be34c83956ae@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add a test for the shrinkable submounts of a synchronous umount:
- P shared, P1 a slave that is shared in turn, P2 its peer
- B on P/options with copies on P1 and P2, R on top of the copy in P2
- P with B moved below P1, R is the working directory
- umount(P1) slides R to where the copy of B is looked up
The umount fails with EBUSY and R stays mounted. Needs the tracefs
automount below debugfs for the shrinkable mounts.
Link: https://patch.msgid.link/20260930-work-mount-fixes-3-v1-3-be34c83956ae@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add a dequeue_remote test that makes moves of tasks in the BPF
scheduler's custody to another CPU's local DSQ the common case, both via
SCX_DSQ_LOCAL_ON dispatch and via scx_bpf_dsq_move_to_local(). The BPF
scheduler tracks each task's custody state and triggers scx_bpf_error()
if a custody period doesn't end with exactly one ops.dequeue() before
the task runs, or if it ends with an SCX_DEQ_CORE_SCHED_EXEC dequeue of
a task without a core cookie.
Without the previous patch, the test fails with:
sched_ext: dequeue_remote: dequeue_remote.bpf.c:141: 15 (rcu_preempt): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=3 cpu=2 seq=1)
...
ops_dequeue+0x114/0x170
set_next_task_scx+0x104/0x1e0
__pick_next_task+0xc7/0x180
__schedule+0x154/0x1870
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
Add a regression test for a gap-filling packet whose acknowledgment has
become old. Linux first sends data to the peer. Deliver two peer data
packets out of order: the later packet acknowledges Linux's data, while
the delayed packet still carries the earlier acknowledgment.
Require the ACK that closes the receive gap to echo the delayed packet's
timestamp, 301000. Without the fix, Linux accepts the data but still echoes
the previously saved timestamp, 1000. Also check that the application can
read all 34 bytes.
Use a large jump in peer timestamps to represent the idle interval, with
no real wait, loss or retransmission. The test directly checks the outgoing
timestamp echo. It requires packetdrill's merged TSecr verification fix
(linked below); older tools incorrectly pass on an unfixed kernel.
Use the existing packetdrill selftest runner for IPv4, IPv6 and
IPv4-mapped IPv6.
Link: https://github.com/google/packetdrill/commit/83f72d3f9085d0e26eb4d206fe4d7cfab5b6d872
Assisted-by: LLM
Signed-off-by: Jeff Jo <jeffjo@openai.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
|
|
Exercise software UDP checksums when splice() transfers multiple pipe
fragments into a corked datagram. Use separate source pipes so short writes
cannot merge, and use odd fragment lengths to expose checksum-position
errors between fragments.
Cover IPv4, IPv6 and IPv4-mapped destinations, connected and unconnected
sockets, resident prefixes, even and odd fragment lengths, and a fragment
starting near a page boundary. Ordinary sends provide a control. Check the
complete payload and reject extra datagrams after a successful send.
Keep receive waits bounded and skip unavailable address families or pipe
capacities. The test uses only loopback sockets with ephemeral ports and
has no dependency on the application that exposed the regression.
Signed-off-by: Alireza Asgari <alireza@asgari.net>
Link: https://patch.msgid.link/20260924-fix-udp-splice-checksum-v1-2-fe61d65447a8@asgari.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Userspace can hide a feature from a guest by clearing its field in a
writable ID register, and KVM then makes the feature's instructions
UNDEFINED in the guest by trapping or disabling them. No selftest checks
that.
Add a test that runs the instruction of each of TLBI OS, MOPS, TCR2_EL1
and FPMR once with its field as advertised and once with it cleared, and
expects an UNDEF only when cleared. A feature the vCPU doesn't advertise
is skipped, as is hidden TLBI OS on a CPU with neither FGT nor
FEAT_EVT2, where KVM can't trap it.
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Link: https://patch.msgid.link/20260929090031.3829185-5-fuad.tabba@linux.dev
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
Add a test for the use-after-free fixed by the previous commit. A
sleepable prog on bpf_fentry_test1() blocks a task on a userfaultfd
page, like bpf_mod_race does, while the two progs that run after it in
the same trampoline are detached and freed one after the other, which
the test knows from their .bss maps going away. After the first detach
the task is in an image that isn't the trampoline's current one
anymore. The task is then released and must not call into the freed
progs. This is done with fentry progs, with fexit progs, where the task
is already past the original function, and with a task sleeping in a
fentry prog while fexit progs are detached, where the original function
must still be called. bpf_mod_race's userfaultfd helper moves to
testing_helpers.c so that both tests can use it.
Without the fix, on a kernel with KASAN:
BUG: KASAN: vmalloc-out-of-bounds in __bpf_prog_enter_recur+0xed/0x1e0
Read of size 8 at addr ffa0000000144040 by task test_progs/171
CPU: 6 UID: 0 PID: 171 Comm: test_progs Tainted: G OE 7.2.0+ #1 PREEMPT(full)
Call Trace:
<TASK>
__bpf_prog_enter_recur+0xed/0x1e0
bpf_trampoline_6442545468+0x72/0xe3
bpf_fentry_test1+0x9/0x20
bpf_prog_test_run_tracing+0x183/0x3e0
__sys_bpf+0xd3f/0x38f0
...
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-4-florent.revest@linux.dev
|
|
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-6-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup
Pull cgroup fix from Tejun Heo:
- A cpuset partition could claim CPUs an ancestor partition already
held exclusively. Restore the rejection an earlier change had turned
into a warning.
* tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fix from Tejun Heo:
- The CPU topology helper for BPF schedulers took no buffer size, so
its structure couldn't grow without breaking schedulers built against
the older layout. Add a size argument.
* tag 'sched_ext-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo can grow
|
|
can grow
scx_bpf_cid_topo() copies struct scx_cid_topo into a buffer the BPF program
sized from its own vmlinux.h while the verifier sizes the write from the
running kernel's BTF. The struct may grow and each growth then breaks every
scheduler built against the older layout, rejected at load or written past
its buffer. This is the usual hole for a struct handed to BPF, closed
elsewhere with a size argument, and it was missed here.
Take the buffer size, copy the smaller of it and the kernel's struct and set
the rest to -1. Accesses to the copy are CO-RE relocated, so the struct can
grow by appending fields, which its comment now states. The kfunc changes in
place: the cid interface is still being finalized and no released scheduler
uses the current form.
Fixes: e9b55af47edf ("sched_ext: Add topological CPU IDs (cids)")
Cc: stable@vger.kernel.org # v7.2+
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
|
|
subpartitions_cpus conflict
When a remote partition is created underneath an existing local partition
via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees
parent->partition_root_state == PRS_MEMBER and calls
remote_partition_enable().
Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency
in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus,
subpartitions_cpus) error check in remote_partition_enable() with
WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning
and proceeds to enable the remote partition on CPUs that are already
owned by the ancestor local partition in subpartitions_cpus.
This can be reproduced on Linux 7.3.0-rc3 with:
mkdir -p /tmp/cg1
mount -t cgroup2 none /tmp/cg1
echo "+cpuset" > /tmp/cg1/cgroup.subtree_control
mkdir /tmp/cg1/A
echo 1 > /tmp/cg1/A/cpuset.cpus
echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive
echo root > /tmp/cg1/A/cpuset.cpus.partition
echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control
mkdir /tmp/cg1/A/B
echo 1 > /tmp/cg1/A/B/cpuset.cpus
echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive
echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control
mkdir /tmp/cg1/A/B/D
echo 1 > /tmp/cg1/A/B/D/cpuset.cpus
echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive
echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition
which triggers:
WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300
and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions
claiming exclusive CPU 1.
Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects
subpartitions_cpus in remote_partition_enable(), matching the error code
used by remote_cpus_update() for the same subpartitions_cpus conflict, and
add a regression test case to
tools/testing/selftests/cgroup/test_cpuset_prs.sh.
Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and
tools/testing/selftests/cgroup/test_cpuset_prs.sh.
Fixes: 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs")
Suggested-by: Guopeng Zhang <guopeng.zhang@linux.dev>
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace
Pull probe fixes from Masami Hiramatsu:
- kprobes: Fix permanent hang when flushing the kprobe optimizer
Fix a deadlock when disabling kprobe optimization via sysctl or
debugfs where flushers hung waiting for optimizer_completion.
Replaced the completion with an optimizer_passes counter and
wait_var_event_mutex() under kprobe_mutex so concurrent flushers can
wait and wake up safely.
- fprobe: Terminate the fgraph_data list when the reservation is not
filled
Fix an issue where unused shadow stack data left uninitialized by
fprobe_fgraph_entry() was misparsed as stale fprobe headers on
return. Explicitly write a zero word to terminate the list and update
read_fprobe_header() to handle the zeroed slot properly.
- ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc
Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
like s390 where a symbol exists once in core kernel but also in
modules. Anchor the /proc/kallsyms search regex to the end of the
line so that module symbols are not incorrectly counted.
* tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
kprobes: Fix permanent hang when flushing the kprobe optimizer
fprobe: Terminate the fgraph_data list when the reservation is not filled
selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
Add overmount_reparent_test:
- a mount that a file descriptor keeps alive after propagate_umount()
reparented its overmount and unmounted it must not be walked into the
freed overmount by move_mount(MOVE_MOUNT_BENEATH)
Link: https://patch.msgid.link/20260926-work-mount-fixes-2-v1-10-f357abf3d17b@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Add unmounted_tree_test:
- a detached nsfs or pidfs bind mount that a file descriptor keeps
alive can't be copied recursively, a copy of the mount itself and
a copy of a mounted one still work
- a detached tree with a child that was vacated behind it can't be
copied recursively either
- a copy of a detached bind mount is an ordinary mount: unlinking its
mountpoint from another mount namespace unmounts it
- mount_setattr() refuses to change the propagation of a detached
tree and still allows it for the root of an anonymous mount
namespace
Every case runs in a user and mount namespace of its own.
Link: https://patch.msgid.link/20260926-work-mount-fixes-2-v1-9-f357abf3d17b@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
in L2
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Add a selftest to verify that KVM intercepts x2APIC MSR accesses for L1
after APICv is inhibited while L2 is active. This is a regression test for
an AVIC bug where KVM would skip updating x2APIC MSR intercepts while L2
is active, thus giving L1 access to a wide swath of L0's x2APIC surface.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260710162052.2188574-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|