| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux
Pull arm64 fixes from Will Deacon:
"It finally seems to have calmed down on the arm64 fixes front, so
please pull these two straightforward fixes for -rc7. One fixes the
EL2 trap configuration for implementation-defined CPU PMU hardware
during boot and the other fixes a kcov selftest failure by excluding
our softirq early entry code:
- Fix PMU EL2 trap configuration for CPUs with an IMPDEF PMU
- Fix kcov boot selftest failure by excluding our early IRQ entry
code"
* tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux:
arm64: irq: exclude the softirq stack switch from KCOV
arm64/boot: Don't set PMUv3p9 FGT2 bits without PMUv3
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from wireless, wireguard, CAN and Bluetooth.
We have one known regression to wrap up in VLAN handling.
Current release - regressions:
- Bluetooth: RFCOMM: fix deadlock on rfcomm_mutex
Previous releases - regressions:
- can: fix regression in handling RPS after migrating metadata to skb_ext
- eth:
- iavf: fix regressions in reconfig impacting bonding
- mana: fix packet forwarding performance regression
- stmmac: remove buggy VLAN acceleration support
Previous releases - always broken:
- a few high prio fixes for tun, and af_packet
- amt: fix a UaF on tunnel teardown
- eth:
- bnxt: fix PCIe AER recovery and FLR handling issues
- macb: don't modify Tx skbs before taking ownership
- axienet: don't leak Tx skbs on interface stop
- wifi:
- nxpwifi: number of LLM-ish fixes
- assorted mt76 fixes"
* tag 'net-7.3-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (128 commits)
net: macb: copy shared skbs before appending the FCS
net: macb: check TX ring before modifying skb
vsock: Fix memory leak in vmci_transport_recv_dgram_cb()
wireguard: noise: reject response consumption after intermediate initiation
wireguard: queueing: preserve tstamp_type when encapsulating packet
net: openvswitch: validate transport header presence in set_ipv6_addr
net/smc: protect clcsock lifetime in smc_getname
ipv6: do not warn on route notification size race
ipv4: do not warn on route notification size race
ipv4: validate checksum_start before completing checksum
ptp: ocp: fix PCIe delay estimation calculation
xen/netfront: don't leak the skb when xennet_fill_frags() fails
net/packet: call packet_parse_headers after virtio_net_hdr_to_skb
xen/netfront: drop RX packets with a short Ethernet header
net: skbuff: don't leave stale bytes in skb_copy_and_csum_bits()
net: sparx5: free the matchall entry on destroy
selftests: mlxsw: Test port range occupancy on template create
mlxsw: spectrum_flower: Fix port range register leak in tmplt_create()
net: dsa: microchip: fix KSZ8765 fiber detection
net/mlx5e: Order ICOSQ cc update after CQ doorbell
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc
Pull SoC fixes from Arnd Bergmann:
- Four distinct issues in TEE firmware, all fairly minor
- Five devicetree mistakes on NXP i.MX8, lx2160a and Qualcomm
based machines, one of these may cause file system corruption
from an incorrect SD card supply voltage
- Five fixes for clk drivers on new Qualcomm platforms,
addressing issues with incorrect enable states
* tag 'soc-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc:
MAINTAINERS: update my email address
tee: optee: ffa: support shared memory offsets on large-page kernels
optee: register TEE devices only once fully initialized
tee: shm: reject zero-sized allocations in tee_dyn_shm_alloc_helper()
arm64: dts: imx8mp-var-dart-sonata: Fix Sonata SD I/O supply
arm64: dts: lx2160a: fix the iic5 spi3 pinmux offset and value
arm64: dts: lx2160a: fix IIC1 pinmux submask rejected by pinctrl-single
arm64: dts: lx2160a: fix incorrect pinmux
clk: qcom: gpucc-kaanapali: Mark the GPU CX GDSC as votable
clk: qcom: gpucc-glymur: Mark the GPU CX GDSC as votable
clk: qcom: gcc-kaanapali: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-hawi: Fix always-enabling PCIE_RSCC clocks
clk: qcom: gcc-eliza: Fix always-enabling PCIE_RSCC clocks
arm64: dts: qcom: x1-denali: Fix microphone distortion
tee: qcomtee: fix kernel-doc warnings
|
|
On IRQ exit, __irq_exit_rcu() drops HARDIRQ_OFFSET before calling
invoke_softirq(). The arm64 do_softirq_own_stack() wrapper and its
____do_softirq() trampoline run before __do_softirq() establishes
softirq context, so KCOV records their PCs in the interrupted task's
coverage buffer. They also run outside softirq accounting when
local_bh_enable() reaches do_softirq().
The IRQ-exit coverage makes the KCOV boot selftest fail even after
the scheduler and timer coverage leaks are suppressed. The generic
softirq.o is already excluded from KCOV, but that exclusion does not
cover the separately compiled arm64 wrappers.
Exclude irq.o from KCOV instrumentation, as is already done for the
arm64 entry code. This covers both wrappers even with compilers that
lack the no_sanitize_coverage attribute. The softirq action functions
remain instrumented for explicit remote coverage.
Fixes: 8eb858c44b98 ("arm64: run softirqs on the per-CPU IRQ stack")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
Commit 9e8db5913264 ("net: avoid false positives in untrusted gso
validation") added a '&& skb->network_header' check before flow-dissecting
GSO packets without VIRTIO_NET_HDR_F_NEEDS_CSUM in
__virtio_net_hdr_to_skb(), because some callers (such as tun_get_user(),
tun_xdp_one(), virtnet_receive_done(), and raw_verify_header()) called
virtio_net_hdr_*_to_skb() before initializing skb->network_header and
skb->dev.
Because __alloc_skb() and __build_skb_around() zero-initialize
skb->network_header to 0 (unlike mac_header and transport_header which
are initialized to ~0U), those four callers always had
skb->network_header == 0 and bypassed flow dissection in
__virtio_net_hdr_to_skb(). More generally, skb->network_header is an
offset from skb->head (where 0 is also a valid offset whenever
skb_headroom(skb) is 0), not a boolean flag.
Whenever the 'if (gso_type && skb->network_header)' branch was skipped,
the fallback 'else if (gso_type)' only pulled nh_min_len + thlen (40 bytes
for TCPv4) without dissecting the packet, without validating ip_proto or
n_proto, and without setting skb->transport_header.
If the packet has a malformed network header, it is not rejected and a
subsequent skb_probe_transport_header() also fails, leaving
skb->transport_header at ~0U (0xffff). Similarly, if an IPv4 packet
carries IP options (ihl > 5) or an IPv6 packet carries extension headers,
pulling only nh_min_len + thlen can leave the TCP header outside
skb->head. In both cases, tcp_hdrlen(skb) in skb_gso_transport_seglen()
reads out-of-bounds:
BUG: KASAN: slab-out-of-bounds in skb_gso_transport_seglen
Read of size 2 by task poc/133
skb_gso_transport_seglen (net/core/gso.c:155)
skb_gso_validate_mac_len (net/core/gso.c:270)
tbf_enqueue (net/sched/sch_tbf.c:260)
dev_qdisc_enqueue (net/core/dev.c:4227)
__dev_queue_xmit (net/core/dev.c:4884)
In addition, checking virtio_net_hdr_match_proto() only inside
'if (!skb->protocol)' before flow dissection both skipped validation when
skb->protocol was pre-set by the caller and rejected VLAN-tagged frames
whose outer L2 protocol is ETH_P_8021Q or ETH_P_8021AD.
Fix this by:
1. Initializing skb->dev and skb->network_header (plus skb->protocol for
IFF_TUN) before virtio_net_hdr_*_to_skb() in tun_get_user(),
tun_xdp_one(), virtnet_receive_done(), and raw_verify_header(). In
tun_get_user(), drop the redundant skb_reset_mac_header(skb) in the
IFF_TUN case since __virtio_net_hdr_to_skb() unconditionally resets
mac_header.
2. Removing '&& skb->network_header' and the unvalidated
'else if (gso_type)' fallback in __virtio_net_hdr_to_skb() so all GSO
packets without VIRTIO_NET_HDR_F_NEEDS_CSUM are flow-dissected, have
their transport header pulled into linear data, and have
skb->transport_header set.
3. Moving the virtio_net_hdr_match_proto() check to after
skb_flow_dissect_flow_keys_basic(), validating keys.basic.n_proto
against hdr_gso_type.
Fixes: 9e8db5913264 ("net: avoid false positives in untrusted gso validation")
Fixes: d5be7f632bad ("net: validate untrusted gso packets without csum offload")
Fixes: 924a9bc362a5 ("net: check if protocol extracted by virtio_net_hdr_set_proto is correct")
Reported-by: Weiming Shi <bestswngs@gmail.com>
Closes: https://lore.kernel.org/netdev/20260927163117.746432-2-bestswngs@gmail.com/
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Cc: Michael S. Tsirkin <mst@redhat.com>
Link: https://patch.msgid.link/20261001191140.2818991-3-edumazet@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Ingo Molnar:
- Don't apply va_align to hugetlb mappings on AMD F15h systems
that have custom va_align.bits values (Laurent Wandrebeck)
- Fix PMD teardown handling regression flagged by lockdep
(Mikhail Gavrilov)
- Hide ptrace header register offset macros behind __ASSEMBLER__ or
__FRAME_OFFSETS, to fix user-space build errors that may trigger
if they happen to shadow these short and generic macro names
(Nick Desaulniers)
* tag 'x86-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
{x86,um}/uapi/ptrace: Guard register offset macros with __ASSEMBLER__ or __FRAME_OFFSETS
x86/mm: Drop unnecessary PMD page copy when freeing
x86/mm: Don't apply va_align to hugetlb mappings on AMD F15h
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fix race between perf_event_exit_task() and perf_pending_task()
(Luo Gengkun)
- Fix perf header output management regressions (Ian Rogers)
- Require kernel access for text poke events (Zhengchuan Liang)
* tag 'perf-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf: Require kernel access for text poke events
perf: Replace perf_event_header__init_id with full header init
perf: Fix race between perf_event_exit_task() and perf_pending_task()
|
|
On a CPU with FGT2 and an IMPLEMENTATION DEFINED PMU (PMUVer 0b1111),
__init_el2_fgt2 sets the PMUv3p9 bits in HDFGRTR2_EL2 and HDFGWTR2_EL2,
which are RES0 there. The macro only checks that PMUVer is at least
PMUv3p9, and 0b1111 passes.
Exclude 0b1111 before the compare, as __init_el2_debug does.
Fixes: 858c7bfcb35e1 ("arm64/boot: Enable EL2 requirements for FEAT_PMUv3p9")
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Bradley Morgan <brads@mainlining.org>
Reviewed-by: Anshuman Khandual <anshuman.khandual@arm.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux
Pull arm64 fixes from Will Deacon:
"Half of this is broken hardware (AMU counters and TLB invalidation)
and the other half is broken software (frequency scaling and signals).
So it seems as though we're all as bad as each other.
The AMU workaround is a little noisy, as it refactors an existing
workaround so that it can more easily be applied to additional CPUs.
Summary:
- Fix handling of CPU erratum #2645198 when batching pte updates
- Fix truncation of CPU frequency calculation by using 64-bit
arithmetic in arch_freq_get_on_cpu()
- Work around AMU erratum #3821522 on Cortex-A725
- Fix panic when trying to restore an SVE sigframe on a CPU that only
supports SME"
* tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux:
arm64/fpsimd: signal: Forbid non-streaming SVE payload on SME-only systems
arm64: errata: Add Cortex-A725 erratum 3821522 workaround
arm64: errata: Factor out broken AMU const counter cap
arm64: topology: fix arch_freq_get_on_cpu() overflow above 4.19 GHz
arm64: mm: Fix the break-before-make flush range for erratum 2645198
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/frank.li/linux into arm/fixes
iMX fixes for 7.3
- Fix incorrect pinmux values for LX2160A.
- Fix Sonata SD I/O supply to avoid filesystem I/O errors because the kernel
was disabling LDO5 during late init.
* tag 'imx-fixes-7.3' of https://git.kernel.org/pub/scm/linux/kernel/git/frank.li/linux:
arm64: dts: imx8mp-var-dart-sonata: Fix Sonata SD I/O supply
arm64: dts: lx2160a: fix the iic5 spi3 pinmux offset and value
arm64: dts: lx2160a: fix IIC1 pinmux submask rejected by pinctrl-single
arm64: dts: lx2160a: fix incorrect pinmux
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/qcom/linux into arm/fixes
Qualcomm Arm64 DeviceTree fix for v7.3
Fix microphone distortion on Hamoa and Purwa-based Surface Pro 11.
* tag 'qcom-arm64-fixes-for-7.3' of https://git.kernel.org/pub/scm/linux/kernel/git/qcom/linux:
arm64: dts: qcom: x1-denali: Fix microphone distortion
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix overflow of backward jump offset in constant blinding
(Alexei Starovoitov)
- Fix packet range of packet pointers sharing an id when var_off
tightens umax of one pointer and not the other (Alexei Starovoitov)
- Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc
(Alexei Starovoitov)
- Fix use-after-free of progs detached from busy trampolines: wait for
an RCU tasks grace period before freeing trampoline progs, and patch
detached progs out of trampoline images that are still in use
(Florent Revest)
- Hold map BTF for the memory allocator destructor record to fix UAF in
deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi)
- Fix missing migration protection in resizable hashtab
lookup_and_delete batch operation (Ömer Mete Kaya)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf:
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
bpf: Fix objects stuck in free_by_rcu_ttrace
bpf: Factor out __do_call_rcu_ttrace()
selftests/bpf: Test packet range of pointers sharing an id
bpf: Fix packet range of pointers sharing an id
selftests/bpf: Detach a trampoline prog while a task sleeps before it
bpf: Skip detached progs in trampoline images that are still in use
bpf: Wait for an RCU tasks grace period before freeing trampoline progs
bpf: Hold map BTF for the memory allocator destructor record
bpf: Fix overflow of jump offset in constant blinding
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/crng/random
Pull random number generator fixes from Jason Donenfeld:
- VMGENID memory needs to be mapped with the decrypted tag, so that
SEV-SNP machines can boot
- A fix for an initialization race in VMGENID, followed by a cleanup
- Trivial kernel doc cleanups in siphash and random.c
- A fix for a new compilation failure with recent clang on PPC and
RISC-V, due to generating an out-of-line memset in the vDSO
* tag 'random-7.3-rc6-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/crng/random:
random: vDSO: avoid call to memset() when zeroing reserved parameter
random: fix vgetrandom_opaque_params kernel-doc
random: vDSO: fix repeated word 'to' in comment
siphash: clean up kernel-doc comments
virt: vmgenid: move to using dev_set/get_drvdata
virt: vmgenid: set driver_data before registering notification handlers
virt: vmgenid: remap memory as decrypted
|
|
After a recent change in LLVM [1], builds with the random vDSO
implementation, such as PowerPC and RISC-V, fail when checking the vDSO:
arch/powerpc/kernel/vdso/vdso32.so.dbg: dynamic relocations are not supported
arch/riscv/kernel/vdso/vdso.so.dbg: dynamic relocations are not supported
memset() is now generated when zeroing params->reserved for some builds
because LLVM has an optimization (now run in more instances) that can
recognize at compile time when it is assigning a static value to a
contiguous area of memory and turn that into a call to memset(). Both
clang and GCC assume memset() is always available [2].
Clang has an internal fiddly hook, -max-store-memset, which we can set
to a high number, to disable generating out of line memset calls [3].
Similarly, GCC has -finline-stringops=memset to do the same [4], should
this issue ever hit future version of GCC. While these options wouldn't
make sense for normal kernel code, it is fine for the extremely limited
and intentionally compact vDSO code.
Link: https://github.com/llvm/llvm-project/commit/90cebef1411617fc3eedd359bdf00cb44b1c2439 [1]
Link: https://gcc.gnu.org/onlinedocs/gcc-16.2.0/gcc/Standards.html#index-ffreestanding [2]
Link: https://github.com/llvm/llvm-project/commit/b28eeb28bea39148738dc375e8a97072a1907e64 [3]
Link: https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#index-finline-stringops [4]
Closes: https://github.com/ClangBuiltLinux/linux/issues/2183
Cc: stable@vger.kernel.org # v6.12+
Signed-off-by: Nathan Chancellor <nathan@kernel.org>
Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/ebiggers/linux
Pull crypto library fixes from Eric Biggers:
- Fix a performance regression in certain AES encryption modes on
certain architectures, introduced this cycle
- Fix a small performance regression in the x86_64 optimized AES-GCM
code, introduced in 6.15
* tag 'libcrypto-fixes-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiggers/linux:
crypto: aes - Fix undesired override of some optimized AES modes
crypto: x86/aes-gcm - fix always true check for last AAD segment
|
|
Pull kvm fixes from Paolo Bonzini:
"The most intrusive change is reverting a commit from 7.3-rc1 that made
struct kvm a bit too large, and fixing the same issue otherwise.
There are again a lot of selftests lines; the sheer number of commits
is not small but I don't expect much more for 7.3 due to people
travelling to Plumbers next week.
ARM:
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
x86:
Various bugfixes where the guest could do stupid things on purpose to
cause problems in the host:
- Failed VMRUNs can cause pending TLB flushes to be dropped, and in
general some actions done through VMCB control fields have to be
redone if VMRUN fails
- Toggling MSR interceptions or eVMCS execution controls can cause
the host to use a stale MSR permission bitmap
- Bad page tables can cause a WARN.
Also fix issues in last week's pull request (my fault, for changing
email workflow and thus missing feedback sent to kvm@ but not LKML).
Generic:
- Take kvm_lock when creating vCPUs. For almost two decades everybody
thought it was not done for some unspecified performance reasons,
but in reality it was only done because kvm_lock was originally a
spinlock.
This is a better fix than 97d65b544f48 ("KVM: Check for duplicate
vcpu_id as early as possible", from the 7.3 merge window), and does
not waste 2K per VM, hence its inclusion here"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits)
KVM: arm64: Use stage-1 leaf size for VM_PFNMAP
KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF
KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM
KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs
KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT
KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor
KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq
KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing
KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state
KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it
KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12
KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt
KVM: selftests: Verify that L0's TPR doesn't get clobbered
KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2
KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified
KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted
KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN
KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded
...
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/kvmarm/kvmarm into HEAD
KVM/arm64 fixes for 7.3, round #2
- Take a reference on the last IRQ loaded into an LR to prevent it
from being freed while running the guest (Marc Zyngier)
- Ensure that the ITS MOVALL command only affects LPIs that were
previously affined to the source redistributor (Marc Zyngier)
- Fix + test for honoring the host's trap configuration when running
non-protected VMs while KVM is in protected mode (Fuad Tabba)
- Use the host stage-1 mapping granularity for VM_PFNMAP mappings at
stage-2 (Mostafa Saleh)
|
|
When commit 2aa53d68cee6 ("KVM: arm64: Try stage2 block mapping for
host device MMIO") added VM_PFNMAP support to get_vma_page_shift(),
stage-1 page tables did not support huge PFNMAP (as in VFIO-PCI
vfio_pci_mmap_huge_fault()) and transparent_hugepage_adjust() was
unsafe for MMIO as it dereferenced struct page.
Since commit 6011cf68c885 ("KVM: arm64: Walk userspace page tables to
compute the THP mapping size"), transparent_hugepage_adjust() instead
walks the host stage-1 page tables via get_user_mapping_size() without
touching struct page. Meanwhile, commit 3e509c9b03f9 ("mm/arm64:
support large pfn mappings") enabled stage-1 huge PFNMAP.
So. we can drop the VMA-based VM_PFNMAP size calculation in
get_vma_page_shift() and let transparent_hugepage_adjust() derive
the stage-2 mapping size directly from the populated stage-1 leaf
for non-cacheable mappings as well.
Assisted-by: LLM
Cc: stable@vger.kernel.org
Fixes: 2aa53d68cee6 ("KVM: arm64: Try stage2 block mapping for host device MMIO")
Reviewed-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
When system_supports_sme() is true but system_supports_sve() is false,
restoring a specifically crafted SVE signal context can result in the
task erroneously having non-streaming SVE state. Subsequent attempts to
manipulate the task's FPSIMD/SVE/SME state can result in a variety of
problems, including fatal EL1 UNDEFs.
In such configurations, the kernel always creates an SVE signal context
when delivering a signal, and this can only be in one of two states:
(1) SVE_SIG_FLAG_SM is set, and an SVE payload is present containing
streaming mode SVE state. The recorded VL is the task's live
streaming VL.
(2) SVE_SIG_FLAG_SM is clear, and an SVE payload is not present. The
FPSIMD context contains the non-streaming mode FPSIMD state. The
recorded VL is 0.
Currently restore_sve_fpsimd_context() correctly rejects cases where
SVE_SIG_FLAG_SM is set and an SVE payload is not present, but fails to
reject cases where SVE_SIG_FLAG_SM is clear and an SVE payload is
present. Consequently, restore_sve_fpsimd_context() can place the task
in a state where it has non-streaming SVE state even when this is not
supported by HW.
For example, this can cause a later EL1 UNDEF when the kernel attempts to
restore the task's ZCR_EL1 value:
| # ./sme-sigcontext-to-sve
| Internal error: Oops - Undefined instruction: 0000000002000000 [#1] SMP
| Modules linked in:
| CPU: 0 UID: 0 PID: 131 Comm: sme-sigcontext- Not tainted 7.3.0-rc1 #1 PREEMPT
| Hardware name: linux,dummy-virt (DT)
| pstate: 61402009 (nZCv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
| pc : fpsimd_restore_current_state+0x258/0x458
| lr : exit_to_user_mode_loop+0xb8/0x188
| sp : ffff80008056be40
| x29: ffff80008056be40 x28: fff00000c1670000 x27: 0000000000000000
| x26: 0000000000000000 x25: 0000000000000000 x24: 0000000000000000
| x23: ffff80008056bec0 x22: 0000000000000008 x21: 0000000000000040
| x20: 0000000000000081 x19: 0000000008800010 x18: 0000000000000000
| x17: 0000fffffd412570 x16: 0000000000001000 x15: 0000fffffd4123b0
| x14: 0000fffffd412780 x13: 0000fffffd412be8 x12: 0000000047435300
| x11: 0000fffffd412570 x10: 0000000000000000 x9 : 0000000045585401
| x8 : fff00000c18a6c44 x7 : 0000000000000000 x6 : 0000000000000002
| x5 : 0000000000000002 x4 : ffff800080568000 x3 : 0000000000000001
| x2 : 0000000008800010 x1 : fff00000c1670000 x0 : 0000000008800000
| Call trace:
| fpsimd_restore_current_state+0x258/0x458 (P)
| exit_to_user_mode_loop+0xb8/0x188
| el0_svc+0x1cc/0x1d0
| el0t_64_sync_handler+0xa0/0xe4
| el0t_64_sync+0x198/0x19c
| Code: d5384101 f9400020 53175c03 36b80de0 (d5381202)
| ---[ end trace 0000000000000000 ]---
| Kernel panic - not syncing: Oops - Undefined instruction: Fatal exception in interrupt
| Kernel Offset: 0x291fb1000000 from 0xffff800080000000
| PHYS_OFFSET: 0x40000000
| CPU features: 0x0,00000000,0052802f,ffb88f43,3afcf73f
| Memory Limit: none
Rework restore_sve_fpsimd_context() to reject cases where
SVE_SIG_FLAG_SM is clear and an SVE payload is not present. As
parse_user_sigframe() rejects SVE signal frames when neither SVE nor SME
are supported, it isn't necessary for restore_sve_fpsimd_context() to
handle the case where neither are supported.
Fixes: 7dde62f0687c ("arm64/signal: Always accept SVE signal frames on SME only systems")
Signed-off-by: Mark Rutland <mark.rutland@arm.com>
Reviewed-by: Mark Brown <broonie@kernel.org>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: stable@vger.kernel.org
Signed-off-by: Will Deacon <will@kernel.org>
|
|
Cortex-A725 erratum 3821522 affects the CNT_CYCLES event, which can
incur a significant increment error when a CPU enters and subsequently
exits WFE or WFI, and may no longer track the system counter
frequency.
The AMEVCNTR01_EL0 counter is used as the AMU constant counter for
frequency invariance and CPPC FFH feedback counters. Wire the affected
Cortex-A725 range into the shared broken AMU constant-counter capability
so the affected counter is treated as unavailable by returning zero in
the AMU counter paths. This prevents the broken counter from being used
as a reference source.
The erratum can also affect PMUv3 users of the CNT_CYCLES event,
but this workaround intentionally does not change PMU event handling.
Hiding or rejecting the PMU event from the erratum code would change
the perf-visible PMU event interface, including raw event selection,
and would need a separate PMU-specific approach rather than being
folded into the AMU reference-counter workaround.
Cc: stable@vger.kernel.org
Signed-off-by: Beata Michalska <beata.michalska@arm.com>
Reviewed-by: Vladimir Murzin <vladimir.murzin@arm.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
Move the workaround from the erratum 2457168-specific cpucap to a generic
broken AMU constant-counter one. This keeps the existing Cortex-A510
handling unchanged while allowing other errata with similar AMU constant
counter issue to share the capability bit and call sites.
Signed-off-by: Beata Michalska <beata.michalska@arm.com>
Reviewed-by: Vladimir Murzin <vladimir.murzin@arm.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
perf_iterate_sb() invokes its callback for each matching perf_event on
the CPU and task context, passing a shared caller-allocated event
structure.
perf_event_header__init_id() mutated header->size in place by adding
event->id_header_size, requiring sideband output callbacks to save and
restore header fields across iterations. Three sideband callbacks failed
to save and restore header.size around perf_event_header__init_id():
- perf_event_ksymbol_output()
- perf_event_bpf_output()
- perf_event_text_poke_output()
When multiple events with attr.ksymbol, attr.bpf_event, or
attr.text_poke and sample_id_all are active on the same CPU, each
subsequent event receives a record whose header.size is inflated by all
preceding events' id_header_size values while only a single id_sample is
written, leaving uninitialized ring-buffer bytes at the end of the
record and causing userspace perf to fail with -EFAULT ("Bad address")
when parsing the sample_id trailer.
Similarly, perf_event_mmap_output() set PERF_RECORD_MISC_MMAP_BUILD_ID
in mmap_event->event_id.header.misc when event->attr.build_id was
enabled, but only saved and restored header.size and header.type. If an
event with attr.build_id was followed by an event with attr.mmap2 and
!attr.build_id, the second event received PERF_RECORD_MISC_MMAP_BUILD_ID
in header.misc while its payload contained maj/min/ino/ino_generation
instead of a build ID.
Rather than splitting header initialization between callers and output
callbacks and saving/restoring mutated header fields, replace
perf_event_header__init_id() with perf_event_header__init(), which
initializes header->type, header->misc, and header->size alongside the
sample_id fields on each invocation.
Fixes: 76193a94522f ("perf, bpf: Introduce PERF_RECORD_KSYMBOL")
Fixes: 6ee52e2a3fe4 ("perf, bpf: Introduce PERF_RECORD_BPF_EVENT")
Fixes: e17d43b93e54 ("perf: Add perf text poke event")
Fixes: 88a16a130933 ("perf: Add build id data in mmap2 event")
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260929222332.973435-1-irogers@google.com
Cc: stable@vger.kernel.org
|
|
In gcm_process_assoc(), a segment that is not the last one must have its
length rounded down to a multiple of 16 bytes, as required by the
assembly. The check for a non-last segment is `if (unlikely(assoclen))
/* Not the last segment yet? */` where assoclen is the number of AAD
bytes remaining after the current segment.
Since the conversion to the new scatterwalk API, assoclen is decremented
at the end of the loop body rather than the beginning, so it still
includes the current segment when the check executes, making the check
always true.
As a result the last segment was rounded down as well, causing some
avoidable extra work: an additional memcpy into the temporary buffer and
an additional call into the assembly after the loop. The GCM
authentication tag is unaffected either way, so the self-tests pass and
this went unnoticed. It is purely an efficiency issue rather than a
correctness one.
Fix this by moving the assoclen decrement back to the beginning of the
loop body.
Fixes: e9787deff49e ("crypto: x86/aes-gcm - use the new scatterwalk functions")
Cc: stable@vger.kernel.org
Suggested-by: Eric Biggers <ebiggers@kernel.org>
Signed-off-by: Mohamad Raizudeen <raizudeen.kerneldev@gmail.com>
Link: https://patch.msgid.link/20260926145903.6061-1-raizudeen.kerneldev@gmail.com
Signed-off-by: Eric Biggers <ebiggers@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vgupta/arc
Pull ARC fix from Vineet Gupta
- cmpxchg snafu spotted by Bradley and other misc fixes
* tag 'arc-fixes-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vgupta/arc:
ARC: arch_cmpxchg_relaxed to use size of pointed type not pointer
arc: kernel: Fix clk reference leak in show_cpuinfo()
arc: remove unused profile.h includes
ARC: cleanup dead ARC_CANT_LLSC option in Kconfig
|
|
For each vCPU, pKVM's EL2 builds its own HCR_EL2 and takes only TWI, TWE
and VSE from the host's value on each entry. For a non-protected VM it
has fallen behind the host's: on a CPU with MTE the guest can read
GMID_EL1, on one with FEAT_EVT2 but no FGT it can execute a TLBI OS its
ID registers hide, an interrupt injected without a vGIC (VI, VF) never
arrives, and set/way emulation loses the TVM trap it relies on.
For a non-protected VM, also take the bits the host varies with the VM's
configuration or at runtime: VI, VF, TVM, TID2, TID4, TID5 and TTLBOS.
They only pend interrupts for, or add traps to, a VM the host controls.
The traps exit to the host, which handles them as it does without pKVM.
The bits that change what EL2 does on entry and exit (E2H, RW, API/APK)
or depend only on the CPU (TEA, TERR, FWB) stay EL2's. A protected VM
still takes only TWI, TWE and VSE.
EL2 now sets TID2 or TID4 only for a protected VM, and never sets ATA,
as the host rejects KVM_CAP_ARM_MTE in pKVM.
Fixes: b56680de9c648 ("KVM: arm64: Initialize trap register values in hyp in pKVM")
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Link: https://patch.msgid.link/20260929090031.3829185-4-fuad.tabba@linux.dev
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
pKVM keeps its own copy of each vCPU's HCR_EL2 at EL2, initialised with
RW set unconditionally. Nothing stops userspace from creating a
non-protected AArch32 VM, and with RW set, entering one of its vCPUs is
an illegal exception return: KVM_RUN fails with KVM_EXIT_FAIL_ENTRY.
Clear RW for a vCPU that is AArch32 at EL1, when the CPU has AArch32
EL1. On a CPU without it, RW stays set and the entry still fails, rather
than EL2 switching *32_EL2 registers that are UNDEFINED there. The host
already rejects such a vCPU, but EL2 takes the vCPU's features from the
host and doesn't rely on that.
Fixes: b56680de9c648 ("KVM: arm64: Initialize trap register values in hyp in pKVM")
Cc: stable@vger.kernel.org
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Link: https://patch.msgid.link/20260929090031.3829185-3-fuad.tabba@linux.dev
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
FGUs don't actually require hardware support, so don't gate fgt trap
config insertion on FGT. This is the only action required to allow FGUs
always, as kvm_calculate_traps() already computes the FGU bits with or
without FEAT_FGT.
This also fixes a spurious WARN when a TLBI OS is run on a guest in vEL1
without FEAT_TLBIOS on non FEAT_FGT hardware. In this configuration the
course-grained HCR_EL2.TTLBOS trap reaches the handler handle_tlbi_el1(),
which expects an EL1 TLBI only from vEL2. With FGUs enabled the UNDEF
will be injected without needing the specific handler.
Fixes: f5a5a406b4b8b ("KVM: arm64: Propagate and handle Fine-Grained UNDEF bits")
Cc: stable@vger.kernel.org
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Wei-Lin Chang <weilin.chang@arm.com>
Link: https://patch.msgid.link/20260929090031.3829185-2-fuad.tabba@linux.dev
[oliver: Take Wei-Lin's suggested changelog]
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
The MOVALL command moves the pending state of all LPIs targeting a given
redistributor to another one. However, we seem to have lost the
filtering on the source RD, which means we move all LPIs to the target.
Not quite what the spec mandates.
Hack update_affinity() to take an optional source vcpu that is used as a
filter when non-NULL, restoring the filtering that was performed by
vgic_copy_lpi_list() back in the days.
Fixes: 11f4f8f3e6e06 ("KVM: arm64: vgic-its: Walk LPI xarray in vgic_its_cmd_handle_movall()")
Signed-off-by: Marc Zyngier <maz@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260929093548.3598547-7-maz@kernel.org
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
Referencing the last interrupt inserted in an LR is rather fragile, as
this interrupt can vanish if a concurrently unmapped LPI.
Solve this by bumping up the refcount on the interrupt when populating
last_lr_irq, and drop it at vgic_prune_ap_list() time, when LPIs are
being reclaimed.
Fixes: 6da5e537f5afe ("KVM: arm64: vgic: Pick EOIcount deactivations from AP-list tail")
Reported-by: Yuchao Zhang <ndaugoing@gmail.com>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Tested-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Marc Zyngier <maz@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260929093548.3598547-5-maz@kernel.org
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
last_lr_irq is always populated when there is any interrupt populated in
the AP list. Not only this is not necessary (it is only useful when we
completely fill the LRs), but this is in the way of further fixes.
Make sure last_lr_irq is kept to NULL when we LRs are not completely
full.
Fixes: 6da5e537f5afe ("KVM: arm64: vgic: Pick EOIcount deactivations from AP-list tail")
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Tested-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Marc Zyngier <maz@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260929093548.3598547-4-maz@kernel.org
Signed-off-by: Oliver Upton <oupton@kernel.org>
|
|
bpf_tramp_image_put() makes sure a trampoline image is not freed while
a task may still be running in it, but nothing similar is done for the
progs called by that image. Detach drops the last prog reference right
away and the prog is freed after grace periods, on the basis that a
task still in the traced function skips the fexit progs once the nop at
ip_after_call is patched to a jump.
That leaves out a task sleeping in a sleepable prog that runs before
the detached one, which no grace period waits for:
CPU 0 CPU 1
in image I, sleeping in prog S
detach P from I's trampoline
-> new image, bpf_tramp_image_put(I)
bpf_prog_put(P), last ref
grace periods, P freed
back from S
__bpf_prog_enter(P)
call P->bpf_func
If S and P are fexit progs the task is already past the patched jump,
and fentry only images don't have one. On x86 this is an int3 in
poisoned bpf_prog_pack memory:
Oops: int3: 0000 [#1] SMP NOPTI
CPU: 18 UID: 0 PID: 94573 Comm: x169 Not tainted 6.18.44 #1 PREEMPT(lazy)
RIP: 0010:0xffffffffc0601d8d
Call Trace:
<TASK>
? bpf_trampoline_6442515411+0x1a4/0x21b
bpf_lsm_bprm_committed_creds+0x5/0x10
security_bprm_committed_creds+0x5f/0x70
begin_new_exec+0x2d6/0x410
...
We hit this in production when progs attached through trampolines got
detached while their hooks were busy, and it was independently found
with a fuzzer and KASAN.
Have the JITs emit a patchable nop in front of each prog call sequence
and record it in the image. When a prog is detached, patch its nop to a
jump over the call sequence. Progs that stay attached keep running for
the tasks that are in the image, and ip_after_call isn't needed anymore.
The task can be in any image that isn't freed yet, not only in the
current one. It sleeps in image I1 that calls S, P and Q, then P is
detached and the trampoline moves to image I2, then Q is detached and
its call is still in I1. So the trampoline keeps a list of its images
until they are freed, and detaching a prog patches its nop in all of
them. Images hold a reference on the trampoline for that long.
On riscv and loongarch a jump of any range takes several instructions,
and a task preempted in the middle of them could resume into half of the
new sequence. The nop is a single instruction there, patched to a near
branch through arch_bpf_trampoline_skip().
With the extra nops, BPF_MAX_TRAMP_LINKS progs no longer fit in a page
on x86 and arm64 (and already didn't on powerpc), so lower the limit
there like s390 does.
Fixes: e21aa341785c ("bpf: Fix fexit trampoline.")
Reported-by: Sechang Lim <rhkrqnwk98@gmail.com>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260926135605.1217928-3-florent.revest@linux.dev
Closes: https://lore.kernel.org/bpf/20260815071927.147049-1-zirajs7@gmail.com/
|
|
Explicitly flush caches after intra-host migration before clearing "SEV
active" on the source VM, as doing cache maintenance afterwards creates a
tiny window where memory reclaim could return memory to the host without
performing a cache flush, e.g. as pointed out by Sashiko:
CPU1 in sev_migrate_from():
src->active = false;
CPU2 running concurrent unmap:
Since active is false, the automatic cache flush in
sev_guest_memory_reclaimed is skipped.
The host frees and reallocates the page.
CPU1 in sev_migrate_from():
sev_writeback_caches(src_kvm);
Executes a hardware cache flush (wbnoinvd), which writes the guest's old
dirty ciphertext over the new page owner's data.
Fixes: 93de2a6a4b91 ("KVM: SEV: Do cache maintenance on the source VM during intra-host migration")
Cc: stable@vger.kernel.org
Reported-by: Sashiko Bot <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/all/20260923165304.1662E1F000FF@smtp.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260928154644.2559454-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Nullify have_run_cpus when freeing the mask, particularly in the error path
of __sev_guest_init(), so that KVM doesn't have to subtly use sev->active
to track whether or not the mask has been freed. As pointed out by Sashiko,
blindly freeing the mask in sev_vm_destroy() results in a double-free if
the mask is freed if __sev_guest_init() fails.
Throw the logic in a helper as nullifying the pointer is frustratingly
difficult and weird due to have_run_cpus being a single-entry array when
CPUMASK_OFFSTACK=n. Deliberately don't use CPUMASK_VAR_NULL, as it's not
directly assignable when the cpumask is on-stack, e.g. requires using a
local variable and a memcpy(), which is beyond ridiculous. Furthermore,
while clearing the on-stack bitmask is an unnecessary and arguably unwanted
side effect, KVM absolutely relies on '0' being the "null" value given that
the struct is zero-allocated.
Opportunistically add an alloc() helper to pair with free(); there are just
enough call sites to make doing so worthwhile.
Fixes: 12c1f6e03f94 ("KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV")
Cc: stable@vger.kernel.org
Reported-by: Sashiko Bot <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/all/20260923165349.CAAF01F000FF@smtp.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260928154644.2559454-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
__FRAME_OFFSETS
The register offset macros in <asm/ptrace-abi.h> are guarded by
`defined(__ASSEMBLER__) || defined(__FRAME_OFFSETS)` for 64-bit, but
were left unguarded for 32-bit. This causes havoc for userspace that
happens to use identifiers colliding with these short macro names
(e.g., EBX, ECX, EAX, DS, ES, FS, GS, CS, SS). Without this guard,
userspace is forced to be super extra careful with include ordering to
minimize the chance of collision.
Wrap both the 32-bit and 64-bit register definitions under
`#if defined(__ASSEMBLER__) || defined(__FRAME_OFFSETS)`, and ensure
User-Mode Linux (UML) defines `__FRAME_OFFSETS` for 32-bit as well.
Closes: https://github.com/llvm/llvm-project/issues/217413
Assisted-by: LLM
Signed-off-by: Nick Desaulniers <ndesaulniers@google.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Johannes Berg <johannes@sipsolutions.net>
Tested-by: Elliott Hughes <enh@google.com>
Link: https://patch.msgid.link/20260821-ptrace_uapi-v1-1-3de8638a29f2@google.com
|
|
On a box with a discrete GPU, lockdep reports a possible deadlock as
soon as kswapd shrinks the TTM page pool. The immediate cause is an
x86 commit that added an mmap_read_lock() to kernel page protection
munging code.
The huge vmap code holds the same lock over a GFP_KERNEL allocation,
which is a no-no now that reclaim can take it. That allocation is in a
page table *free* path and ends up being for dubious purposes[1].
Basically, it tries to avoid hardware setting Accessed=1 in page table
entries that are unreachable by the hardware, a non-issue.
Remove the PMD copy. Detach the original PMD page at the PUD, flush
the mid-level caches, and free the PTE tables straight from the
detached PMD page. With no allocation left, the locking issue is gone.
Lockdep splat/analysis:
WARNING: possible circular locking dependency detected
7.3.0-rc3-f6e7b42bf05b+ #183 Tainted: G U
------------------------------------------------------
kswapd0/269 is trying to acquire lock:
((init_mm).mmap_lock){++++}-{4:4}, at: change_page_attr_set_clr+0x29a/0x4a0
but task is already holding lock:
(pool_shrink_rwsem){.+.+}-{4:4}, at: ttm_pool_shrink+0xb2/0x330 [ttm]
Chain exists of:
(init_mm).mmap_lock --> fs_reclaim --> pool_shrink_rwsem
The cycle is built from three edges:
1) pool_shrink_rwsem -> (init_mm).mmap_lock
The TTM shrinker restores the caching attribute of every page it
frees, while holding pool_shrink_rwsem:
ttm_pool_shrink()
-> ttm_pool_dispose_list()
-> ttm_pool_free_page()
-> set_pages_wb()
-> change_page_attr_set_clr() [ init_mm mmap read lock ]
2) fs_reclaim -> pool_shrink_rwsem
The same shrinker, called from reclaim.
3) (init_mm).mmap_lock -> fs_reclaim
ioremap() installing a huge PUD mapping over an existing PMD table:
ioremap_page_range()
-> vmap_range_noflush()
-> vmap_try_huge_pud() [ init_mm mmap read lock ]
-> pud_free_pmd_page()
-> __get_free_page(GFP_KERNEL) [ enters reclaim ]
[ dhansen: Lots of changelog munging/trimming and merged comments from my
version of the fix. ]
Fixes: d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF")
Suggested-by: Pedro Falcato <pfalcato@suse.de>
Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Link: https://lore.kernel.org/20260916062222.27347-1-mikhail.v.gavrilov@gmail.com
Link: https://lore.kernel.org/all/e11449f0-d9ad-4d1b-ab21-2be7d71fe335@intel.com/ [1]
Link: https://patch.msgid.link/20260923223116.20090-1-mikhail.v.gavrilov@gmail.com
Cc: stable@vger.kernel.org
|
|
get_align_mask() returns huge_page_mask_align() for hugetlbfs, but get_align_bits()
adds va_align.bits regardless, so vm_unmapped_area() returns an address off the huge
page boundary and __unmap_hugepage_range() hits BUG_ON(start & ~huge_page_mask(h)) at
teardown.
This can be triggered on Carrizo and FX-8370E, both hstates.
Pass the file to get_align_bits() and skip the randomisation for hugetlbfs.
[ bp: Massage commit message. ]
Fixes: 1317a5e7f7b1 ("arch/x86: teach arch_get_unmapped_area_vmflags to handle hugetlb mappings")
Suggested-by: Dave Hansen <dave.hansen@intel.com>
Acked-by: Dave Hansen <dave.hansen@intel.com>
Signed-off-by: Laurent Wandrebeck <l.wandrebeck@quelquesmots.fr>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: stable@vger.kernel.org # 6.13+
Link: https://patch.msgid.link/20260922085032.46144-1-l.wandrebeck@quelquesmots.fr
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Ingo Molnar:
- Fix preemption bugs in the SVSM vTPM guest implementation
(Melody Wang)
- Fix MCE-triggered hardware debug register corruption on
task migration (Masami Hiramatsu)
* tag 'x86-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/mce: Fix hardware debug register corruption on task migration
x86/sev: Make vTPM SVSM calls preemption-safe
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci
Pull PCI fixes from Bjorn Helgaas:
- Make BAR resize work even for devices where no upstream bridge is
visible to the OS, which fixes an amdgpu regression on SolidRun
HoneyComb, which doesn't expose Root Ports to the OS (Liz Fong-Jones)
- Omit bus properties in dynamic OF nodes when a bridge has no
subordinate bus, which fixes early boot hangs caused by NULL pointer
dereferences with CONFIG_PCI_DYNAMIC_OF_NODES enabled (Angel J)
- Disable enhanced atomics on AMD NBIO 7.7 and 7.11 to avoid silent
data corruption on 64-bit DMAs (Mario Limonciello)
* tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11
PCI: of_property: Omit bus properties without a subordinate bus
PCI: Fix BAR resize for devices on a root bus
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
Force a refresh of the vmcs02 MSR bitmap during nested VM-Enter if the
runtime eVMCS controls (pin, primary, secondary, etc.) are being updated.
If L1 isn't intercepting TPR writes, runs L2 with TPR virtualization, and
then runs the same L2 with TPR virtualization disabled, KVM will fail to
refresh msr_bitmap02 and leave TPR in passthrough mode even though TPR
virtualization is disabled. I.e. failure to refresh the bitmap lets L2 (or
L1 by proxy) read and write L0's TPR.
Fixes: 502d2bf5f2fd ("KVM: nVMX: Implement Enlightened MSR Bitmap feature")
Cc: stable@vger.kernel.org
Reviewed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Use the MSR permission bitmap of the active VMCB instead of assuming that
KVM is always using vmcb02's bitmap when L2 is active, as KVM uses msrpm02
if and only if L1 wants to intercept MSR accesses, i.e. if and only if KVM
needs to merge msprm01 with msrpm12.
Don't bother tracking the virtual address of the bitmap that's being used,
as __va() is cheap on x86, and caching the virtual address would introduce
yet another source of potentially stale information.
Fixes: b2ac58f90540 ("KVM/SVM: Allow direct access to MSR_IA32_SPEC_CTRL")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260826195833.844526-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Don't (re)read PERF_CNTR_GLOBAL_CTL from hardware on a failed VMRUN, as the
purpose of the read is to synchronize KVM's cache with any writes done by
the guest, and the guest can't possibly have modified the MSR if it never
got a chance to run.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Don't mark the ASID as dirty in the VMCB when requesting a TLB flush via
control.tlb_ctl. Per "15.15.3 VMCB Clean Field" of the July 2026, Revision
3.45 version of the APM:
The following are explicitly not cached and not represented by Clean bits:
* TLB_Control
Fixes: 7e8e6eed75e2 ("KVM: SVM: Move asid to vcpu_svm")
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Leave control.erap_ctl and control.clean as-is in the VMCS if VMRUN fails,
because as per AMD:
there's no explicit architectural guarantee about the behavior in the
presence of VMRUN failures. So the best thing to do would be to assume
that if VMRUN fails, the actions requested in the control fields may not
have been performed.
Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Don't reset the VMCB's TLB control back to "do nothing" on a failed VMRUN,
as empirical testing shows that the CPU performs the requested TLB flush if
and only if VMRUN is successful, i.e. clearing TLB control on a failed
VMRUN effectively drops a TLB flush.
Explicitly track the need to flush all ASIDs on a per-CPU basis, as the
ASID reuse condition is tied to the pCPU, not to the vCPU. As a bonus,
this also obviates the need to avoid clobbering FLUSH_ALL_ASID with
TLB_CONTROL_FLUSH_ASID, e.g. in svm_flush_tlb_asid().
Deliberately don't bother saving/restoring the "old" tlb_ctl on failure,
in quotes because it's not exactly the old tlb_ctl, it's the tlb_ctl from
after pre_svm_run(), but before updating tlb_ctl for flush_all_asids. If
VMRUN fails and TLB_CONTROL_FLUSH_ALL_ASID is forced, then the next
successful run of the VMCB *may* unnecessarily flush all ASIDs, which
strictly speaking could result in noisy neighbor issues. However, the
fact that new_asid() is already guest-triggerable, because of KVM's flawed
behavior of clearing the ASID on emulated INIT, means that a guest can
already trigger a flush of all ASIDs at roughly the same rate. And once
KVM stops clobbering the ASID on emulated INIT, *or* assigns a static ASID
to each vCPU, this flaw goes away.
Fixes: 38e5e92fe8c0 ("KVM: SVM: Implement Flush-By-Asid feature")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Cc: Tom Lendacky <thomas.lendacky@amd.com>
Cc: Jim Mattson <jmattson@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
When walking shadow page tables, immediately terminate the walk if the root
is a "dummy" root, i.e. a root whose top-level page table is backed by the
zero page, but otherwise doesn't exist. If memslot creation races with a
stage-2 page fault (EPT violation or #NPF) from L2, then if the stars align,
KVM will attempt to walk shadow page tables using the zero page and hit a
NULL pointer deref.
BUG: kernel NULL pointer dereference, address: 0000000000000021
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
PGD 0 P4D 0
Oops: Oops: 0000 [#1] SMP
CPU: 30 UID: 1000 PID: 941 Comm: qemu Not tainted 7.2.0-rc2-1b731e5ded48-next-vm #1741 PREEMPTLAZY
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 0.0.0 02/06/2015
RIP: 0010:__kvm_mmu_invalidate_addr+0xea/0x210 [kvm]
Call Trace:
<TASK>
kvm_mmu_invalidate_addr+0x92/0xe0 [kvm]
__kvm_inject_emulated_page_fault+0x67/0x80 [kvm]
ept_page_fault+0x160/0x850 [kvm]
kvm_mmu_do_page_fault+0x102/0x1f0 [kvm]
kvm_mmu_page_fault+0x8e/0x6b0 [kvm]
vmx_handle_exit+0x163/0x640 [kvm_intel]
kvm_arch_vcpu_ioctl_run+0x960/0x2120 [kvm]
kvm_vcpu_ioctl+0x2c7/0x970 [kvm]
__x64_sys_ioctl+0x90/0xd0
do_syscall_64+0x67/0x5f0
entry_SYSCALL_64_after_hwframe+0x4b/0x53
RIP: 0033:0x7f606d0b53bb
Opportunistically harden the shadow walks against fully invalid roots, but
WARN, as all callers are expected/required to pre-check for a valid root.
Don't WARN in the dummy root case as the whole point of using a dummy root
is to provide a root that's valid enough to enter the guest, i.e. it should
Just Work for all flows except those that *need* to know about dummy roots.
Alternatively, KVM could allocate a dedicated page and associated shadow
page structure for the dummy root, which is very tempting as it such an
approach should be more resilient against unexpected behavior. But that
would be a much larger and thus riskier change than simply terminating
walks of dummy roots.
Fixes: 0e3223d8d00a ("KVM: x86/mmu: Use dummy root, backed by zero page, for !visible guest roots")
Reported-by: Gabriel Schneider <gbrls@osec.io>
Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260925234256.2384816-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
When creating a vCPU, don't drop kvm->lock to when doing the bulk of actual
vCPU creation, as allowing multiple vCPUs to be created in parallel adds
significant complexity in KVM (as evidenced by the many related bugs), and
all known VMMs fully serialize vCPU creation. Remove all manually locking
of kvm->lock from kvm_arch_vcpu_{post,}create() for obvious reasons.
For many years, "everyone" has assumed that dropping kvm->lock was done for
performance reasons optimization, e.g. to allow userspace to create all
vCPUs concurrently for latency purposes. But as above, no known VMM does
that. Looking at the history of this code, before commit 11ec28047118
("KVM: Convert vm lock to a mutex"), kvm->lock was a spinlock. I.e. KVM
*had* to drop kvm->lock when doing the bulk of vCPU creation, otherwise KVM
couldn't do normal memory allocations. When kvm->lock got turned into a
mutex for unrelated reasons, no one took advantage updated of the change to
simplify vCPU creation. And 19 years later, everyone just assumed that KVM
continued to deal with the complexity for performance reasons.
Furthermore, naively parallelizing vCPU creation in userspace is likely a
net negative due to the overheads of task creation. Unless a VMM carefully
avoids the extra overhead related to parallelization, e.g. spawns each
vCPU's thread before creating the vCPU, creating vCPUs concurrently is a
net *negative* up until about ~64 vCPUs, after which the times are a wash.
The absolute speed of light _is_ faster if KVM doesn't hold kvm-lock, but
at vCPU counts of ~16 or less, it's probably in the noise when considering
total VM creation time, as the added latency is less than 1ms up until 16
or so vCPUs.
On top of all that, KVM has had a *lot* of fatal bugs (most often found by
syzkaller) related to vCPUs being created while trying to do per-VM
operations (basically, see every flow that locks all vCPUs). I.e. the
parallel vCPU creation "support" is actively harmful as the only "use case"
is for misbehaving userspace to exploit KVM bugs.
Serializing vCPU creation will allow reverting commit 97d65b544f48 ("KVM:
Check for duplicate vcpu_id as early as possible"), which had "minor" math
error: the worst case scenario isn't "256 bytes per VM", it's "256 unsigned
longs per VM", i.e. 2048 bytes per VM, which doubles the size of each VM
and pushes several architectures into order-1 allocations.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
equivalent
Use kvm_is_vcpu_creation_in_progress() instead of an open-coded equivalent
during AIA initialization. Unlike similar vGIC code in arm64, the relevant
RISC-V code runs under kvm->lock, i.e. can use the standard API without
hitting lockdep false positive.
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
Now that KVM's APIs for locking all vCPUs return -EBUSY if vCPU creation is
in-progress, drop the manual check for the same from vGIC creation, and
update the comments accordingly.
Note, while KVM arm64 guards many vGIC operations with its arch-specific
config_lock, holding kvm->lock is sufficient to guarantee a stable result
for "is vCPU creation in-progress". So, no functional change intended.
Note #2, the open coded check in vgic_init() is racy when called without
kvm->lock held, e.g. via vgic_lazy_init(). I.e. that check needs to stay
open coded to avoid triggering a lockdep assert. Whether or not the race
is "fine" is a problem for a different day.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net>
Tested-by: Naveen N Rao (AMD) <naveen@kernel.org>
Message-ID: <20260921174445.911676-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|