summaryrefslogtreecommitdiff
path: root/arch/x86
AgeCommit message (Collapse)Author
5 daysMerge tag 'x86-urgent-2026-10-04' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Ingo Molnar: - Don't apply va_align to hugetlb mappings on AMD F15h systems that have custom va_align.bits values (Laurent Wandrebeck) - Fix PMD teardown handling regression flagged by lockdep (Mikhail Gavrilov) - Hide ptrace header register offset macros behind __ASSEMBLER__ or __FRAME_OFFSETS, to fix user-space build errors that may trigger if they happen to shadow these short and generic macro names (Nick Desaulniers) * tag 'x86-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: {x86,um}/uapi/ptrace: Guard register offset macros with __ASSEMBLER__ or __FRAME_OFFSETS x86/mm: Drop unnecessary PMD page copy when freeing x86/mm: Don't apply va_align to hugetlb mappings on AMD F15h
7 daysMerge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpfLinus Torvalds
Pull bpf fixes from Alexei Starovoitov: - Fix overflow of backward jump offset in constant blinding (Alexei Starovoitov) - Fix packet range of packet pointers sharing an id when var_off tightens umax of one pointer and not the other (Alexei Starovoitov) - Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc (Alexei Starovoitov) - Fix use-after-free of progs detached from busy trampolines: wait for an RCU tasks grace period before freeing trampoline progs, and patch detached progs out of trampoline images that are still in use (Florent Revest) - Hold map BTF for the memory allocator destructor record to fix UAF in deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi) - Fix missing migration protection in resizable hashtab lookup_and_delete batch operation (Ömer Mete Kaya) * tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch() selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace bpf: Fix objects stuck in free_by_rcu_ttrace bpf: Factor out __do_call_rcu_ttrace() selftests/bpf: Test packet range of pointers sharing an id bpf: Fix packet range of pointers sharing an id selftests/bpf: Detach a trampoline prog while a task sleeps before it bpf: Skip detached progs in trampoline images that are still in use bpf: Wait for an RCU tasks grace period before freeing trampoline progs bpf: Hold map BTF for the memory allocator destructor record bpf: Fix overflow of jump offset in constant blinding
7 daysMerge tag 'random-7.3-rc6-for-linus' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/crng/random Pull random number generator fixes from Jason Donenfeld: - VMGENID memory needs to be mapped with the decrypted tag, so that SEV-SNP machines can boot - A fix for an initialization race in VMGENID, followed by a cleanup - Trivial kernel doc cleanups in siphash and random.c - A fix for a new compilation failure with recent clang on PPC and RISC-V, due to generating an out-of-line memset in the vDSO * tag 'random-7.3-rc6-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/crng/random: random: vDSO: avoid call to memset() when zeroing reserved parameter random: fix vgetrandom_opaque_params kernel-doc random: vDSO: fix repeated word 'to' in comment siphash: clean up kernel-doc comments virt: vmgenid: move to using dev_set/get_drvdata virt: vmgenid: set driver_data before registering notification handlers virt: vmgenid: remap memory as decrypted
8 daysrandom: vDSO: avoid call to memset() when zeroing reserved parameterNathan Chancellor
After a recent change in LLVM [1], builds with the random vDSO implementation, such as PowerPC and RISC-V, fail when checking the vDSO: arch/powerpc/kernel/vdso/vdso32.so.dbg: dynamic relocations are not supported arch/riscv/kernel/vdso/vdso.so.dbg: dynamic relocations are not supported memset() is now generated when zeroing params->reserved for some builds because LLVM has an optimization (now run in more instances) that can recognize at compile time when it is assigning a static value to a contiguous area of memory and turn that into a call to memset(). Both clang and GCC assume memset() is always available [2]. Clang has an internal fiddly hook, -max-store-memset, which we can set to a high number, to disable generating out of line memset calls [3]. Similarly, GCC has -finline-stringops=memset to do the same [4], should this issue ever hit future version of GCC. While these options wouldn't make sense for normal kernel code, it is fine for the extremely limited and intentionally compact vDSO code. Link: https://github.com/llvm/llvm-project/commit/90cebef1411617fc3eedd359bdf00cb44b1c2439 [1] Link: https://gcc.gnu.org/onlinedocs/gcc-16.2.0/gcc/Standards.html#index-ffreestanding [2] Link: https://github.com/llvm/llvm-project/commit/b28eeb28bea39148738dc375e8a97072a1907e64 [3] Link: https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#index-finline-stringops [4] Closes: https://github.com/ClangBuiltLinux/linux/issues/2183 Cc: stable@vger.kernel.org # v6.12+ Signed-off-by: Nathan Chancellor <nathan@kernel.org> Signed-off-by: Jason A. Donenfeld <Jason@zx2c4.com>
8 daysMerge tag 'libcrypto-fixes-for-linus' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/ebiggers/linux Pull crypto library fixes from Eric Biggers: - Fix a performance regression in certain AES encryption modes on certain architectures, introduced this cycle - Fix a small performance regression in the x86_64 optimized AES-GCM code, introduced in 6.15 * tag 'libcrypto-fixes-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/ebiggers/linux: crypto: aes - Fix undesired override of some optimized AES modes crypto: x86/aes-gcm - fix always true check for last AAD segment
8 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "The most intrusive change is reverting a commit from 7.3-rc1 that made struct kvm a bit too large, and fixing the same issue otherwise. There are again a lot of selftests lines; the sheer number of commits is not small but I don't expect much more for 7.3 due to people travelling to Plumbers next week. ARM: - Take a reference on the last IRQ loaded into an LR to prevent it from being freed while running the guest (Marc Zyngier) - Ensure that the ITS MOVALL command only affects LPIs that were previously affined to the source redistributor (Marc Zyngier) - Fix + test for honoring the host's trap configuration when running non-protected VMs while KVM is in protected mode (Fuad Tabba) - Use the host stage-1 mapping granularity for VM_PFNMAP mappings at stage-2 (Mostafa Saleh) x86: Various bugfixes where the guest could do stupid things on purpose to cause problems in the host: - Failed VMRUNs can cause pending TLB flushes to be dropped, and in general some actions done through VMCB control fields have to be redone if VMRUN fails - Toggling MSR interceptions or eVMCS execution controls can cause the host to use a stale MSR permission bitmap - Bad page tables can cause a WARN. Also fix issues in last week's pull request (my fault, for changing email workflow and thus missing feedback sent to kvm@ but not LKML). Generic: - Take kvm_lock when creating vCPUs. For almost two decades everybody thought it was not done for some unspecified performance reasons, but in reality it was only done because kvm_lock was originally a spinlock. This is a better fix than 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as possible", from the 7.3 merge window), and does not waste 2K per VM, hence its inclusion here" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (29 commits) KVM: arm64: Use stage-1 leaf size for VM_PFNMAP KVM: arm64: selftests: Check a feature hidden in an ID register is UNDEF KVM: arm64: Use the host's HCR_EL2 for non-protected VMs in pKVM KVM: arm64: Clear HCR_EL2.RW for 32-bit non-protected vCPUs KVM: arm64: Apply the fine-grained UNDEFs without FEAT_FGT KVM: arm64: vgic-its: Fix MOVALL handling of source redistributor KVM: arm64: vgic: Take a refcount on IRQs referenced by last_lr_irq KVM: arm64: vgic: Allow last_lr_irq to be NULL when LRs are not overflowing KVM: SEV: Do cache maintenance on the source VM *before* clearing SEV state KVM: SEV: Nullify "have run CPUs" mask pointer when freeing it KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt KVM: selftests: Verify that L0's TPR doesn't get clobbered KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded ...
10 dayscrypto: x86/aes-gcm - fix always true check for last AAD segmentMohamad Raizudeen
In gcm_process_assoc(), a segment that is not the last one must have its length rounded down to a multiple of 16 bytes, as required by the assembly. The check for a non-last segment is `if (unlikely(assoclen)) /* Not the last segment yet? */` where assoclen is the number of AAD bytes remaining after the current segment. Since the conversion to the new scatterwalk API, assoclen is decremented at the end of the loop body rather than the beginning, so it still includes the current segment when the check executes, making the check always true. As a result the last segment was rounded down as well, causing some avoidable extra work: an additional memcpy into the temporary buffer and an additional call into the assembly after the loop. The GCM authentication tag is unaffected either way, so the self-tests pass and this went unnoticed. It is purely an efficiency issue rather than a correctness one. Fix this by moving the assoclen decrement back to the beginning of the loop body. Fixes: e9787deff49e ("crypto: x86/aes-gcm - use the new scatterwalk functions") Cc: stable@vger.kernel.org Suggested-by: Eric Biggers <ebiggers@kernel.org> Signed-off-by: Mohamad Raizudeen <raizudeen.kerneldev@gmail.com> Link: https://patch.msgid.link/20260926145903.6061-1-raizudeen.kerneldev@gmail.com Signed-off-by: Eric Biggers <ebiggers@kernel.org>
11 daysbpf: Skip detached progs in trampoline images that are still in useFlorent Revest (Anthropic)
bpf_tramp_image_put() makes sure a trampoline image is not freed while a task may still be running in it, but nothing similar is done for the progs called by that image. Detach drops the last prog reference right away and the prog is freed after grace periods, on the basis that a task still in the traced function skips the fexit progs once the nop at ip_after_call is patched to a jump. That leaves out a task sleeping in a sleepable prog that runs before the detached one, which no grace period waits for: CPU 0 CPU 1 in image I, sleeping in prog S detach P from I's trampoline -> new image, bpf_tramp_image_put(I) bpf_prog_put(P), last ref grace periods, P freed back from S __bpf_prog_enter(P) call P->bpf_func If S and P are fexit progs the task is already past the patched jump, and fentry only images don't have one. On x86 this is an int3 in poisoned bpf_prog_pack memory: Oops: int3: 0000 [#1] SMP NOPTI CPU: 18 UID: 0 PID: 94573 Comm: x169 Not tainted 6.18.44 #1 PREEMPT(lazy) RIP: 0010:0xffffffffc0601d8d Call Trace: <TASK> ? bpf_trampoline_6442515411+0x1a4/0x21b bpf_lsm_bprm_committed_creds+0x5/0x10 security_bprm_committed_creds+0x5f/0x70 begin_new_exec+0x2d6/0x410 ... We hit this in production when progs attached through trampolines got detached while their hooks were busy, and it was independently found with a fuzzer and KASAN. Have the JITs emit a patchable nop in front of each prog call sequence and record it in the image. When a prog is detached, patch its nop to a jump over the call sequence. Progs that stay attached keep running for the tasks that are in the image, and ip_after_call isn't needed anymore. The task can be in any image that isn't freed yet, not only in the current one. It sleeps in image I1 that calls S, P and Q, then P is detached and the trampoline moves to image I2, then Q is detached and its call is still in I1. So the trampoline keeps a list of its images until they are freed, and detaching a prog patches its nop in all of them. Images hold a reference on the trampoline for that long. On riscv and loongarch a jump of any range takes several instructions, and a task preempted in the middle of them could resume into half of the new sequence. The nop is a single instruction there, patched to a near branch through arch_bpf_trampoline_skip(). With the extra nops, BPF_MAX_TRAMP_LINKS progs no longer fit in a page on x86 and arm64 (and already didn't on powerpc), so lower the limit there like s390 does. Fixes: e21aa341785c ("bpf: Fix fexit trampoline.") Reported-by: Sechang Lim <rhkrqnwk98@gmail.com> Suggested-by: Alexei Starovoitov <ast@kernel.org> Signed-off-by: Florent Revest (Anthropic) <florent.revest@linux.dev> Signed-off-by: Alexei Starovoitov <ast@kernel.org> Link: https://patch.msgid.link/20260926135605.1217928-3-florent.revest@linux.dev Closes: https://lore.kernel.org/bpf/20260815071927.147049-1-zirajs7@gmail.com/
11 daysKVM: SEV: Do cache maintenance on the source VM *before* clearing SEV stateSean Christopherson
Explicitly flush caches after intra-host migration before clearing "SEV active" on the source VM, as doing cache maintenance afterwards creates a tiny window where memory reclaim could return memory to the host without performing a cache flush, e.g. as pointed out by Sashiko: CPU1 in sev_migrate_from(): src->active = false; CPU2 running concurrent unmap: Since active is false, the automatic cache flush in sev_guest_memory_reclaimed is skipped. The host frees and reallocates the page. CPU1 in sev_migrate_from(): sev_writeback_caches(src_kvm); Executes a hardware cache flush (wbnoinvd), which writes the guest's old dirty ciphertext over the new page owner's data. Fixes: 93de2a6a4b91 ("KVM: SEV: Do cache maintenance on the source VM during intra-host migration") Cc: stable@vger.kernel.org Reported-by: Sashiko Bot <sashiko-bot@kernel.org> Closes: https://lore.kernel.org/all/20260923165304.1662E1F000FF@smtp.kernel.org Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260928154644.2559454-3-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
11 daysKVM: SEV: Nullify "have run CPUs" mask pointer when freeing itSean Christopherson
Nullify have_run_cpus when freeing the mask, particularly in the error path of __sev_guest_init(), so that KVM doesn't have to subtly use sev->active to track whether or not the mask has been freed. As pointed out by Sashiko, blindly freeing the mask in sev_vm_destroy() results in a double-free if the mask is freed if __sev_guest_init() fails. Throw the logic in a helper as nullifying the pointer is frustratingly difficult and weird due to have_run_cpus being a single-entry array when CPUMASK_OFFSTACK=n. Deliberately don't use CPUMASK_VAR_NULL, as it's not directly assignable when the cpumask is on-stack, e.g. requires using a local variable and a memcpy(), which is beyond ridiculous. Furthermore, while clearing the on-stack bitmask is an unnecessary and arguably unwanted side effect, KVM absolutely relies on '0' being the "null" value given that the struct is zero-allocated. Opportunistically add an alloc() helper to pair with free(); there are just enough call sites to make doing so worthwhile. Fixes: 12c1f6e03f94 ("KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV") Cc: stable@vger.kernel.org Reported-by: Sashiko Bot <sashiko-bot@kernel.org> Closes: https://lore.kernel.org/all/20260923165349.CAAF01F000FF@smtp.kernel.org Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260928154644.2559454-2-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
11 days{x86,um}/uapi/ptrace: Guard register offset macros with __ASSEMBLER__ or ↵Nick Desaulniers
__FRAME_OFFSETS The register offset macros in <asm/ptrace-abi.h> are guarded by `defined(__ASSEMBLER__) || defined(__FRAME_OFFSETS)` for 64-bit, but were left unguarded for 32-bit. This causes havoc for userspace that happens to use identifiers colliding with these short macro names (e.g., EBX, ECX, EAX, DS, ES, FS, GS, CS, SS). Without this guard, userspace is forced to be super extra careful with include ordering to minimize the chance of collision. Wrap both the 32-bit and 64-bit register definitions under `#if defined(__ASSEMBLER__) || defined(__FRAME_OFFSETS)`, and ensure User-Mode Linux (UML) defines `__FRAME_OFFSETS` for 32-bit as well. Closes: https://github.com/llvm/llvm-project/issues/217413 Assisted-by: LLM Signed-off-by: Nick Desaulniers <ndesaulniers@google.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Acked-by: Oleg Nesterov <oleg@redhat.com> Acked-by: Johannes Berg <johannes@sipsolutions.net> Tested-by: Elliott Hughes <enh@google.com> Link: https://patch.msgid.link/20260821-ptrace_uapi-v1-1-3de8638a29f2@google.com
11 daysx86/mm: Drop unnecessary PMD page copy when freeingMikhail Gavrilov
On a box with a discrete GPU, lockdep reports a possible deadlock as soon as kswapd shrinks the TTM page pool. The immediate cause is an x86 commit that added an mmap_read_lock() to kernel page protection munging code. The huge vmap code holds the same lock over a GFP_KERNEL allocation, which is a no-no now that reclaim can take it. That allocation is in a page table *free* path and ends up being for dubious purposes[1]. Basically, it tries to avoid hardware setting Accessed=1 in page table entries that are unreachable by the hardware, a non-issue. Remove the PMD copy. Detach the original PMD page at the PUD, flush the mid-level caches, and free the PTE tables straight from the detached PMD page. With no allocation left, the locking issue is gone. Lockdep splat/analysis: WARNING: possible circular locking dependency detected 7.3.0-rc3-f6e7b42bf05b+ #183 Tainted: G U ------------------------------------------------------ kswapd0/269 is trying to acquire lock: ((init_mm).mmap_lock){++++}-{4:4}, at: change_page_attr_set_clr+0x29a/0x4a0 but task is already holding lock: (pool_shrink_rwsem){.+.+}-{4:4}, at: ttm_pool_shrink+0xb2/0x330 [ttm] Chain exists of: (init_mm).mmap_lock --> fs_reclaim --> pool_shrink_rwsem The cycle is built from three edges: 1) pool_shrink_rwsem -> (init_mm).mmap_lock The TTM shrinker restores the caching attribute of every page it frees, while holding pool_shrink_rwsem: ttm_pool_shrink() -> ttm_pool_dispose_list() -> ttm_pool_free_page() -> set_pages_wb() -> change_page_attr_set_clr() [ init_mm mmap read lock ] 2) fs_reclaim -> pool_shrink_rwsem The same shrinker, called from reclaim. 3) (init_mm).mmap_lock -> fs_reclaim ioremap() installing a huge PUD mapping over an existing PMD table: ioremap_page_range() -> vmap_range_noflush() -> vmap_try_huge_pud() [ init_mm mmap read lock ] -> pud_free_pmd_page() -> __get_free_page(GFP_KERNEL) [ enters reclaim ] [ dhansen: Lots of changelog munging/trimming and merged comments from my version of the fix. ] Fixes: d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF") Suggested-by: Pedro Falcato <pfalcato@suse.de> Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com> Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com> Reviewed-by: Pedro Falcato <pfalcato@suse.de> Link: https://lore.kernel.org/20260916062222.27347-1-mikhail.v.gavrilov@gmail.com Link: https://lore.kernel.org/all/e11449f0-d9ad-4d1b-ab21-2be7d71fe335@intel.com/ [1] Link: https://patch.msgid.link/20260923223116.20090-1-mikhail.v.gavrilov@gmail.com Cc: stable@vger.kernel.org
12 daysx86/mm: Don't apply va_align to hugetlb mappings on AMD F15hLaurent Wandrebeck
get_align_mask() returns huge_page_mask_align() for hugetlbfs, but get_align_bits() adds va_align.bits regardless, so vm_unmapped_area() returns an address off the huge page boundary and __unmap_hugepage_range() hits BUG_ON(start & ~huge_page_mask(h)) at teardown. This can be triggered on Carrizo and FX-8370E, both hstates. Pass the file to get_align_bits() and skip the randomisation for hugetlbfs. [ bp: Massage commit message. ] Fixes: 1317a5e7f7b1 ("arch/x86: teach arch_get_unmapped_area_vmflags to handle hugetlb mappings") Suggested-by: Dave Hansen <dave.hansen@intel.com> Acked-by: Dave Hansen <dave.hansen@intel.com> Signed-off-by: Laurent Wandrebeck <l.wandrebeck@quelquesmots.fr> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Cc: stable@vger.kernel.org # 6.13+ Link: https://patch.msgid.link/20260922085032.46144-1-l.wandrebeck@quelquesmots.fr
12 daysMerge tag 'x86-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Ingo Molnar: - Fix preemption bugs in the SVSM vTPM guest implementation (Melody Wang) - Fix MCE-triggered hardware debug register corruption on task migration (Masami Hiramatsu) * tag 'x86-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/mce: Fix hardware debug register corruption on task migration x86/sev: Make vTPM SVSM calls preemption-safe
12 daysMerge tag 'perf-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fixes for KVM guest PEBS virtualization (Sean Christopherson) - Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi) - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi) - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi) - Rename two confusingly named PMU attributes (Dapeng Mi) - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim) - Fix NULL pointer dereference crash in __perf_pmu_sched_task() (Puranjay Mohan) - Fix CPU-wide event scheduling (Puranjay Mohan) - Fix x86 LBR branch entry generation (Puranjay Mohan) * tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf/core: Fill branch entries with a single assignment perf/core: Run sched_task() for PMUs with only CPU-wide events perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() perf/core: Fix a refcount leak in attach_perf_ctx_data() perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3 perf/x86/intel: Delete dead NVL PEBS data-source initcall perf/x86/intel: Fix Panther Cove PEBS data-source snoop states perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints perf/x86/intel: Update arw_latency_data() mem-op direction handling perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
13 daysMerge tag 'pci-v7.3-fixes-2' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci Pull PCI fixes from Bjorn Helgaas: - Make BAR resize work even for devices where no upstream bridge is visible to the OS, which fixes an amdgpu regression on SolidRun HoneyComb, which doesn't expose Root Ports to the OS (Liz Fong-Jones) - Omit bus properties in dynamic OF nodes when a bridge has no subordinate bus, which fixes early boot hangs caused by NULL pointer dereferences with CONFIG_PCI_DYNAMIC_OF_NODES enabled (Angel J) - Disable enhanced atomics on AMD NBIO 7.7 and 7.11 to avoid silent data corruption on 64-bit DMAs (Mario Limonciello) * tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci: x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11 PCI: of_property: Omit bus properties without a subordinate bus PCI: Fix BAR resize for devices on a root bus
13 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "Arm: - Invalidate the ITS translation cache when the guest changes the base address of the ITS tables (Fuad Tabba) - Skip saving ITS devices with device IDs that are out-of-bounds rather than failing the entire ITS save ioctl (Fuad Tabba) - Close race between VM teardown and invalidations of nested MMUs when handling MMU operations that are allowed to block (Lorenzo Stoakes) - Various fixes for the handling of the host's untrusted SVE configuration in pKVM (Fuad Tabba) - Make sure that empty SMCCC ranges based at 0 are rejected by the kvm_smccc_set_filter() (Karl Mehltretter) - Revoke the host mapping for pKVM's private stack pages, along with a new sanity check that all mappings in the hyp's private VA range have been correctly marked as hyp-owned (Fuad Tabba) - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that concurrent vCPU initialization cannot relocate in-use MMUs. Defer the freeing of shadow stage-2 MMUs to the point that no other users (e.g. MMU notifier) could reference them (Marc Zyngier) - Drop useless WARN when rejecting an unsupported ioctl for pKVM (Fuad Tabba) - Fix the steal_time selftest to install correctly-sized mappings for non-4K hosts (Sebastian Ott) - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark Brown) - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit guests (Karl Mehltretter) RISC-V: - Synchronize hrtimer during VCPU teardown - Fix the conversion between vsip and hvip values - Serialize IMSIC attributes with vCPU migration - Release unused page after MMU invalidation - Propagate interrupted G-stage faults to KVM user-space as EINTR - Fix nested acceleration hfence entry update order - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem - Preserve firmware counter value across PMU counter stop/start - Report PMU snapshot write failure to the guest - Fix perf-backed counter accounting across PMU stop and read - Correctly propagate error of a hart status SBI call s390: - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty the pages that contain indicator and summary bits - Fix compile warning for kvm_s390_update_cmma_dirty() - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to userspace - Move s390_kvm_mmu_commit_memory_region() into s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN - Add missing srcu in kvm_s390_set_irq_state() - Fix potential races in storage functions - Fix race in _destroy_pages_crste() - Fix issues in the handling of KVM interrupt and page resources, when a queue that is assigned to a mediated device (mdev) is removed from the host's AP configuration - Fix loop condition in uv_find_secrets - Prevent potential out-of-bounds read x86: - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl) - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely) - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior) - Fix a memory leak and a cache maintenance issue related to doing intra-host migration on an SEV guest" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits) KVM: SEV: Do cache maintenance on the source VM during intra-host migration KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg KVM: Don't treat reserved xarray entries as having memory attributes KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails KVM: arm64: Fix AArch32 DBGBXVR<n> handling KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP KVM: selftests: fix steal_time for arm64 with host page size > 4K KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction KVM: arm64: nv: Fix life cycle of the nested_mmus array KVM: arm64: Check every private mapping is hyp-owned at pKVM init KVM: arm64: Move the private VA allocation cursor to __io_map_next KVM: arm64: Match hyp text by physical address in fix_host_ownership() KVM: arm64: Transfer the hyp stack pages out of the host stage-2 KVM: arm64: selftests: Test empty SMCCC filter range at base 0 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 ...
14 daysKVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modifiedSean Christopherson
Force a refresh of the vmcs02 MSR bitmap during nested VM-Enter if the runtime eVMCS controls (pin, primary, secondary, etc.) are being updated. If L1 isn't intercepting TPR writes, runs L2 with TPR virtualization, and then runs the same L2 with TPR virtualization disabled, KVM will fail to refresh msr_bitmap02 and leave TPR in passthrough mode even though TPR virtualization is disabled. I.e. failure to refresh the bitmap lets L2 (or L1 by proxy) read and write L0's TPR. Fixes: 502d2bf5f2fd ("KVM: nVMX: Implement Enlightened MSR Bitmap feature") Cc: stable@vger.kernel.org Reviewed-by: Vitaly Kuznetsov <vkuznets@redhat.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is interceptedSean Christopherson
Use the MSR permission bitmap of the active VMCB instead of assuming that KVM is always using vmcb02's bitmap when L2 is active, as KVM uses msrpm02 if and only if L1 wants to intercept MSR accesses, i.e. if and only if KVM needs to merge msprm01 with msrpm12. Don't bother tracking the virtual address of the bitmap that's being used, as __va() is cheap on x86, and caching the virtual address would introduce yet another source of potentially stale information. Fixes: b2ac58f90540 ("KVM/SVM: Allow direct access to MSR_IA32_SPEC_CTRL") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu <fane@google.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260826195833.844526-1-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUNSean Christopherson
Don't (re)read PERF_CNTR_GLOBAL_CTL from hardware on a failed VMRUN, as the purpose of the read is to synchronize KVM's cache with any writes done by the guest, and the guest can't possibly have modified the MSR if it never got a chance to run. Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260904170642.3291466-5-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctlSean Christopherson
Don't mark the ASID as dirty in the VMCB when requesting a TLB flush via control.tlb_ctl. Per "15.15.3 VMCB Clean Field" of the July 2026, Revision 3.45 version of the APM: The following are explicitly not cached and not represented by Clean bits: * TLB_Control Fixes: 7e8e6eed75e2 ("KVM: SVM: Move asid to vcpu_svm") Suggested-by: Yosry Ahmed <yosry@kernel.org> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260904170642.3291466-4-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeededSean Christopherson
Leave control.erap_ctl and control.clean as-is in the VMCS if VMRUN fails, because as per AMD: there's no explicit architectural guarantee about the behavior in the presence of VMRUN failures. So the best thing to do would be to assume that if VMRUN fails, the actions requested in the control fields may not have been performed. Cc: stable@vger.kernel.org Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260904170642.3291466-3-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUNSean Christopherson
Don't reset the VMCB's TLB control back to "do nothing" on a failed VMRUN, as empirical testing shows that the CPU performs the requested TLB flush if and only if VMRUN is successful, i.e. clearing TLB control on a failed VMRUN effectively drops a TLB flush. Explicitly track the need to flush all ASIDs on a per-CPU basis, as the ASID reuse condition is tied to the pCPU, not to the vCPU. As a bonus, this also obviates the need to avoid clobbering FLUSH_ALL_ASID with TLB_CONTROL_FLUSH_ASID, e.g. in svm_flush_tlb_asid(). Deliberately don't bother saving/restoring the "old" tlb_ctl on failure, in quotes because it's not exactly the old tlb_ctl, it's the tlb_ctl from after pre_svm_run(), but before updating tlb_ctl for flush_all_asids. If VMRUN fails and TLB_CONTROL_FLUSH_ALL_ASID is forced, then the next successful run of the VMCB *may* unnecessarily flush all ASIDs, which strictly speaking could result in noisy neighbor issues. However, the fact that new_asid() is already guest-triggerable, because of KVM's flawed behavior of clearing the ASID on emulated INIT, means that a guest can already trigger a flush of all ASIDs at roughly the same rate. And once KVM stops clobbering the ASID on emulated INIT, *or* assigns a static ASID to each vCPU, this flaw goes away. Fixes: 38e5e92fe8c0 ("KVM: SVM: Implement Flush-By-Asid feature") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu <fane@google.com> Suggested-by: Yosry Ahmed <yosry@kernel.org> Cc: Tom Lendacky <thomas.lendacky@amd.com> Cc: Jim Mattson <jmattson@google.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260904170642.3291466-2-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: x86/mmu: Bail from shadow walks if the root is invalid or a dummySean Christopherson
When walking shadow page tables, immediately terminate the walk if the root is a "dummy" root, i.e. a root whose top-level page table is backed by the zero page, but otherwise doesn't exist. If memslot creation races with a stage-2 page fault (EPT violation or #NPF) from L2, then if the stars align, KVM will attempt to walk shadow page tables using the zero page and hit a NULL pointer deref. BUG: kernel NULL pointer dereference, address: 0000000000000021 #PF: supervisor read access in kernel mode #PF: error_code(0x0000) - not-present page PGD 0 P4D 0 Oops: Oops: 0000 [#1] SMP CPU: 30 UID: 1000 PID: 941 Comm: qemu Not tainted 7.2.0-rc2-1b731e5ded48-next-vm #1741 PREEMPTLAZY Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 0.0.0 02/06/2015 RIP: 0010:__kvm_mmu_invalidate_addr+0xea/0x210 [kvm] Call Trace: <TASK> kvm_mmu_invalidate_addr+0x92/0xe0 [kvm] __kvm_inject_emulated_page_fault+0x67/0x80 [kvm] ept_page_fault+0x160/0x850 [kvm] kvm_mmu_do_page_fault+0x102/0x1f0 [kvm] kvm_mmu_page_fault+0x8e/0x6b0 [kvm] vmx_handle_exit+0x163/0x640 [kvm_intel] kvm_arch_vcpu_ioctl_run+0x960/0x2120 [kvm] kvm_vcpu_ioctl+0x2c7/0x970 [kvm] __x64_sys_ioctl+0x90/0xd0 do_syscall_64+0x67/0x5f0 entry_SYSCALL_64_after_hwframe+0x4b/0x53 RIP: 0033:0x7f606d0b53bb Opportunistically harden the shadow walks against fully invalid roots, but WARN, as all callers are expected/required to pre-check for a valid root. Don't WARN in the dummy root case as the whole point of using a dummy root is to provide a root that's valid enough to enter the guest, i.e. it should Just Work for all flows except those that *need* to know about dummy roots. Alternatively, KVM could allocate a dedicated page and associated shadow page structure for the dummy root, which is very tempting as it such an approach should be more resilient against unexpected behavior. But that would be a much larger and thus riskier change than simply terminating walks of dummy roots. Fixes: 0e3223d8d00a ("KVM: x86/mmu: Use dummy root, backed by zero page, for !visible guest roots") Reported-by: Gabriel Schneider <gbrls@osec.io> Cc: stable@vger.kernel.org Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260925234256.2384816-1-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: Reject attempts to lock all vCPUs if vCPU creation is in-progressSean Christopherson
Reject locking of all vCPUs if vCPU creation is in-progress, i.e. if the number of "created" vCPUs doesn't match the number of "onlined" vCPUs. It's simply not possible to guarantee that KVM has truly locked all vCPUs if one or more vCPUs are actively being created. Holding kvm->lock does prevent in-flight vCPUs from being fully onlined, but it's infeasible for common KVM to know whether or not that provides sufficient protection. In practice, this is likely a minor bug fix for the ARM and RISC-V usage of kvm_trylock_all_vcpus(), and a glorified nop for everything else. E.g. ARM's kvm_timer_vcpu_init() can race kvm_vm_ioctl_set_counter_offset() with respect to observing KVM_ARCH_FLAG_VM_COUNTER_OFFSET. Opportunistically drop x86's existing manual checks on vCPU creation being in-progress as all of x86's checks immediately precede or follow locking of all vCPUs. Leave arm64 and RISC-V alone for the moment, as their checks aren't as obviously redundant/equivalent. Signed-off-by: Sean Christopherson <seanjc@google.com> Tested-by: Jean-Christophe Guillain <jean-christophe@guillain.net> Tested-by: Naveen N Rao (AMD) <naveen@kernel.org> Message-ID: <20260921174445.911676-2-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysMerge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEADPaolo Bonzini
KVM fixes for 7.3-rcN - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD. - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl). - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs. - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages. - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed. - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes. - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely). - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior).
14 daysKVM: SEV: Do cache maintenance on the source VM during intra-host migrationSean Christopherson
Manually perform cache maintenance on the source VM during intra-host migration to ensure no stale data is left in CPU caches after the VM is destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's memory reclaim flows won't trigger cache maintenance, e.g. when all guest memory is reclaimed in response to detaching from the mmu_notifier. Note, relying on the destination VM to do cache maintenance isn't an option as KVM doesn't require identical guest memory configurations, i.e. the source VM may have access to memory that the destination VM does not. Enforcing equivalent memory configurations is infeasible, as it would require a *deep* comparison of memslots, e.g. to verify that not only are the memslot identical, but what the memslots point at is also identical. Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu <fane@google.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260923163721.1584779-3-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
14 daysKVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEVSean Christopherson
Unconditionally free SEV's "have run CPUs" cpumask in the VM destroy path, i.e. even for what appear to be non-SEV VMs, as an SEV VM becomes a non-SEV VM if its state is intra-host migrated. Alternatively, the mask could be freed in sev_migrate_from() when "converting" the source VM, but that gets annoying because ideally KVM would nullify the mask to guard against UAF, and nullifying the mask would need be conditioned on CPUMASK_OFFSTACK=y. Freeing the mask during sev_migrate_from() is also not robust against other KVM bugs, though that's kind of a moot point since any such bugs would show up even if sev->active is never set. I.e. KVM must get that side of things correct. But, that's not a great reason to add more code just to make things marginally less robust. Fixes: 6f38f8c57464 ("KVM: SVM: Flush cache only on CPUs running SEV guest") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu <fane@google.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260923163721.1584779-2-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-09-23x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11Mario Limonciello
Multiple users report data corruption during 64-bit DMA transfers on systems with AMD NBIO 7.7 and 7.11 controllers. This occurs when BIOS enables AMD "enhanced atomic operations" on PCIe Root Ports. When enhanced atomics are enabled, any 64-bit DMA access may be corrupted. Disable enhanced atomics using SMN for NBIO 7.7 and 7.11 based models. Reported-by: Mikael Etienne <mikael1022bzh@gmail.com> Closes: https://lore.kernel.org/178789300872.392066.15963676631650361573@gmail.com/ Reported-by: Arthur Husband <artmoty@gmail.com> Closes: https://lore.kernel.org/20260406222335.379935-1-artmoty@gmail.com/ Reported-by: Alvin Lim <alvinwylim@gmail.com> Closes: https://lore.kernel.org/20260621100844.1224301-1-alvinwylim@gmail.com/ Signed-off-by: Mario Limonciello <mario.limonciello@amd.com> [bhelgaas: commit log, s/IOVA/DMA/ in comment] Signed-off-by: Bjorn Helgaas <bhelgaas@google.com> Cc: stable@vger.kernel.org Cc: David Laight <david.laight.linux@gmail.com> Cc: John Smith <imjohnsmith4000@gmail.com> Cc: Lennert Buytenhek <kernel@wantstofly.org> Cc: Niklas Cassel <cassel@kernel.org> Cc: Roland Waltersson <roland.waltersson@netinsight.net> Link: https://patch.msgid.link/20260908190600.226485-2-mario.limonciello@amd.com
2026-09-23x86/mce: Fix hardware debug register corruption on task migrationMasami Hiramatsu (Google)
In exc_machine_check_user(), local_db_save() and local_db_restore() are invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER, DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding exc_machine_check_user(). However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If the task migrates to another CPU during schedule(), local_db_restore() runs on the new CPU with the dr7 state saved from the old CPU. This corrupts the new CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In short, local_db_save() and local_db_restore() pair must be run on the same CPU. To fix this, move local_db_save() and local_db_restore() inside exc_machine_check_user() and exc_machine_check_kernel(). In exc_machine_check_user(), DR7 is saved and restored strictly around do_machine_check() to avoid schedule() during migration. In exc_machine_check_kernel(), local_db_save() is called at the entry point to prevent early memory accesses from triggering nested #DB exceptions, and restored on all exits. Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC") Assisted-by: LLM Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org> Cc: <stable@kernel.org> Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
2026-09-23perf/core: Fill branch entries with a single assignmentPuranjay Mohan
perf_clear_branch_entry_bitfields() clears the bitfields of struct perf_branch_entry one by one and leaves from/to alone, since callers overwrite those straight away. The list has to be kept in sync with the struct by hand and has already fallen behind: new_type and priv were added to perf_branch_entry and never added here. Only BRBE writes those two, and neither for every record. brbe_set_perf_entry_type() leaves new_type alone for a branch type it does not recognise, and priv is not set for source-only records. arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a record reaches userspace with whatever the slot held: uninitialised kmalloc() data on the first pass over the buffer, the previous record's values after that. Nothing under arch/x86/events/ writes either field, so only arm64 is affected. Assign the whole entry at each site instead. Everything not named is then zero, and there is no list to keep in sync. The bitfields add up to exactly 64 bits, so the struct has no padding to leave undefined. perf_clear_branch_entry_bitfields() has no callers left, so remove it. perf_entry_from_brbe_regset() assigns an empty literal instead, since it fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the explicit spec assignment changes nothing. Fixes: b190bc4ac9e6 ("perf: Extend branch type classification") Fixes: 5402d25aa571 ("perf: Capture branch privilege information") Suggested-by: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: Yifan Wu <wuyifan50@huawei.com> Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
2026-09-22KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error ↵Sean Christopherson
paths Fill kvm_run with "internal error, emulation" in the common error handling paths for getting nested state pages, as requiring each check to manually fill kvm_run is error prone and requires a non-trivial amount of copy+paste. Specifically, both SVM and VMX fail to fill kvm_run if load_pdptrs() fails, and SVM fails to fill kvm_run if kvm_hv_verify_vp_assist() fails. If those flows fail, the *best* case scenario is that KVM will exit to userspace with KVM_EXIT_UNKNOWN. The worst case scenario is that KVM exits with a stale exit_reason and confuses userspace. Note, SVM never exits to userspace if something goes sideways when dealing with vmcb12 assets while emulating VMRUN, i.e. lack of SVM-specific code is not a bug. Fixes: 0f85722341b0 ("KVM: nVMX: delay loading of PDPTRs to KVM_REQ_GET_NESTED_STATE_PAGES") Fixes: 232f75d3b4b5 ("KVM: nSVM: call nested_svm_load_cr3 on nested state load") Fixes: 3f4a812edf5c ("KVM: nSVM: hyper-v: Enable L2 TLB flush") Cc: stable@vger.kernel.org Reported-by: Jinwoo Lee <rkskek9254@gmail.com> Closes: https://lore.kernel.org/all/20260813043932.3214460-1-rkskek9254@gmail.com Reported-by: Stefan Teodorescu <fane@google.com> Link: https://patch.msgid.link/20260921211608.1030158-3-seanjc@google.com Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-09-22KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages failsSean Christopherson
Re-pend GET_NESTED_STATE_PAGES before exiting to userspace if getting the nested pages fails in the KVM_RUN path. If userspace re-runs the vCPU, and vmcs02 holds valid PFNs from the *previous* run of L2, then KVM could re-enter L2 with stale, unpinned PFNs mapped into e.g. the vAPIC page. Note, both SVM and VMX (as of commit 11722439fb20 ("KVM: nVMX: Ensure KVM_REQ_GET_NESTED_STATE_PAGES is cleared on VM-Exit") ensure the request is cleared on VM-Exit (including the "forced" case), i.e. there is no risk of double-mapping due to emulated VMLAUNCH/VMRESUME/VMRUN *and* the request trying to map the nested pages. Fixes: 671ddc700fd0 ("KVM: nVMX: Don't leak L1 MMIO regions to L2") Cc: stable@vger.kernel.org Reported-by: Jinwoo Lee <rkskek9254@gmail.com> Closes: https://lore.kernel.org/all/20260813043932.3214460-1-rkskek9254@gmail.com Reported-by: Stefan Teodorescu <fane@google.com> Link: https://patch.msgid.link/20260921211608.1030158-2-seanjc@google.com Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-09-22perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rspDapeng Mi
NVL introduces Offmodule Response events in place of the legacy Offcore Response events, but it still exposes the inherited offcore_rsp PMU attribute for programming the corresponding MSR data. Rename the NVL PMU attribute to offmodule_rsp so the sysfs interface matches the underlying event name and avoids user & tooling confusion. Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-13-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rspDapeng Mi
DMR introduces Offmodule Response events in place of the legacy Offcore Response events, but it still exposes the inherited offcore_rsp PMU attribute for programming the corresponding MSR data. Rename the DMR PMU attribute to offmodule_rsp so the sysfs interface matches the underlying event name and avoids user & tooling confusion. Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-12-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Fix precise OMR event scheduling for DMR/NVLDapeng Mi
The latest perfmon event database introduces below precise OMR event support for DMR/NVL: - MEM_LOAD_L2_MISS_RETIRED.* (event 0xd6) - MEM_STORE_L2_MISS_RETIRED.* (event 0x4f) These events use the same OMR MSRs as the existing OMR events, but they are not listed in intel_pnc_extra_regs[]. As a result, perf cannot assign the required OMR extra registers when scheduling them. Add the new precise OMR events to intel_pnc_extra_regs[] so they can be scheduled with the correct OMR MSRs. MEM_LOAD_L2_MISS_RETIRED.* remains limited to GP counters 0-3, while MEM_STORE_L2_MISS_RETIRED.* is available on all GP counters. Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-11-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3Dapeng Mi
Per the latest Panther Cove event definitions, the following events are only supported on PMCs 0-3: - UOPS_DISPATCHED.INT_EU_ALL (0x1b2) - UOPS_DISPATCHED.ALU (0x2b2) Add explicit event constraints for these two events so scheduling does not place them on unsupported counters. Fixes: d345b6bb8860 ("perf/x86/intel: Add core PMU support for DMR") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.0+ Link: https://patch.msgid.link/20260917015234.981153-10-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Delete dead NVL PEBS data-source initcallDapeng Mi
Nova Lake now uses the OMR data-source table for PEBS data-source decoding and no longer depends on the legacy static pebs_data_source[] mapping. Remove the dead intel_pmu_pebs_data_source_lnl() initialization call for NVL. Fixes: c847a208f43b ("perf/x86/intel: Add core PMU support for Novalake") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-9-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Fix Panther Cove PEBS data-source snoop statesDapeng Mi
For Panther Cove, the snoop states for the data source encodings "Prefetch Promotion" and "Cross Core Prefetch Promotion" should be SNOOP_NONE instead of SNOOP_MISS. Fix the incorrect snooping states for Panther Cove. Fixes: d2bdcde9626c ("perf/x86/intel: Add support for PEBS memory auxiliary info field in DMR") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.0+ Link: https://patch.msgid.link/20260917015234.981153-8-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraintsDapeng Mi
Same issue exists on Panther Cove, PEBS data source is valid only for these events: - MEM_TRANS_RETIRED.LOAD_LATENCY (0x1cd) - MEM_TRANS_RETIRED.STORE_SAMPLE (0x2cd) The perfmon database (https://github.com/intel/perfmon) previously tagged additional memory events such as MEM_INST_RETIRED.STLB_MISS_LOADS with L1_Hit_Indication, implying PEBS data-source support, which is incorrect. The database has since been fixed, but intel_pnc_pebs_event_constraints[] still follows the old definition and marks those events as data-source capable. As a result, get_data_src() may decode data-source information for events that do not provide valid PEBS data-source data and mislead users. Remove those non-data-source memory events from the Pather Cove PEBS constraint table so matching falls back to the regular non-PEBS constraints, which already provide the same counter constraints. Also update pnc_latency_data() to decode LOAD/STORE flags explicitly when setting memory operation direction, for consistency with other *_latency_data() helpers. Fixes: d345b6bb8860 ("perf/x86/intel: Add core PMU support for DMR") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.0+ Link: https://patch.msgid.link/20260917015234.981153-7-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Remove incorrect LionCove PEBS data-source constraintsDapeng Mi
On Lion Cove, PEBS data source is valid only for these events: - MEM_TRANS_RETIRED.LOAD_LATENCY (0x1cd) - MEM_TRANS_RETIRED.STORE_SAMPLE (0x2cd) The perfmon database (https://github.com/intel/perfmon) previously tagged additional memory events such as MEM_INST_RETIRED.STLB_MISS_LOADS with L1_Hit_Indication, implying PEBS data-source support, which is incorrect. The database has since been fixed, but intel_lnc_pebs_event_constraints[] still follows the old definition and marks those events as data-source capable. As a result, get_data_src() may decode data-source information for events that do not provide valid PEBS data-source data and mislead users. Remove those non-data-source memory events from the Lion Cove PEBS constraint table so matching falls back to the regular non-PEBS constraints, which already provide the same counter constraints. Also update lnc_latency_data() to decode LOAD/STORE flags explicitly when setting memory operation direction, for consistency with other *_latency_data() helpers. Fixes: a932aa0e868f ("perf/x86: Add Lunar Lake and Arrow Lake support") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> Link: https://patch.msgid.link/20260917015234.981153-6-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Update arw_latency_data() mem-op direction handlingDapeng Mi
Align arw_latency_data() with other *_latency_data() helpers by explicitly decoding LOAD/STORE event flags when setting the sampled memory operation direction. This keeps the latency data path behavior consistent across platforms and avoids relying on implicit direction inference. Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-5-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix ↵Dapeng Mi
sample classification Same bug exists on Darkmont as on Gracemont: intel_dkt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit LOAD/STORE flags for those events. The PEBS latency path (pebs_latency_data(), via cmt_latency_data) uses the event flags to determine memory operation direction. Without an explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be misclassified as LOADs. Set explicit LOAD/STORE flags in intel_dkt_pebs_event_constraints[] for: - MEM_UOPS_RETIRED.LOAD_LATENCY - MEM_UOPS_RETIRED.STORE_LATENCY This fixes incorrect STORE sample classification. Additionally remove INTEL_HYBRID_LAT_CONSTRAINT() since no one uses it anymore. Fixes: 65fd435095bb ("perf/x86/intel: Update event constraints for PTL") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.2+ Link: https://patch.msgid.link/20260917015234.981153-4-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix ↵Dapeng Mi
sample classification The same bug exists on Crestmont as on Gracemont: intel_cmt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit LOAD/STORE flags for those events. The PEBS latency path (pebs_latency_data(), via cmt_latency_data) uses the event flags to determine memory operation direction. Without an explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be misclassified as LOADs. Set explicit LOAD/STORE flags in intel_cmt_pebs_event_constraints[] for: - MEM_UOPS_RETIRED.LOAD_LATENCY - MEM_UOPS_RETIRED.STORE_LATENCY This fixes incorrect STORE sample classification. Fixes: e99fb45436ea ("perf/x86/intel: Update event constraints and cache_extra_regsfor MTL") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.2+ Link: https://patch.msgid.link/20260917015234.981153-3-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix ↵Dapeng Mi
sample classification On Gracemont, intel_grt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit LOAD/STORE flags for those events. The PEBS latency path (pebs_latency_data(), via __grt_latency_data()) uses the event flags to determine memory operation direction. Without an explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be misclassified as LOADs. Set explicit LOAD/STORE flags in intel_grt_pebs_event_constraints[] for: - MEM_UOPS_RETIRED.LOAD_LATENCY - MEM_UOPS_RETIRED.STORE_LATENCY Also update __grt_latency_data() to explicitly interpret these flags when assigning the sampled memory operation direction. This fixes incorrect STORE sample classification. Fixes: 39a41278f041 ("perf/x86/intel: Fix PEBS memory access info encoding for ADL") Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@vger.kernel.org> # v7.2+ Link: https://patch.msgid.link/20260917015234.981153-2-dapeng1.mi@linux.intel.com
2026-09-22perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()Sean Christopherson
Drop "support" for passing a NULL @data/@kvm_pmu param when getting guest MSRs. KVM, the only in-tree user, unconditionally passes a non-NULL pointer, and carrying code that suggests @data may be NULL is confusing, e.g. incorrectly implies that there are scenarios where KVM doesn't pass a PMU context. Fixes: 8183a538cd95 ("KVM: x86/pmu: Add IA32_DS_AREA MSR emulation to support guest DS") Signed-off-by: Sean Christopherson <seanjc@google.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Jim Mattson <jmattson@google.com> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Link: https://patch.msgid.link/20260921191418.950933-5-seanjc@google.com
2026-09-22perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) ↵Sean Christopherson
if PEBS is unused When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit, load the guest values for DS_AREA and (conditionally) MSR_PEBS_DATA_CFG if and only if PEBS will be active in the guest, i.e. only if a PEBS record may be generated while running the guest. As shown by the !pebs_ept path, it's perfectly safe to run with the host's DS_AREA, so long as PEBS-enabled counters are disabled via PERF_GLOBAL_CTRL. Omitting DS_AREA and MSR_PEBS_DATA_CFG when PEBS is unused saves two MSR writes per MSR on each VMX transition, i.e. eliminates two/four pointless MSR writes on each VMX roundtrip when PEBS isn't being used by the guest. Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson <seanjc@google.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Jim Mattson <jmattson@google.com> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Link: https://patch.msgid.link/20260921191418.950933-4-seanjc@google.com
2026-09-22perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has ↵Sean Christopherson
PEBS isolation, to fix stuck PEBS_ENABLED When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit, *never* insert an entry for PEBS_ENABLED if the CPU properly isolates PEBS events, in which case disabling counters via PERF_GLOBAL_CTRL is sufficient to prevent unwanted PEBS events in the guest (or host). Because perf loads PEBS_ENABLE with the unfiltered cpu_hw_events.pebs_enabled, i.e. with both host and guest masks, there is no need to load different values for the guest versus host, perf+KVM can and should simply control which counters are enabled/disabled via PERF_GLOBAL_CTRL. Avoiding touching PEBS_ENABLED "fixes" a bug where PEBS_ENABLED can end up with "stuck" bits if a PEBS event is throttled between generating the list and actually entering the guest (Intel CPUs can't arbtitrarily block NMIs). Fixes in quotes because leaving PEBS_ENABLED as-is doesn't fix the underlying problem of perf (via PMIs) being able to modify state after the perf<=>KVM handoff. But not writing PEBS_ENABLED is desirable no matter what, as stating the obvious, leaving PEBS_ENABLED as-is avoids three MSR writes on every VMX transition: one each on entry/exit, and one more explicit WRMSR to zero PEBS_ENABLED before VM-Entry (KVM assumes the only reason PEBS_ENABLED is in the load list is if the CPU lacks PEBS isolation and thus needs a quiescent period). Opportunistically add comments to (better) explain the rules for generating the set of PEBS counters that will be active while the guest is running, along with a FIXME for the suspected hack-a-fix where perf disables guest PEBS if _any_ PEBS event is configured to count in the host (commit 854250329c02 ("KVM: x86/pmu: Disable guest PEBS temporarily in two rare situations") doesn't explain the motivation, at all). Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson <seanjc@google.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Link: https://patch.msgid.link/20260921191418.950933-3-seanjc@google.com
2026-09-22perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted ↵Sean Christopherson
PERF_GLOBAL_CTRL bits When reinstating PEBS counters into PERF_GLOBAL_CTRL for a KVM guest, mask the value with perf's desired/original PERF_GLOBAL_CTRL value to ensure KVM doesn't unintentionally set reserved bits in PERF_GLOBAL_CTRL. E.g. if the guest's PEBS_ENABLE value had bit 63, "Enable Precise Store", set, then using the raw guest PEBS value would propagate bit 63 to the guest's PERF_GLOBAL_CTRL value (which thankfully would be a failed VM-Entry, not a VMX Abort). The only reason this bug isn't reachable is because KVM doesn't support "Enable Precise Store" (which is probably a KVM bug?), i.e. bit 63 can't be set in kvm_pmu->pebs_enable and thus not in arr[pebs_enable].guest. In other words, this _should_ be a glorified NOP in the current code base. Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson <seanjc@google.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Link: https://patch.msgid.link/20260921191418.950933-2-seanjc@google.com
2026-09-21x86/sev: Make vTPM SVSM calls preemption-safeMelody Wang
Two functions in the SVSM vTPM guest implementation do not disable preemption when fetching the SVSM Calling Area Address (CAA). The SVSM CAA is a per-CPU structure. When a thread is preempted and migrated to a different CPU after fetching the per-CPU CAA, the SVSM call will execute on the new CPU with the original CPU's CAA. Which is wrong. Move the CAA fetching operation inside svsm_perform_call_protocol() which disables interrupts around the SVSM call and thus runs preemption-safe. Fixes: 770de678bc28 ("x86/sev: Add SVSM vTPM probe/send_command functions") Signed-off-by: Melody Wang <huibo.wang@amd.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Stefano Garzarella <sgarzare@redhat.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/a5bc0d4a2c462a0089109e145c21626b244b2ff0.1789345277.git.huibo.wang@amd.com