linux-toradex.git/include/linux/percpu.h, branch v6.16-rc6

alloc_tag: allocate percpu counters for module tags dynamically

2025-05-25T07:53:48+00:00

When a module gets unloaded it checks whether any of its tags are still in
use and if so, we keep the memory containing module's allocation tags
alive until all tags are unused.  However percpu counters referenced by
the tags are freed by free_module().  This will lead to UAF if the memory
allocated by a module is accessed after module was unloaded.

To fix this we allocate percpu counters for module allocation tags
dynamically and we keep it alive for tags which are still in use after
module unloading.  This also removes the requirement of a larger
PERCPU_MODULE_RESERVE when memory allocation profiling is enabled because
percpu memory for counters does not need to be reserved anymore.

Link: https://lkml.kernel.org/r/20250517000739.5930-1-surenb@google.com
Fixes: 0db6f8d7820a ("alloc_tag: load module tags into separate contiguous memory")
Signed-off-by: Suren Baghdasaryan 
Reported-by: David Wang <00107082@163.com>
Closes: https://lore.kernel.org/all/20250516131246.6244-1-00107082@163.com/
Tested-by: David Wang <00107082@163.com>
Cc: Christoph Lameter (Ampere) 
Cc: Dennis Zhou 
Cc: Kent Overstreet 
Cc: Pasha Tatashin 
Cc: Tejun Heo 
Cc: 
Signed-off-by: Andrew Morton

mm: percpu: increase PERCPU_DYNAMIC_SIZE_SHIFT on certain builds.

2024-10-17T07:28:07+00:00

Arnd reported a build failure due to the BUILD_BUG_ON() statement in
alloc_kmem_cache_cpus().  The test

  PERCPU_DYNAMIC_EARLY_SIZE < NR_KMALLOC_TYPES * KMALLOC_SHIFT_HIGH * sizeof(struct kmem_cache_cpu)

The factors that increase the right side of the equation:
- PAGE_SIZE > 4KiB increases KMALLOC_SHIFT_HIGH
- For the local_lock_t in kmem_cache_cpu:
  - PREEMPT_RT adds an actual lock.
  - LOCKDEP increases the size of the lock.
  - LOCK_STAT adds additional bytes plus padding to the lockdep
    structure.

The net difference with and without PREEMPT_RT is 88 bytes for the
lock_lock_t, 96 bytes for kmem_cache_cpu due to additional padding.  This
is enough to exceed the 80KiB limit with 16KiB page size - the 8KiB page
size is fine.

Increase PERCPU_DYNAMIC_SIZE_SHIFT to 13 on configs with PAGE_SIZE larger
than 4KiB and LOCKDEP enabled.

Link: https://lkml.kernel.org/r/20241007143049.gyMpEu89@linutronix.de
Fixes: d8fccd9ca5f9 ("arm64: Allow to enable PREEMPT_RT.")
Signed-off-by: Sebastian Andrzej Siewior 
Reported-by: kernel test robot 
Closes: https://lore.kernel.org/oe-kbuild-all/202410020326.iaZIteIx-lkp@intel.com/
Reported-by: Arnd Bergmann 
Closes: https://lore.kernel.org/20241004095702.637528-1-arnd@kernel.org
Acked-by: Arnd Bergmann 
Acked-by: Vlastimil Babka 
Acked-by: David Rientjes 
Cc: Christoph Lameter 
Cc: Dennis Zhou 
Cc: Hyeonggon Yoo <42.hyeyoo@gmail.com>
Cc: Joonsoo Kim 
Cc: Pekka Enberg 
Cc: Roman Gushchin 
Cc: Tejun Heo 
Cc: Thomas Gleixner 
Signed-off-by: Andrew Morton

percpu: remove pcpu_alloc_size()

2024-09-02T03:26:04+00:00

pcpu_alloc_size() was added in 7ac5c53e0073 "mm/percpu.c: introduce
pcpu_alloc_size()", which is used to get the allocated memory size in bpf.
However, pcpu_alloc_size() is no longer used in "bpf: Use c->unit_size to
select target cache during free" because its actuall allocated memory size
may change at runtime due to its slab merging mechanism.  Therefore,
pcpu_alloc_size() can be removed.

Link: https://lkml.kernel.org/r/tencent_AD5C50E8D78C07A3CE539BD5F6BF39706507@qq.com
Signed-off-by: Jianhui Zhou <912460177@qq.com>
Cc: Christoph Lameter 
Cc: Dennis Zhou 
Cc: JonasZhou 
Cc: Tejun Heo 
Signed-off-by: Andrew Morton

cpumask: cleanup core headers inclusion

2024-06-25T05:25:02+00:00

Many core headers include cpumask.h for nothing. Drop it.

Link: https://lkml.kernel.org/r/20240528005648.182376-6-yury.norov@gmail.com
Signed-off-by: Yury Norov 
Cc: Amit Daniel Kachhap 
Cc: Anna-Maria Behnsen 
Cc: Christoph Lameter 
Cc: Daniel Lezcano 
Cc: Dennis Zhou 
Cc: Frederic Weisbecker 
Cc: Johannes Weiner 
Cc: Juri Lelli 
Cc: Kees Cook 
Cc: Mathieu Desnoyers 
Cc: Paul E. McKenney 
Cc: Peter Zijlstra 
Cc: Rafael J. Wysocki 
Cc: Rasmus Villemoes 
Cc: Tejun Heo 
Cc: Thomas Gleixner 
Cc: Ulf Hansson 
Cc: Vincent Guittot 
Cc: Viresh Kumar 
Cc: Yury Norov 
Signed-off-by: Andrew Morton

mm: change inlined allocation helpers to account at the call site

2024-04-26T03:55:59+00:00

Main goal of memory allocation profiling patchset is to provide accounting
that is cheap enough to run in production.  To achieve that we inject
counters using codetags at the allocation call sites to account every time
allocation is made.  This injection allows us to perform accounting
efficiently because injected counters are immediately available as opposed
to the alternative methods, such as using _RET_IP_, which would require
counter lookup and appropriate locking that makes accounting much more
expensive.  This method requires all allocation functions to inject
separate counters at their call sites so that their callers can be
individually accounted.  Counter injection is implemented by allocation
hooks which should wrap all allocation functions.

Inlined functions which perform allocations but do not use allocation
hooks are directly charged for the allocations they perform.  In most
cases these functions are just specialized allocation wrappers used from
multiple places to allocate objects of a specific type.  It would be more
useful to do the accounting at their call sites instead.  Instrument these
helpers to do accounting at the call site.  Simple inlined allocation
wrappers are converted directly into macros.  More complex allocators or
allocators with documentation are converted into _noprof versions and
allocation hooks are added.  This allows memory allocation profiling
mechanism to charge allocations to the callers of these functions.

Link: https://lkml.kernel.org/r/20240415020731.1152108-1-surenb@google.com
Signed-off-by: Suren Baghdasaryan 
Acked-by: Jan Kara 		[jbd2]
Cc: Anna Schumaker 
Cc: Arnd Bergmann 
Cc: Benjamin Tissoires 
Cc: Christoph Lameter 
Cc: David Rientjes 
Cc: David S. Miller 
Cc: Dennis Zhou 
Cc: Eric Dumazet 
Cc: Herbert Xu 
Cc: Jakub Kicinski 
Cc: Jakub Sitnicki 
Cc: Jiri Kosina 
Cc: Joerg Roedel 
Cc: Joonsoo Kim 
Cc: Kent Overstreet 
Cc: Matthew Wilcox (Oracle) 
Cc: Paolo Abeni 
Cc: Pekka Enberg 
Cc: Tejun Heo 
Cc: Theodore Ts'o 
Cc: Trond Myklebust 
Cc: Vlastimil Babka 
Cc: Will Deacon 
Signed-off-by: Andrew Morton

mm: percpu: enable per-cpu allocation tagging

2024-04-26T03:55:56+00:00

Redefine __alloc_percpu, __alloc_percpu_gfp and __alloc_reserved_percpu
to record allocations and deallocations done by these functions.

[surenb@google.com: undo _noprof additions in the documentation]
  Link: https://lkml.kernel.org/r/20240326231453.1206227-6-surenb@google.com
Link: https://lkml.kernel.org/r/20240321163705.3067592-30-surenb@google.com
Signed-off-by: Kent Overstreet 
Signed-off-by: Suren Baghdasaryan 
Tested-by: Kees Cook 
Cc: Alexander Viro 
Cc: Alex Gaynor 
Cc: Alice Ryhl 
Cc: Andreas Hindborg 
Cc: Benno Lossin 
Cc: "Björn Roy Baron" 
Cc: Boqun Feng 
Cc: Christoph Lameter 
Cc: Dennis Zhou 
Cc: Gary Guo 
Cc: Miguel Ojeda 
Cc: Pasha Tatashin 
Cc: Peter Zijlstra 
Cc: Tejun Heo 
Cc: Vlastimil Babka 
Cc: Wedson Almeida Filho 
Signed-off-by: Andrew Morton

mm: percpu: increase PERCPU_MODULE_RESERVE to accommodate allocation tags

2024-04-26T03:55:53+00:00

As each allocation tag generates a per-cpu variable, more space is
required to store them.  Increase PERCPU_MODULE_RESERVE to provide enough
area.  A better long-term solution would be to allocate this memory
dynamically.

[surenb@google.com: increase PERCPU_MODULE_RESERVE to accommodate allocation tags]
  Link: https://lkml.kernel.org/r/20240406214044.1114406-1-surenb@google.com
Link: https://lkml.kernel.org/r/20240321163705.3067592-17-surenb@google.com
Signed-off-by: Suren Baghdasaryan 
Signed-off-by: Kent Overstreet 
Tested-by: Kees Cook 
Cc: Peter Zijlstra 
Cc: Tejun Heo 
Cc: Alexander Viro 
Cc: Alex Gaynor 
Cc: Alice Ryhl 
Cc: Andreas Hindborg 
Cc: Benno Lossin 
Cc: "Björn Roy Baron" 
Cc: Boqun Feng 
Cc: Christoph Lameter 
Cc: Dennis Zhou 
Cc: Gary Guo 
Cc: Miguel Ojeda 
Cc: Pasha Tatashin 
Cc: Vlastimil Babka 
Cc: Wedson Almeida Filho 
Signed-off-by: Andrew Morton

mm/percpu.c: introduce pcpu_alloc_size()

2023-10-20T21:15:06+00:00

Introduce pcpu_alloc_size() to get the size of the dynamic per-cpu
area. It will be used by bpf memory allocator in the following patches.
BPF memory allocator maintains per-cpu area caches for multiple area
sizes and its free API only has the to-be-freed per-cpu pointer, so it
needs the size of dynamic per-cpu area to select the corresponding cache
when bpf program frees the dynamic per-cpu pointer.

Acked-by: Dennis Zhou 
Signed-off-by: Hou Tao 
Link: https://lore.kernel.org/r/20231020133202.4043247-3-houtao@huaweicloud.com
Signed-off-by: Alexei Starovoitov

Randomized slab caches for kmalloc()

2023-07-18T08:07:47+00:00

When exploiting memory vulnerabilities, "heap spraying" is a common
technique targeting those related to dynamic memory allocation (i.e. the
"heap"), and it plays an important role in a successful exploitation.
Basically, it is to overwrite the memory area of vulnerable object by
triggering allocation in other subsystems or modules and therefore
getting a reference to the targeted memory location. It's usable on
various types of vulnerablity including use after free (UAF), heap out-
of-bound write and etc.

There are (at least) two reasons why the heap can be sprayed: 1) generic
slab caches are shared among different subsystems and modules, and
2) dedicated slab caches could be merged with the generic ones.
Currently these two factors cannot be prevented at a low cost: the first
one is a widely used memory allocation mechanism, and shutting down slab
merging completely via `slub_nomerge` would be overkill.

To efficiently prevent heap spraying, we propose the following approach:
to create multiple copies of generic slab caches that will never be
merged, and random one of them will be used at allocation. The random
selection is based on the address of code that calls `kmalloc()`, which
means it is static at runtime (rather than dynamically determined at
each time of allocation, which could be bypassed by repeatedly spraying
in brute force). In other words, the randomness of cache selection will
be with respect to the code address rather than time, i.e. allocations
in different code paths would most likely pick different caches,
although kmalloc() at each place would use the same cache copy whenever
it is executed. In this way, the vulnerable object and memory allocated
in other subsystems and modules will (most probably) be on different
slab caches, which prevents the object from being sprayed.

Meanwhile, the static random selection is further enhanced with a
per-boot random seed, which prevents the attacker from finding a usable
kmalloc that happens to pick the same cache with the vulnerable
subsystem/module by analyzing the open source code. In other words, with
the per-boot seed, the random selection is static during each time the
system starts and runs, but not across different system startups.

The overhead of performance has been tested on a 40-core x86 server by
comparing the results of `perf bench all` between the kernels with and
without this patch based on the latest linux-next kernel, which shows
minor difference. A subset of benchmarks are listed below:

                sched/  sched/  syscall/       mem/       mem/
             messaging    pipe     basic     memcpy     memset
                 (sec)   (sec)     (sec)   (GB/sec)   (GB/sec)

control1         0.019   5.459     0.733  15.258789  51.398026
control2         0.019   5.439     0.730  16.009221  48.828125
control3         0.019   5.282     0.735  16.009221  48.828125
control_avg      0.019   5.393     0.733  15.759077  49.684759

experiment1      0.019   5.374     0.741  15.500992  46.502976
experiment2      0.019   5.440     0.746  16.276042  51.398026
experiment3      0.019   5.242     0.752  15.258789  51.398026
experiment_avg   0.019   5.352     0.746  15.678608  49.766343

The overhead of memory usage was measured by executing `free` after boot
on a QEMU VM with 1GB total memory, and as expected, it's positively
correlated with # of cache copies:

           control  4 copies  8 copies  16 copies

total       969.8M    968.2M    968.2M     968.2M
used         20.0M     21.9M     24.1M      26.7M
free        936.9M    933.6M    931.4M     928.6M
available   932.2M    928.8M    926.6M     923.9M

Co-developed-by: Xiu Jianfeng 
Signed-off-by: Xiu Jianfeng 
Signed-off-by: GONG, Ruiqi 
Reviewed-by: Kees Cook 
Reviewed-by: Hyeonggon Yoo <42.hyeyoo@gmail.com>
Acked-by: Dennis Zhou  # percpu
Signed-off-by: Vlastimil Babka

Merge tag 'core_guards_for_6.5_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/peterz/queue

2023-07-04T20:50:38+00:00

Pull scope-based resource management infrastructure from Peter Zijlstra:
 "These are the first few patches in the Scope-based Resource Management
  series that introduce the infrastructure but not any conversions as of
  yet.

  Adding the infrastructure now allows multiple people to start using
  them.

  Of note is that Sparse will need some work since it doesn't yet
  understand this attribute and might have decl-after-stmt issues"

* tag 'core_guards_for_6.5_rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/peterz/queue:
  kbuild: Drop -Wdeclaration-after-statement
  locking: Introduce __cleanup() based infrastructure
  apparmor: Free up __cleanup() name
  dmaengine: ioat: Free up __cleanup() name