linux-toradex.git/include/linux/memcontrol.h, branch v6.0-rc1

mm: memcontrol: introduce mem_cgroup_ino() and mem_cgroup_get_from_ino()

2022-07-04T01:08:40+00:00

Patch series "mm: introduce shrinker debugfs interface", v5.

The only existing debugging mechanism is a couple of tracepoints in
do_shrink_slab(): mm_shrink_slab_start and mm_shrink_slab_end.  They
aren't covering everything though: shrinkers which report 0 objects will
never show up, there is no support for memcg-aware shrinkers.  Shrinkers
are identified by their scan function, which is not always enough (e.g. 
hard to guess which super block's shrinker it is having only
"super_cache_scan").

To provide a better visibility and debug options for memory shrinkers this
patchset introduces a /sys/kernel/debug/shrinker interface, to some extent
similar to /sys/kernel/slab.

For each shrinker registered in the system a directory is created.  As
now, the directory will contain only a "scan" file, which allows to get
the number of managed objects for each memory cgroup (for memcg-aware
shrinkers) and each numa node (for numa-aware shrinkers on a numa
machine).  Other interfaces might be added in the future.

To make debugging more pleasant, the patchset also names all shrinkers, so
that debugfs entries can have meaningful names.


This patch (of 5):

Shrinker debugfs requires a way to represent memory cgroups without using
full paths, both for displaying information and getting input from a user.

Cgroup inode number is a perfect way, already used by bpf.

This commit adds a couple of helper functions which will be used to handle
memcg-aware shrinkers.

Link: https://lkml.kernel.org/r/20220601032227.4076670-1-roman.gushchin@linux.dev
Link: https://lkml.kernel.org/r/20220601032227.4076670-2-roman.gushchin@linux.dev
Signed-off-by: Roman Gushchin 
Acked-by: Muchun Song 
Cc: Dave Chinner 
Cc: Kent Overstreet 
Cc: Hillf Danton 
Cc: Christophe JAILLET 
Cc: Roman Gushchin 
Signed-off-by: Andrew Morton

net: set proper memcg for net_init hooks allocations

2022-06-17T02:48:31+00:00

__register_pernet_operations() executes init hook of registered
pernet_operation structure in all existing net namespaces.

Typically, these hooks are called by a process associated with the
specified net namespace, and all __GFP_ACCOUNT marked allocation are
accounted for corresponding container/memcg.

However __register_pernet_operations() calls the hooks in the same
context, and as a result all marked allocations are accounted to one memcg
for all processed net namespaces.

This patch adjusts active memcg for each net namespace and helps to
account memory allocated inside ops_init() into the proper memcg.

Link: https://lkml.kernel.org/r/f9394752-e272-9bf9-645f-a18c56d1c4ec@openvz.org
Signed-off-by: Vasily Averin 
Acked-by: Roman Gushchin 
Acked-by: Shakeel Butt 
Cc: Michal Koutný 
Cc: Vlastimil Babka 
Cc: Michal Hocko 
Cc: Florian Westphal 
Cc: David S. Miller 
Cc: Jakub Kicinski 
Cc: Paolo Abeni 
Cc: Eric Dumazet 
Cc: Johannes Weiner 
Cc: Kefeng Wang 
Cc: Linux Kernel Functional Testing 
Cc: Muchun Song 
Cc: Naresh Kamboju 
Cc: Qian Cai 
Signed-off-by: Andrew Morton

mm: kmem: make mem_cgroup_from_obj() vmalloc()-safe

2022-06-17T02:48:31+00:00

Currently mem_cgroup_from_obj() is not working properly with objects
allocated using vmalloc().  It creates problems in some cases, when it's
called for static objects belonging to modules or generally allocated
using vmalloc().

This patch makes mem_cgroup_from_obj() safe to be called on objects
allocated using vmalloc().

It also introduces mem_cgroup_from_slab_obj(), which is a faster version
to use in places when we know the object is either a slab object or a
generic slab page (e.g.  when adding an object to a lru list).

Link: https://lkml.kernel.org/r/20220610180310.1725111-1-roman.gushchin@linux.dev
Suggested-by: Kefeng Wang 
Signed-off-by: Roman Gushchin 
Tested-by: Linux Kernel Functional Testing 
Acked-by: Shakeel Butt 
Tested-by: Vasily Averin 
Acked-by: Michal Hocko 
Acked-by: Muchun Song 
Cc: Johannes Weiner 
Cc: Naresh Kamboju 
Cc: Qian Cai 
Cc: Kefeng Wang 
Cc: David S. Miller 
Cc: Eric Dumazet 
Cc: Florian Westphal 
Cc: Jakub Kicinski 
Cc: Michal Koutný 
Cc: Paolo Abeni 
Cc: Vlastimil Babka 
Signed-off-by: Andrew Morton

zswap: memcg accounting

2022-05-19T21:08:53+00:00

Applications can currently escape their cgroup memory containment when
zswap is enabled.  This patch adds per-cgroup tracking and limiting of
zswap backend memory to rectify this.

The existing cgroup2 memory.stat file is extended to show zswap statistics
analogous to what's in meminfo and vmstat.  Furthermore, two new control
files, memory.zswap.current and memory.zswap.max, are added to allow
tuning zswap usage on a per-workload basis.  This is important since not
all workloads benefit from zswap equally; some even suffer compared to
disk swap when memory contents don't compress well.  The optimal size of
the zswap pool, and the threshold for writeback, also depends on the size
of the workload's warm set.

The implementation doesn't use a traditional page_counter transaction. 
zswap is unconventional as a memory consumer in that we only know the
amount of memory to charge once expensive compression has occurred.  If
zwap is disabled or the limit is already exceeded we obviously don't want
to compress page upon page only to reject them all.  Instead, the limit is
checked against current usage, then we compress and charge.  This allows
some limit overrun, but not enough to matter in practice.

[hannes@cmpxchg.org: fix for CONFIG_SLOB builds]
  Link: https://lkml.kernel.org/r/YnwD14zxYjUJPc2w@cmpxchg.org
[hannes@cmpxchg.org: opt out of cgroups v1]
  Link: https://lkml.kernel.org/r/Yn6it9mBYFA+/lTb@cmpxchg.org
Link: https://lkml.kernel.org/r/20220510152847.230957-7-hannes@cmpxchg.org
Signed-off-by: Johannes Weiner 
Cc: Michal Hocko 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Seth Jennings 
Cc: Dan Streetman 
Cc: Minchan Kim 
Signed-off-by: Andrew Morton

vmscan: convert lazy freeing to folios

2022-05-13T14:20:15+00:00

Remove a hidden call to compound_head(), and account nr_pages instead of a
single page.  This matches the code in lru_lazyfree_fn() that accounts
nr_pages to PGLAZYFREE.

Link: https://lkml.kernel.org/r/20220504182857.4013401-12-willy@infradead.org
Signed-off-by: Matthew Wilcox (Oracle) 
Reviewed-by: Christoph Hellwig 
Signed-off-by: Andrew Morton

mm/memcontrol.c: make cgroup_memory_noswap static

2022-04-29T06:16:00+00:00

cgroup_memory_noswap is only used in mm/memcontrol.c, therefore just make
it static, and remove export in include/linux/memcontrol.h

Link: https://lkml.kernel.org/r/20220421124736.62180-1-lujialin4@huawei.com
Signed-off-by: Lu Jialin 
Acked-by: Johannes Weiner 
Acked-by: Shakeel Butt 
Acked-by: Roman Gushchin 
Reviewed-by: Muchun Song 
Cc: Michal Hocko 
Signed-off-by: Andrew Morton

memcg: sync flush only if periodic flush is delayed

2022-04-22T03:01:09+00:00

Daniel Dao has reported [1] a regression on workloads that may trigger a
lot of refaults (anon and file).  The underlying issue is that flushing
rstat is expensive.  Although rstat flush are batched with (nr_cpus *
MEMCG_BATCH) stat updates, it seems like there are workloads which
genuinely do stat updates larger than batch value within short amount of
time.  Since the rstat flush can happen in the performance critical
codepaths like page faults, such workload can suffer greatly.

This patch fixes this regression by making the rstat flushing
conditional in the performance critical codepaths.  More specifically,
the kernel relies on the async periodic rstat flusher to flush the stats
and only if the periodic flusher is delayed by more than twice the
amount of its normal time window then the kernel allows rstat flushing
from the performance critical codepaths.

Now the question: what are the side-effects of this change? The worst
that can happen is the refault codepath will see 4sec old lruvec stats
and may cause false (or missed) activations of the refaulted page which
may under-or-overestimate the workingset size.  Though that is not very
concerning as the kernel can already miss or do false activations.

There are two more codepaths whose flushing behavior is not changed by
this patch and we may need to come to them in future.  One is the
writeback stats used by dirty throttling and second is the deactivation
heuristic in the reclaim.  For now keeping an eye on them and if there
is report of regression due to these codepaths, we will reevaluate then.

Link: https://lore.kernel.org/all/CA+wXwBSyO87ZX5PVwdHm-=dBjZYECGmfnydUicUyrQqndgX2MQ@mail.gmail.com [1]
Link: https://lkml.kernel.org/r/20220304184040.1304781-1-shakeelb@google.com
Fixes: 1f828223b799 ("memcg: flush lruvec stats in the refault")
Signed-off-by: Shakeel Butt 
Reported-by: Daniel Dao 
Tested-by: Ivan Babrou 
Cc: Michal Hocko 
Cc: Roman Gushchin 
Cc: Johannes Weiner 
Cc: Michal Koutný 
Cc: Frank Hofmann 
Cc: 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: memcontrol: rename memcg_cache_id to memcg_kmem_id

2022-03-22T22:57:04+00:00

The memcg_cache_id() introduced by commit 2633d7a02823 ("slab/slub:
consider a memcg parameter in kmem_create_cache") is used to index in the
kmem_cache->memcg_params->memcg_caches array.  Since
kmem_cache->memcg_params.memcg_caches has been removed by commit
9855609bde03 ("mm: memcg/slab: use a single set of kmem_caches for all
accounted allocations").  So the name does not need to reflect cache
related.  Just rename it to memcg_kmem_id.  And it can reflect kmem
related.

Link: https://lkml.kernel.org/r/20220228122126.37293-17-songmuchun@bytedance.com
Signed-off-by: Muchun Song 
Cc: Alex Shi 
Cc: Anna Schumaker 
Cc: Chao Yu 
Cc: Dave Chinner 
Cc: Fam Zheng 
Cc: Jaegeuk Kim 
Cc: Johannes Weiner 
Cc: Kari Argillander 
Cc: Matthew Wilcox (Oracle) 
Cc: Michal Hocko 
Cc: Qi Zheng 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Theodore Ts'o 
Cc: Trond Myklebust 
Cc: Vladimir Davydov 
Cc: Vlastimil Babka 
Cc: Wei Yang 
Cc: Xiongchun Duan 
Cc: Yang Shi 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: list_lru: replace linear array with xarray

2022-03-22T22:57:03+00:00

If we run 10k containers in the system, the size of the
list_lru_memcg->lrus can be ~96KB per list_lru.  When we decrease the
number containers, the size of the array will not be shrinked.  It is
not scalable.  The xarray is a good choice for this case.  We can save a
lot of memory when there are tens of thousands continers in the system.
If we use xarray, we also can remove the logic code of resizing array,
which can simplify the code.

[akpm@linux-foundation.org: remove unused local]

Link: https://lkml.kernel.org/r/20220228122126.37293-13-songmuchun@bytedance.com
Signed-off-by: Muchun Song 
Cc: Alex Shi 
Cc: Anna Schumaker 
Cc: Chao Yu 
Cc: Dave Chinner 
Cc: Fam Zheng 
Cc: Jaegeuk Kim 
Cc: Johannes Weiner 
Cc: Kari Argillander 
Cc: Matthew Wilcox (Oracle) 
Cc: Michal Hocko 
Cc: Qi Zheng 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Theodore Ts'o 
Cc: Trond Myklebust 
Cc: Vladimir Davydov 
Cc: Vlastimil Babka 
Cc: Wei Yang 
Cc: Xiongchun Duan 
Cc: Yang Shi 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: introduce kmem_cache_alloc_lru

2022-03-22T22:57:03+00:00

We currently allocate scope for every memcg to be able to tracked on
every superblock instantiated in the system, regardless of whether that
superblock is even accessible to that memcg.

These huge memcg counts come from container hosts where memcgs are
confined to just a small subset of the total number of superblocks that
instantiated at any given point in time.

For these systems with huge container counts, list_lru does not need the
capability of tracking every memcg on every superblock.  What it comes
down to is that adding the memcg to the list_lru at the first insert.
So introduce kmem_cache_alloc_lru to allocate objects and its list_lru.
In the later patch, we will convert all inode and dentry allocation from
kmem_cache_alloc to kmem_cache_alloc_lru.

Link: https://lkml.kernel.org/r/20220228122126.37293-3-songmuchun@bytedance.com
Signed-off-by: Muchun Song 
Cc: Alex Shi 
Cc: Anna Schumaker 
Cc: Chao Yu 
Cc: Dave Chinner 
Cc: Fam Zheng 
Cc: Jaegeuk Kim 
Cc: Johannes Weiner 
Cc: Kari Argillander 
Cc: Matthew Wilcox (Oracle) 
Cc: Michal Hocko 
Cc: Qi Zheng 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Theodore Ts'o 
Cc: Trond Myklebust 
Cc: Vladimir Davydov 
Cc: Vlastimil Babka 
Cc: Wei Yang 
Cc: Xiongchun Duan 
Cc: Yang Shi 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds