linux-toradex.git/fs/proc, branch v4.4.93

mm: larger stack guard gap, between vmas

2017-06-26T05:13:11+00:00

commit 1be7107fbe18eed3e319a6c3e83c78254b693acb upstream.

Stack guard page is a useful feature to reduce a risk of stack smashing
into a different mapping. We have been using a single page gap which
is sufficient to prevent having stack adjacent to a different mapping.
But this seems to be insufficient in the light of the stack usage in
userspace. E.g. glibc uses as large as 64kB alloca() in many commonly
used functions. Others use constructs liks gid_t buffer[NGROUPS_MAX]
which is 256kB or stack strings with MAX_ARG_STRLEN.

This will become especially dangerous for suid binaries and the default
no limit for the stack size limit because those applications can be
tricked to consume a large portion of the stack and a single glibc call
could jump over the guard page. These attacks are not theoretical,
unfortunatelly.

Make those attacks less probable by increasing the stack guard gap
to 1MB (on systems with 4k pages; but make it depend on the page size
because systems with larger base pages might cap stack allocations in
the PAGE_SIZE units) which should cover larger alloca() and VLA stack
allocations. It is obviously not a full fix because the problem is
somehow inherent, but it should reduce attack space a lot.

One could argue that the gap size should be configurable from userspace,
but that can be done later when somebody finds that the new 1MB is wrong
for some special case applications.  For now, add a kernel command line
option (stack_guard_gap) to specify the stack gap size (in page units).

Implementation wise, first delete all the old code for stack guard page:
because although we could get away with accounting one extra page in a
stack vma, accounting a larger gap can break userspace - case in point,
a program run with "ulimit -S -v 20000" failed when the 1MB gap was
counted for RLIMIT_AS; similar problems could come with RLIMIT_MLOCK
and strict non-overcommit mode.

Instead of keeping gap inside the stack vma, maintain the stack guard
gap as a gap between vmas: using vm_start_gap() in place of vm_start
(or vm_end_gap() in place of vm_end if VM_GROWSUP) in just those few
places which need to respect the gap - mainly arch_get_unmapped_area(),
and and the vma tree's subtree_gap support for that.

Original-patch-by: Oleg Nesterov 
Original-patch-by: Michal Hocko 
Signed-off-by: Hugh Dickins 
Acked-by: Michal Hocko 
Tested-by: Helge Deller  # parisc
Signed-off-by: Linus Torvalds 
[wt: backport to 4.11: adjust context]
[wt: backport to 4.9: adjust context ; kernel doc was not in admin-guide]
[wt: backport to 4.4: adjust context ; drop ppc hugetlb_radix changes]
Signed-off-by: Willy Tarreau 
[gkh: minor build fixes for 4.4]
Signed-off-by: Greg Kroah-Hartman

proc: add a schedule point in proc_pid_readdir()

2017-06-17T04:39:38+00:00

[ Upstream commit 3ba4bceef23206349d4130ddf140819b365de7c8 ]

We have seen proc_pid_readdir() invocations holding cpu for more than 50
ms.  Add a cond_resched() to be gentle with other tasks.

[akpm@linux-foundation.org: coding style fix]
Link: http://lkml.kernel.org/r/1484238380.15816.42.camel@edumazet-glaptop3.roam.corp.google.com
Signed-off-by: Eric Dumazet 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 

Signed-off-by: Sasha Levin 
Signed-off-by: Greg Kroah-Hartman

proc: Fix unbalanced hard link numbers

2017-05-25T12:30:10+00:00

commit d66bb1607e2d8d384e53f3d93db5c18483c8c4f7 upstream.

proc_create_mount_point() forgot to increase the parent's nlink, and
it resulted in unbalanced hard link numbers, e.g. /proc/fs shows one
less than expected.

Fixes: eb6d38d5427b ("proc: Allow creating permanently empty directories...")
Reported-by: Tristan Ye 
Signed-off-by: Takashi Iwai 
Signed-off-by: Eric W. Biederman 
Signed-off-by: Greg Kroah-Hartman

thp: fix MADV_DONTNEED vs clear soft dirty race

2017-04-21T07:30:04+00:00

commit 5b7abeae3af8c08c577e599dd0578b9e3ee6687b upstream.

Yet another instance of the same race.

Fix is identical to change_huge_pmd().

See "thp: fix MADV_DONTNEED vs.  numa balancing race" for more details.

Link: http://lkml.kernel.org/r/20170302151034.27829-5-kirill.shutemov@linux.intel.com
Signed-off-by: Kirill A. Shutemov 
Cc: Andrea Arcangeli 
Cc: Hillf Danton 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 
Signed-off-by: Greg Kroah-Hartman

sysctl: Drop reference added by grab_header in proc_sys_readdir

2017-01-19T19:17:21+00:00

commit 93362fa47fe98b62e4a34ab408c4a418432e7939 upstream.

Fixes CVE-2016-9191, proc_sys_readdir doesn't drop reference
added by grab_header when return from !dir_emit_dots path.
It can cause any path called unregister_sysctl_table will
wait forever.

The calltrace of CVE-2016-9191:

[ 5535.960522] Call Trace:
[ 5535.963265]  [] schedule+0x3f/0xa0
[ 5535.968817]  [] schedule_timeout+0x3db/0x6f0
[ 5535.975346]  [] ? wait_for_completion+0x45/0x130
[ 5535.982256]  [] wait_for_completion+0xc3/0x130
[ 5535.988972]  [] ? wake_up_q+0x80/0x80
[ 5535.994804]  [] drop_sysctl_table+0xc4/0xe0
[ 5536.001227]  [] drop_sysctl_table+0x77/0xe0
[ 5536.007648]  [] unregister_sysctl_table+0x4d/0xa0
[ 5536.014654]  [] unregister_sysctl_table+0x7f/0xa0
[ 5536.021657]  [] unregister_sched_domain_sysctl+0x15/0x40
[ 5536.029344]  [] partition_sched_domains+0x44/0x450
[ 5536.036447]  [] ? __mutex_unlock_slowpath+0x111/0x1f0
[ 5536.043844]  [] rebuild_sched_domains_locked+0x64/0xb0
[ 5536.051336]  [] update_flag+0x11d/0x210
[ 5536.057373]  [] ? mutex_lock_nested+0x2df/0x450
[ 5536.064186]  [] ? cpuset_css_offline+0x1b/0x60
[ 5536.070899]  [] ? trace_hardirqs_on+0xd/0x10
[ 5536.077420]  [] ? mutex_lock_nested+0x2df/0x450
[ 5536.084234]  [] ? css_killed_work_fn+0x25/0x220
[ 5536.091049]  [] cpuset_css_offline+0x35/0x60
[ 5536.097571]  [] css_killed_work_fn+0x5c/0x220
[ 5536.104207]  [] process_one_work+0x1df/0x710
[ 5536.110736]  [] ? process_one_work+0x160/0x710
[ 5536.117461]  [] worker_thread+0x12b/0x4a0
[ 5536.123697]  [] ? process_one_work+0x710/0x710
[ 5536.130426]  [] kthread+0xfe/0x120
[ 5536.135991]  [] ret_from_fork+0x1f/0x40
[ 5536.142041]  [] ? kthread_create_on_node+0x230/0x230

One cgroup maintainer mentioned that "cgroup is trying to offline
a cpuset css, which takes place under cgroup_mutex.  The offlining
ends up trying to drain active usages of a sysctl table which apprently
is not happening."
The real reason is that proc_sys_readdir doesn't drop reference added
by grab_header when return from !dir_emit_dots path. So this cpuset
offline path will wait here forever.

See here for details: http://www.openwall.com/lists/oss-security/2016/11/04/13

Fixes: f0c3b5093add ("[readdir] convert procfs")
Reported-by: CAI Qian 
Tested-by: Yang Shukui 
Signed-off-by: Zhou Chengming 
Acked-by: Al Viro 
Signed-off-by: Eric W. Biederman 
Signed-off-by: Greg Kroah-Hartman

mm: introduce get_task_exe_file

2016-09-24T08:07:36+00:00

commit cd81a9170e69e018bbaba547c1fd85a585f5697a upstream.

For more convenient access if one has a pointer to the task.

As a minor nit take advantage of the fact that only task lock + rcu are
needed to safely grab ->exe_file. This saves mm refcount dance.

Use the helper in proc_exe_link.

Signed-off-by: Mateusz Guzik 
Acked-by: Konstantin Khlebnikov 
Acked-by: Richard Guy Briggs 
Signed-off-by: Paul Moore 
Signed-off-by: Greg Kroah-Hartman

proc: revert /proc//maps [stack:TID] annotation

2016-09-15T06:27:46+00:00

[ Upstream commit 65376df582174ffcec9e6471bf5b0dd79ba05e4a ]

Commit b76437579d13 ("procfs: mark thread stack correctly in
proc//maps") added [stack:TID] annotation to /proc//maps.

Finding the task of a stack VMA requires walking the entire thread list,
turning this into quadratic behavior: a thousand threads means a
thousand stacks, so the rendering of /proc//maps needs to look at a
million combinations.

The cost is not in proportion to the usefulness as described in the
patch.

Drop the [stack:TID] annotation to make /proc//maps (and
/proc//numa_maps) usable again for higher thread counts.

The [stack] annotation inside /proc//task//maps is retained, as
identifying the stack VMA there is an O(1) operation.

Siddesh said:
 "The end users needed a way to identify thread stacks programmatically and
  there wasn't a way to do that.  I'm afraid I no longer remember (or have
  access to the resources that would aid my memory since I changed
  employers) the details of their requirement.  However, I did do this on my
  own time because I thought it was an interesting project for me and nobody
  really gave any feedback then as to its utility, so as far as I am
  concerned you could roll back the main thread maps information since the
  information is available in the thread-specific files"

Signed-off-by: Johannes Weiner 
Cc: "Kirill A. Shutemov" 
Cc: Siddhesh Poyarekar 
Cc: Shaohua Li 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 
Signed-off-by: Sasha Levin 
Signed-off-by: Greg Kroah-Hartman

proc: prevent stacking filesystems on top

2016-06-24T17:18:20+00:00

commit e54ad7f1ee263ffa5a2de9c609d58dfa27b21cd9 upstream.

This prevents stacking filesystems (ecryptfs and overlayfs) from using
procfs as lower filesystem.  There is too much magic going on inside
procfs, and there is no good reason to stack stuff on top of procfs.

(For example, procfs does access checks in VFS open handlers, and
ecryptfs by design calls open handlers from a kernel thread that doesn't
drop privileges or so.)

Signed-off-by: Jann Horn 
Signed-off-by: Linus Torvalds 
Signed-off-by: Greg Kroah-Hartman

proc: prevent accessing /proc//environ until it's ready

2016-05-11T09:21:16+00:00

commit 8148a73c9901a8794a50f950083c00ccf97d43b3 upstream.

If /proc//environ gets read before the envp[] array is fully set up
in create_{aout,elf,elf_fdpic,flat}_tables(), we might end up trying to
read more bytes than are actually written, as env_start will already be
set but env_end will still be zero, making the range calculation
underflow, allowing to read beyond the end of what has been written.

Fix this as it is done for /proc//cmdline by testing env_end for
zero.  It is, apparently, intentionally set last in create_*_tables().

This bug was found by the PaX size_overflow plugin that detected the
arithmetic underflow of 'this_len = env_end - (env_start + src)' when
env_end is still zero.

The expected consequence is that userland trying to access
/proc//environ of a not yet fully set up process may get
inconsistent data as we're in the middle of copying in the environment
variables.

Fixes: https://forums.grsecurity.net/viewtopic.php?f=3&t=4363
Fixes: https://bugzilla.kernel.org/show_bug.cgi?id=116461
Signed-off-by: Mathias Krause 
Cc: Emese Revfy 
Cc: Pax Team 
Cc: Al Viro 
Cc: Mateusz Guzik 
Cc: Alexey Dobriyan 
Cc: Cyrill Gorcunov 
Cc: Jarod Wilson 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 
Signed-off-by: Greg Kroah-Hartman

numa: fix /proc//numa_maps for THP

2016-05-04T21:48:49+00:00

commit 28093f9f34cedeaea0f481c58446d9dac6dd620f upstream.

In gather_pte_stats() a THP pmd is cast into a pte, which is wrong
because the layouts may differ depending on the architecture.  On s390
this will lead to inaccurate numa_maps accounting in /proc because of
misguided pte_present() and pte_dirty() checks on the fake pte.

On other architectures pte_present() and pte_dirty() may work by chance,
but there may be an issue with direct-access (dax) mappings w/o
underlying struct pages when HAVE_PTE_SPECIAL is set and THP is
available.  In vm_normal_page() the fake pte will be checked with
pte_special() and because there is no "special" bit in a pmd, this will
always return false and the VM_PFNMAP | VM_MIXEDMAP checking will be
skipped.  On dax mappings w/o struct pages, an invalid struct page
pointer would then be returned that can crash the kernel.

This patch fixes the numa_maps THP handling by introducing new "_pmd"
variants of the can_gather_numa_stats() and vm_normal_page() functions.

Signed-off-by: Gerald Schaefer 
Cc: Naoya Horiguchi 
Cc: "Kirill A . Shutemov" 
Cc: Konstantin Khlebnikov 
Cc: Michal Hocko 
Cc: Vlastimil Babka 
Cc: Jerome Marchand 
Cc: Johannes Weiner 
Cc: Dave Hansen 
Cc: Mel Gorman 
Cc: Dan Williams 
Cc: Martin Schwidefsky 
Cc: Heiko Carstens 
Cc: Michael Holzheu 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 
Signed-off-by: Greg Kroah-Hartman