linux-toradex.git/kernel/sched.c, branch v2.6.32.58

ftrace: Fix memory leak with function graph and cpu hotplug

2011-03-23T20:16:39+00:00

commit 868baf07b1a259f5f3803c1dc2777b6c358f83cf upstream.

When the fuction graph tracer starts, it needs to make a special
stack for each task to save the real return values of the tasks.
All running tasks have this stack created, as well as any new
tasks.

On CPU hot plug, the new idle task will allocate a stack as well
when init_idle() is called. The problem is that cpu hotplug does
not create a new idle_task. Instead it uses the idle task that
existed when the cpu went down.

ftrace_graph_init_task() will add a new ret_stack to the task
that is given to it. Because a clone will make the task
have a stack of its parent it does not check if the task's
ret_stack is already NULL or not. When the CPU hotplug code
starts a CPU up again, it will allocate a new stack even
though one already existed for it.

The solution is to treat the idle_task specially. In fact, the
function_graph code already does, just not at init_idle().
Instead of using the ftrace_graph_init_task() for the idle task,
which that function expects the task to be a clone, have a
separate ftrace_graph_init_idle_task(). Also, we will create a
per_cpu ret_stack that is used by the idle task. When we call
ftrace_graph_init_idle_task() it will check if the idle task's
ret_stack is NULL, if it is, then it will assign it the per_cpu
ret_stack.

Reported-by: Benjamin Herrenschmidt 
Suggested-by: Peter Zijlstra 
Signed-off-by: Steven Rostedt 
Signed-off-by: Greg Kroah-Hartman

sched: Fix wake_affine() vs RT tasks

2011-02-17T23:37:30+00:00

Commit: e51fd5e22e12b39f49b1bb60b37b300b17378a43 upstream

Mike reports that since e9e9250b (sched: Scale down cpu_power due to RT
tasks), wake_affine() goes funny on RT tasks due to them still having a
!0 weight and wake_affine() still subtracts that from the rq weight.

Since nobody should be using se->weight for RT tasks, set the value to
zero. Also, since we now use ->cpu_power to normalize rq weights to
account for RT cpu usage, add that factor into the imbalance computation.

Reported-by: Mike Galbraith 
Tested-by: Mike Galbraith 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1275316109.27810.22969.camel@twins>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Fix idle balancing

2011-02-17T23:37:30+00:00

Commit: d5ad140bc1505a98c0f040937125bfcbb508078f upstream

An earlier commit reverts idle balancing throttling reset to fix a 30%
regression in volanomark throughput. We still need to reset idle_stamp
when we pull a task in newidle balance.

Reported-by: Alex Shi 
Signed-off-by: Nikhil Rao 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1290022924-3548-1-git-send-email-ncrao@google.com>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Fix volanomark performance regression

2011-02-17T23:37:29+00:00

Commit: b5482cfa1c95a188b3054fa33274806add91bbe5 upstream

Commit fab4762 triggers excessive idle balancing, causing a ~30% loss in
volanomark throughput. Remove idle balancing throttle reset.

Originally-by: Alex Shi 
Signed-off-by: Mike Galbraith 
Acked-by: Nikhil Rao 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1289928732.5169.211.camel@maggy.simson.net>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Fix cross-sched-class wakeup preemption

2011-02-17T23:37:29+00:00

Commit: 1e5a74059f9059d330744eac84873b1b99657008 upstream

Instead of dealing with sched classes inside each check_preempt_curr()
implementation, pull out this logic into the generic wakeup preemption
path.

This fixes a hang in KVM (and others) where we are waiting for the
stop machine thread to run ...

Reported-by: Markus Trippelsdorf 
Tested-by: Marcelo Tosatti 
Tested-by: Sergey Senozhatsky 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1288891946.2039.31.camel@laptop>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Use group weight, idle cpu metrics to fix imbalances during idle

2011-02-17T23:37:28+00:00

Commit: aae6d3ddd8b90f5b2c8d79a2b914d1706d124193 upstream

Currently we consider a sched domain to be well balanced when the imbalance
is less than the domain's imablance_pct. As the number of cores and threads
are increasing, current values of imbalance_pct (for example 25% for a
NUMA domain) are not enough to detect imbalances like:

a) On a WSM-EP system (two sockets, each having 6 cores and 12 logical threads),
24 cpu-hogging tasks get scheduled as 13 on one socket and 11 on another
socket. Leading to an idle HT cpu.

b) On a hypothetial 2 socket NHM-EX system (each socket having 8 cores and
16 logical threads), 16 cpu-hogging tasks can get scheduled as 9 on one
socket and 7 on another socket. Leaving one core in a socket idle
whereas in another socket we have a core having both its HT siblings busy.

While this issue can be fixed by decreasing the domain's imbalance_pct
(by making it a function of number of logical cpus in the domain), it
can potentially cause more task migrations across sched groups in an
overloaded case.

Fix this by using imbalance_pct only during newly_idle and busy
load balancing. And during idle load balancing, check if there
is an imbalance in number of idle cpu's across the busiest and this
sched_group or if the busiest group has more tasks than its weight that
the idle cpu in this_group can pull.

Reported-by: Nikhil Rao 
Signed-off-by: Suresh Siddha 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1284760952.2676.11.camel@sbsiddha-MOBL3.sc.intel.com>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched, cgroup: Fixup broken cgroup movement

2011-02-17T23:37:28+00:00

Commit: b2b5ce022acf5e9f52f7b78c5579994fdde191d4 upstream

Dima noticed that we fail to correct the ->vruntime of sleeping tasks
when we move them between cgroups.

Reported-by: Dima Zavin 
Signed-off-by: Peter Zijlstra 
Tested-by: Mike Galbraith 
LKML-Reference: <1287150604.29097.1513.camel@twins>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Export account_system_vtime()

2011-02-17T23:37:27+00:00

Commit: b7dadc38797584f6203386da1947ed5edf516646 upstream

KVM uses it for example:

 ERROR: "account_system_vtime" [arch/x86/kvm/kvm.ko] undefined!

Cc: Venkatesh Pallipadi 
Cc: Peter Zijlstra 
LKML-Reference: <1286237003-12406-3-git-send-email-venki@google.com>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Call tick_check_idle before __irq_enter

2011-02-17T23:37:27+00:00

Commit: d267f87fb8179c6dba03d08b91952e81bc3723c7 upstream

When CPU is idle and on first interrupt, irq_enter calls tick_check_idle()
to notify interruption from idle. But, there is a problem if this call
is done after __irq_enter, as all routines in __irq_enter may find
stale time due to yet to be done tick_check_idle.

Specifically, trace calls in __irq_enter when they use global clock and also
account_system_vtime change in this patch as it wants to use sched_clock_cpu()
to do proper irq timing.

But, tick_check_idle was moved after __irq_enter intentionally to
prevent problem of unneeded ksoftirqd wakeups by the commit ee5f80a:

    irq: call __irq_enter() before calling the tick_idle_check
    Impact: avoid spurious ksoftirqd wakeups

Moving tick_check_idle() before __irq_enter and wrapping it with
local_bh_enable/disable would solve both the problems.

Fixed-by: Yong Zhang 
Signed-off-by: Venkatesh Pallipadi 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1286237003-12406-9-git-send-email-venki@google.com>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman

sched: Remove irq time from available CPU power

2011-02-17T23:37:27+00:00

Commit: aa483808516ca5cacfa0e5849691f64fec25828e upstream

The idea was suggested by Peter Zijlstra here:

  http://marc.info/?l=linux-kernel&m=127476934517534&w=2

irq time is technically not available to the tasks running on the CPU.
This patch removes irq time from CPU power piggybacking on
sched_rt_avg_update().

Tested this by keeping CPU X busy with a network intensive task having 75%
oa a single CPU irq processing (hard+soft) on a 4-way system. And start seven
cycle soakers on the system. Without this change, there will be two tasks on
each CPU. With this change, there is a single task on irq busy CPU X and
remaining 7 tasks are spread around among other 3 CPUs.

Signed-off-by: Venkatesh Pallipadi 
Signed-off-by: Peter Zijlstra 
LKML-Reference: <1286237003-12406-8-git-send-email-venki@google.com>
Signed-off-by: Ingo Molnar 
Signed-off-by: Mike Galbraith 
Acked-by: Peter Zijlstra 
Signed-off-by: Greg Kroah-Hartman