<feed xmlns='http://www.w3.org/2005/Atom'>
<title>linux-toradex.git/Documentation/RCU/trace.txt, branch v3.17-rc6</title>
<subtitle>Linux kernel for Apalis and Colibri modules</subtitle>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/'/>
<entry>
<title>Merge branches 'doc.2013.12.03a', 'fixes.2013.12.12a', 'rcutorture.2013.12.03a' and 'sparse.2013.12.12a' into HEAD</title>
<updated>2013-12-12T20:35:38+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paulmck@linux.vnet.ibm.com</email>
</author>
<published>2013-12-12T20:35:38+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=0d3c55bc9fd58393bd3bd9974991ec1f815e1326'/>
<id>0d3c55bc9fd58393bd3bd9974991ec1f815e1326</id>
<content type='text'>
doc.2013.12.03a: Topic branch for documentation changes.
fixes.2013.12.12a: Topic branch for miscellaneous fixes.
rcutorture.2013.12.03a: Topic branch for new rcutorture/KVM scripting.
sparse.2013.12.12a: Topic branch for sparse-RCU changes.
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
doc.2013.12.03a: Topic branch for documentation changes.
fixes.2013.12.12a: Topic branch for miscellaneous fixes.
rcutorture.2013.12.03a: Topic branch for new rcutorture/KVM scripting.
sparse.2013.12.12a: Topic branch for sparse-RCU changes.
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Break call_rcu() deadlock involving scheduler and perf</title>
<updated>2013-12-03T18:10:18+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paulmck@linux.vnet.ibm.com</email>
</author>
<published>2013-10-04T21:33:34+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=96d3fd0d315a949e30adc80f086031c5cdf070d1'/>
<id>96d3fd0d315a949e30adc80f086031c5cdf070d1</id>
<content type='text'>
Dave Jones got the following lockdep splat:

&gt;  ======================================================
&gt;  [ INFO: possible circular locking dependency detected ]
&gt;  3.12.0-rc3+ #92 Not tainted
&gt;  -------------------------------------------------------
&gt;  trinity-child2/15191 is trying to acquire lock:
&gt;   (&amp;rdp-&gt;nocb_wq){......}, at: [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;
&gt; but task is already holding lock:
&gt;   (&amp;ctx-&gt;lock){-.-...}, at: [&lt;ffffffff81154c19&gt;] perf_event_exit_task+0x109/0x230
&gt;
&gt; which lock already depends on the new lock.
&gt;
&gt;
&gt; the existing dependency chain (in reverse order) is:
&gt;
&gt; -&gt; #3 (&amp;ctx-&gt;lock){-.-...}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff81733f90&gt;] _raw_spin_lock+0x40/0x80
&gt;         [&lt;ffffffff811500ff&gt;] __perf_event_task_sched_out+0x2df/0x5e0
&gt;         [&lt;ffffffff81091b83&gt;] perf_event_task_sched_out+0x93/0xa0
&gt;         [&lt;ffffffff81732052&gt;] __schedule+0x1d2/0xa20
&gt;         [&lt;ffffffff81732f30&gt;] preempt_schedule_irq+0x50/0xb0
&gt;         [&lt;ffffffff817352b6&gt;] retint_kernel+0x26/0x30
&gt;         [&lt;ffffffff813eed04&gt;] tty_flip_buffer_push+0x34/0x50
&gt;         [&lt;ffffffff813f0504&gt;] pty_write+0x54/0x60
&gt;         [&lt;ffffffff813e900d&gt;] n_tty_write+0x32d/0x4e0
&gt;         [&lt;ffffffff813e5838&gt;] tty_write+0x158/0x2d0
&gt;         [&lt;ffffffff811c4850&gt;] vfs_write+0xc0/0x1f0
&gt;         [&lt;ffffffff811c52cc&gt;] SyS_write+0x4c/0xa0
&gt;         [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2
&gt;
&gt; -&gt; #2 (&amp;rq-&gt;lock){-.-.-.}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff81733f90&gt;] _raw_spin_lock+0x40/0x80
&gt;         [&lt;ffffffff810980b2&gt;] wake_up_new_task+0xc2/0x2e0
&gt;         [&lt;ffffffff81054336&gt;] do_fork+0x126/0x460
&gt;         [&lt;ffffffff81054696&gt;] kernel_thread+0x26/0x30
&gt;         [&lt;ffffffff8171ff93&gt;] rest_init+0x23/0x140
&gt;         [&lt;ffffffff81ee1e4b&gt;] start_kernel+0x3f6/0x403
&gt;         [&lt;ffffffff81ee1571&gt;] x86_64_start_reservations+0x2a/0x2c
&gt;         [&lt;ffffffff81ee1664&gt;] x86_64_start_kernel+0xf1/0xf4
&gt;
&gt; -&gt; #1 (&amp;p-&gt;pi_lock){-.-.-.}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;         [&lt;ffffffff810979d1&gt;] try_to_wake_up+0x31/0x350
&gt;         [&lt;ffffffff81097d62&gt;] default_wake_function+0x12/0x20
&gt;         [&lt;ffffffff81084af8&gt;] autoremove_wake_function+0x18/0x40
&gt;         [&lt;ffffffff8108ea38&gt;] __wake_up_common+0x58/0x90
&gt;         [&lt;ffffffff8108ff59&gt;] __wake_up+0x39/0x50
&gt;         [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;         [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;         [&lt;ffffffff81111b8d&gt;] call_rcu+0x1d/0x20
&gt;         [&lt;ffffffff81093697&gt;] cpu_attach_domain+0x287/0x360
&gt;         [&lt;ffffffff81099d7e&gt;] build_sched_domains+0xe5e/0x10a0
&gt;         [&lt;ffffffff81efa7fc&gt;] sched_init_smp+0x3b7/0x47a
&gt;         [&lt;ffffffff81ee1f4e&gt;] kernel_init_freeable+0xf6/0x202
&gt;         [&lt;ffffffff817200be&gt;] kernel_init+0xe/0x190
&gt;         [&lt;ffffffff8173d22c&gt;] ret_from_fork+0x7c/0xb0
&gt;
&gt; -&gt; #0 (&amp;rdp-&gt;nocb_wq){......}:
&gt;         [&lt;ffffffff810cb7ca&gt;] __lock_acquire+0x191a/0x1be0
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;         [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;         [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;         [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;         [&lt;ffffffff81111bb0&gt;] kfree_call_rcu+0x20/0x30
&gt;         [&lt;ffffffff81149abf&gt;] put_ctx+0x4f/0x70
&gt;         [&lt;ffffffff81154c3e&gt;] perf_event_exit_task+0x12e/0x230
&gt;         [&lt;ffffffff81056b8d&gt;] do_exit+0x30d/0xcc0
&gt;         [&lt;ffffffff8105893c&gt;] do_group_exit+0x4c/0xc0
&gt;         [&lt;ffffffff810589c4&gt;] SyS_exit_group+0x14/0x20
&gt;         [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2
&gt;
&gt; other info that might help us debug this:
&gt;
&gt; Chain exists of:
&gt;   &amp;rdp-&gt;nocb_wq --&gt; &amp;rq-&gt;lock --&gt; &amp;ctx-&gt;lock
&gt;
&gt;   Possible unsafe locking scenario:
&gt;
&gt;         CPU0                    CPU1
&gt;         ----                    ----
&gt;    lock(&amp;ctx-&gt;lock);
&gt;                                 lock(&amp;rq-&gt;lock);
&gt;                                 lock(&amp;ctx-&gt;lock);
&gt;    lock(&amp;rdp-&gt;nocb_wq);
&gt;
&gt;  *** DEADLOCK ***
&gt;
&gt; 1 lock held by trinity-child2/15191:
&gt;  #0:  (&amp;ctx-&gt;lock){-.-...}, at: [&lt;ffffffff81154c19&gt;] perf_event_exit_task+0x109/0x230
&gt;
&gt; stack backtrace:
&gt; CPU: 2 PID: 15191 Comm: trinity-child2 Not tainted 3.12.0-rc3+ #92
&gt;  ffffffff82565b70 ffff880070c2dbf8 ffffffff8172a363 ffffffff824edf40
&gt;  ffff880070c2dc38 ffffffff81726741 ffff880070c2dc90 ffff88022383b1c0
&gt;  ffff88022383aac0 0000000000000000 ffff88022383b188 ffff88022383b1c0
&gt; Call Trace:
&gt;  [&lt;ffffffff8172a363&gt;] dump_stack+0x4e/0x82
&gt;  [&lt;ffffffff81726741&gt;] print_circular_bug+0x200/0x20f
&gt;  [&lt;ffffffff810cb7ca&gt;] __lock_acquire+0x191a/0x1be0
&gt;  [&lt;ffffffff810c6439&gt;] ? get_lock_stats+0x19/0x60
&gt;  [&lt;ffffffff8100b2f4&gt;] ? native_sched_clock+0x24/0x80
&gt;  [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;  [&lt;ffffffff8108ff43&gt;] ? __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;  [&lt;ffffffff8108ff43&gt;] ? __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;  [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;  [&lt;ffffffff8109bc8f&gt;] ? local_clock+0x3f/0x50
&gt;  [&lt;ffffffff81111bb0&gt;] kfree_call_rcu+0x20/0x30
&gt;  [&lt;ffffffff81149abf&gt;] put_ctx+0x4f/0x70
&gt;  [&lt;ffffffff81154c3e&gt;] perf_event_exit_task+0x12e/0x230
&gt;  [&lt;ffffffff81056b8d&gt;] do_exit+0x30d/0xcc0
&gt;  [&lt;ffffffff810c9af5&gt;] ? trace_hardirqs_on_caller+0x115/0x1e0
&gt;  [&lt;ffffffff810c9bcd&gt;] ? trace_hardirqs_on+0xd/0x10
&gt;  [&lt;ffffffff8105893c&gt;] do_group_exit+0x4c/0xc0
&gt;  [&lt;ffffffff810589c4&gt;] SyS_exit_group+0x14/0x20
&gt;  [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2

The underlying problem is that perf is invoking call_rcu() with the
scheduler locks held, but in NOCB mode, call_rcu() will with high
probability invoke the scheduler -- which just might want to use its
locks.  The reason that call_rcu() needs to invoke the scheduler is
to wake up the corresponding rcuo callback-offload kthread, which
does the job of starting up a grace period and invoking the callbacks
afterwards.

One solution (championed on a related problem by Lai Jiangshan) is to
simply defer the wakeup to some point where scheduler locks are no longer
held.  Since we don't want to unnecessarily incur the cost of such
deferral, the task before us is threefold:

1.	Determine when it is likely that a relevant scheduler lock is held.

2.	Defer the wakeup in such cases.

3.	Ensure that all deferred wakeups eventually happen, preferably
	sooner rather than later.

We use irqs_disabled_flags() as a proxy for relevant scheduler locks
being held.  This works because the relevant locks are always acquired
with interrupts disabled.  We may defer more often than needed, but that
is at least safe.

The wakeup deferral is tracked via a new field in the per-CPU and
per-RCU-flavor rcu_data structure, namely -&gt;nocb_defer_wakeup.

This flag is checked by the RCU core processing.  The __rcu_pending()
function now checks this flag, which causes rcu_check_callbacks()
to initiate RCU core processing at each scheduling-clock interrupt
where this flag is set.  Of course this is not sufficient because
scheduling-clock interrupts are often turned off (the things we used to
be able to count on!).  So the flags are also checked on entry to any
state that RCU considers to be idle, which includes both NO_HZ_IDLE idle
state and NO_HZ_FULL user-mode-execution state.

This approach should allow call_rcu() to be invoked regardless of what
locks you might be holding, the key word being "should".

Reported-by: Dave Jones &lt;davej@redhat.com&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Cc: Peter Zijlstra &lt;peterz@infradead.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Dave Jones got the following lockdep splat:

&gt;  ======================================================
&gt;  [ INFO: possible circular locking dependency detected ]
&gt;  3.12.0-rc3+ #92 Not tainted
&gt;  -------------------------------------------------------
&gt;  trinity-child2/15191 is trying to acquire lock:
&gt;   (&amp;rdp-&gt;nocb_wq){......}, at: [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;
&gt; but task is already holding lock:
&gt;   (&amp;ctx-&gt;lock){-.-...}, at: [&lt;ffffffff81154c19&gt;] perf_event_exit_task+0x109/0x230
&gt;
&gt; which lock already depends on the new lock.
&gt;
&gt;
&gt; the existing dependency chain (in reverse order) is:
&gt;
&gt; -&gt; #3 (&amp;ctx-&gt;lock){-.-...}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff81733f90&gt;] _raw_spin_lock+0x40/0x80
&gt;         [&lt;ffffffff811500ff&gt;] __perf_event_task_sched_out+0x2df/0x5e0
&gt;         [&lt;ffffffff81091b83&gt;] perf_event_task_sched_out+0x93/0xa0
&gt;         [&lt;ffffffff81732052&gt;] __schedule+0x1d2/0xa20
&gt;         [&lt;ffffffff81732f30&gt;] preempt_schedule_irq+0x50/0xb0
&gt;         [&lt;ffffffff817352b6&gt;] retint_kernel+0x26/0x30
&gt;         [&lt;ffffffff813eed04&gt;] tty_flip_buffer_push+0x34/0x50
&gt;         [&lt;ffffffff813f0504&gt;] pty_write+0x54/0x60
&gt;         [&lt;ffffffff813e900d&gt;] n_tty_write+0x32d/0x4e0
&gt;         [&lt;ffffffff813e5838&gt;] tty_write+0x158/0x2d0
&gt;         [&lt;ffffffff811c4850&gt;] vfs_write+0xc0/0x1f0
&gt;         [&lt;ffffffff811c52cc&gt;] SyS_write+0x4c/0xa0
&gt;         [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2
&gt;
&gt; -&gt; #2 (&amp;rq-&gt;lock){-.-.-.}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff81733f90&gt;] _raw_spin_lock+0x40/0x80
&gt;         [&lt;ffffffff810980b2&gt;] wake_up_new_task+0xc2/0x2e0
&gt;         [&lt;ffffffff81054336&gt;] do_fork+0x126/0x460
&gt;         [&lt;ffffffff81054696&gt;] kernel_thread+0x26/0x30
&gt;         [&lt;ffffffff8171ff93&gt;] rest_init+0x23/0x140
&gt;         [&lt;ffffffff81ee1e4b&gt;] start_kernel+0x3f6/0x403
&gt;         [&lt;ffffffff81ee1571&gt;] x86_64_start_reservations+0x2a/0x2c
&gt;         [&lt;ffffffff81ee1664&gt;] x86_64_start_kernel+0xf1/0xf4
&gt;
&gt; -&gt; #1 (&amp;p-&gt;pi_lock){-.-.-.}:
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;         [&lt;ffffffff810979d1&gt;] try_to_wake_up+0x31/0x350
&gt;         [&lt;ffffffff81097d62&gt;] default_wake_function+0x12/0x20
&gt;         [&lt;ffffffff81084af8&gt;] autoremove_wake_function+0x18/0x40
&gt;         [&lt;ffffffff8108ea38&gt;] __wake_up_common+0x58/0x90
&gt;         [&lt;ffffffff8108ff59&gt;] __wake_up+0x39/0x50
&gt;         [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;         [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;         [&lt;ffffffff81111b8d&gt;] call_rcu+0x1d/0x20
&gt;         [&lt;ffffffff81093697&gt;] cpu_attach_domain+0x287/0x360
&gt;         [&lt;ffffffff81099d7e&gt;] build_sched_domains+0xe5e/0x10a0
&gt;         [&lt;ffffffff81efa7fc&gt;] sched_init_smp+0x3b7/0x47a
&gt;         [&lt;ffffffff81ee1f4e&gt;] kernel_init_freeable+0xf6/0x202
&gt;         [&lt;ffffffff817200be&gt;] kernel_init+0xe/0x190
&gt;         [&lt;ffffffff8173d22c&gt;] ret_from_fork+0x7c/0xb0
&gt;
&gt; -&gt; #0 (&amp;rdp-&gt;nocb_wq){......}:
&gt;         [&lt;ffffffff810cb7ca&gt;] __lock_acquire+0x191a/0x1be0
&gt;         [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;         [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;         [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;         [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;         [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;         [&lt;ffffffff81111bb0&gt;] kfree_call_rcu+0x20/0x30
&gt;         [&lt;ffffffff81149abf&gt;] put_ctx+0x4f/0x70
&gt;         [&lt;ffffffff81154c3e&gt;] perf_event_exit_task+0x12e/0x230
&gt;         [&lt;ffffffff81056b8d&gt;] do_exit+0x30d/0xcc0
&gt;         [&lt;ffffffff8105893c&gt;] do_group_exit+0x4c/0xc0
&gt;         [&lt;ffffffff810589c4&gt;] SyS_exit_group+0x14/0x20
&gt;         [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2
&gt;
&gt; other info that might help us debug this:
&gt;
&gt; Chain exists of:
&gt;   &amp;rdp-&gt;nocb_wq --&gt; &amp;rq-&gt;lock --&gt; &amp;ctx-&gt;lock
&gt;
&gt;   Possible unsafe locking scenario:
&gt;
&gt;         CPU0                    CPU1
&gt;         ----                    ----
&gt;    lock(&amp;ctx-&gt;lock);
&gt;                                 lock(&amp;rq-&gt;lock);
&gt;                                 lock(&amp;ctx-&gt;lock);
&gt;    lock(&amp;rdp-&gt;nocb_wq);
&gt;
&gt;  *** DEADLOCK ***
&gt;
&gt; 1 lock held by trinity-child2/15191:
&gt;  #0:  (&amp;ctx-&gt;lock){-.-...}, at: [&lt;ffffffff81154c19&gt;] perf_event_exit_task+0x109/0x230
&gt;
&gt; stack backtrace:
&gt; CPU: 2 PID: 15191 Comm: trinity-child2 Not tainted 3.12.0-rc3+ #92
&gt;  ffffffff82565b70 ffff880070c2dbf8 ffffffff8172a363 ffffffff824edf40
&gt;  ffff880070c2dc38 ffffffff81726741 ffff880070c2dc90 ffff88022383b1c0
&gt;  ffff88022383aac0 0000000000000000 ffff88022383b188 ffff88022383b1c0
&gt; Call Trace:
&gt;  [&lt;ffffffff8172a363&gt;] dump_stack+0x4e/0x82
&gt;  [&lt;ffffffff81726741&gt;] print_circular_bug+0x200/0x20f
&gt;  [&lt;ffffffff810cb7ca&gt;] __lock_acquire+0x191a/0x1be0
&gt;  [&lt;ffffffff810c6439&gt;] ? get_lock_stats+0x19/0x60
&gt;  [&lt;ffffffff8100b2f4&gt;] ? native_sched_clock+0x24/0x80
&gt;  [&lt;ffffffff810cc243&gt;] lock_acquire+0x93/0x200
&gt;  [&lt;ffffffff8108ff43&gt;] ? __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8173419b&gt;] _raw_spin_lock_irqsave+0x4b/0x90
&gt;  [&lt;ffffffff8108ff43&gt;] ? __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8108ff43&gt;] __wake_up+0x23/0x50
&gt;  [&lt;ffffffff8110d4f8&gt;] __call_rcu_nocb_enqueue+0xa8/0xc0
&gt;  [&lt;ffffffff81111450&gt;] __call_rcu+0x140/0x820
&gt;  [&lt;ffffffff8109bc8f&gt;] ? local_clock+0x3f/0x50
&gt;  [&lt;ffffffff81111bb0&gt;] kfree_call_rcu+0x20/0x30
&gt;  [&lt;ffffffff81149abf&gt;] put_ctx+0x4f/0x70
&gt;  [&lt;ffffffff81154c3e&gt;] perf_event_exit_task+0x12e/0x230
&gt;  [&lt;ffffffff81056b8d&gt;] do_exit+0x30d/0xcc0
&gt;  [&lt;ffffffff810c9af5&gt;] ? trace_hardirqs_on_caller+0x115/0x1e0
&gt;  [&lt;ffffffff810c9bcd&gt;] ? trace_hardirqs_on+0xd/0x10
&gt;  [&lt;ffffffff8105893c&gt;] do_group_exit+0x4c/0xc0
&gt;  [&lt;ffffffff810589c4&gt;] SyS_exit_group+0x14/0x20
&gt;  [&lt;ffffffff8173d4e4&gt;] tracesys+0xdd/0xe2

The underlying problem is that perf is invoking call_rcu() with the
scheduler locks held, but in NOCB mode, call_rcu() will with high
probability invoke the scheduler -- which just might want to use its
locks.  The reason that call_rcu() needs to invoke the scheduler is
to wake up the corresponding rcuo callback-offload kthread, which
does the job of starting up a grace period and invoking the callbacks
afterwards.

One solution (championed on a related problem by Lai Jiangshan) is to
simply defer the wakeup to some point where scheduler locks are no longer
held.  Since we don't want to unnecessarily incur the cost of such
deferral, the task before us is threefold:

1.	Determine when it is likely that a relevant scheduler lock is held.

2.	Defer the wakeup in such cases.

3.	Ensure that all deferred wakeups eventually happen, preferably
	sooner rather than later.

We use irqs_disabled_flags() as a proxy for relevant scheduler locks
being held.  This works because the relevant locks are always acquired
with interrupts disabled.  We may defer more often than needed, but that
is at least safe.

The wakeup deferral is tracked via a new field in the per-CPU and
per-RCU-flavor rcu_data structure, namely -&gt;nocb_defer_wakeup.

This flag is checked by the RCU core processing.  The __rcu_pending()
function now checks this flag, which causes rcu_check_callbacks()
to initiate RCU core processing at each scheduling-clock interrupt
where this flag is set.  Of course this is not sufficient because
scheduling-clock interrupts are often turned off (the things we used to
be able to count on!).  So the flags are also checked on entry to any
state that RCU considers to be idle, which includes both NO_HZ_IDLE idle
state and NO_HZ_FULL user-mode-execution state.

This approach should allow call_rcu() to be invoked regardless of what
locks you might be holding, the key word being "should".

Reported-by: Dave Jones &lt;davej@redhat.com&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Cc: Peter Zijlstra &lt;peterz@infradead.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Fix typo in Documentation/RCU/trace.txt</title>
<updated>2013-12-03T18:08:56+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paulmck@linux.vnet.ibm.com</email>
</author>
<published>2013-10-09T21:34:00+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=c9592ecb98d1e6d6ab2b54895314a46abfa69db4'/>
<id>c9592ecb98d1e6d6ab2b54895314a46abfa69db4</id>
<content type='text'>
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Remove TINY_PREEMPT_RCU tracing documentation</title>
<updated>2013-06-10T20:45:52+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paulmck@linux.vnet.ibm.com</email>
</author>
<published>2013-03-27T18:32:12+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=7807acdb6b794b3af42dfa1912edd4fba8d0b622'/>
<id>7807acdb6b794b3af42dfa1912edd4fba8d0b622</id>
<content type='text'>
Because TINY_PREEMPT_RCU is no more, this commit removes its tracing
formats from the documentation.

Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Because TINY_PREEMPT_RCU is no more, this commit removes its tracing
formats from the documentation.

Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Add documentation for the new rcuexp debugfs trace file</title>
<updated>2012-11-16T17:55:31+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paul.mckenney@linaro.org</email>
</author>
<published>2012-10-31T20:22:08+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=d484a215139cf556cb718a7ec7042260b7fc2d28'/>
<id>d484a215139cf556cb718a7ec7042260b7fc2d28</id>
<content type='text'>
This commit adds the documentation of the rcuexp debugfs trace file
that records statistics for expedited grace periods.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
This commit adds the documentation of the rcuexp debugfs trace file
that records statistics for expedited grace periods.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Update documentation for TREE_RCU debugfs tracing</title>
<updated>2012-11-16T17:54:02+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paul.mckenney@linaro.org</email>
</author>
<published>2012-10-31T20:00:15+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=40e80c469f1b52a68e09da3808a1228cf9947fa7'/>
<id>40e80c469f1b52a68e09da3808a1228cf9947fa7</id>
<content type='text'>
This commit updates the tracing documentation to reflect the new
format that has per-RCU-flavor directories.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
This commit updates the tracing documentation to reflect the new
format that has per-RCU-flavor directories.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Adjust debugfs tracing for kthread-based quiescent-state forcing</title>
<updated>2012-09-23T14:41:54+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paulmck@linux.vnet.ibm.com</email>
</author>
<published>2012-06-26T21:00:48+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=4605c0143c6d611b3076025ba3a7e04293c01d69'/>
<id>4605c0143c6d611b3076025ba3a7e04293c01d69</id>
<content type='text'>
Moving quiescent-state forcing into a kthread dispenses with the need
for the -&gt;n_rp_need_fqs field, so this commit removes it.

Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Moving quiescent-state forcing into a kthread dispenses with the need
for the -&gt;n_rp_need_fqs field, so this commit removes it.

Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Rework detection of use of RCU by offline CPUs</title>
<updated>2012-02-21T17:06:07+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paul.mckenney@linaro.org</email>
</author>
<published>2012-01-31T01:02:47+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=2036d94a7b61ca5032ce90f2bda06afec0fe713e'/>
<id>2036d94a7b61ca5032ce90f2bda06afec0fe713e</id>
<content type='text'>
Because newly offlined CPUs continue executing after completing the
CPU_DYING notifiers, they legitimately enter the scheduler and use
RCU while appearing to be offline.  This calls for a more sophisticated
approach as follows:

1.	RCU marks the CPU online during the CPU_UP_PREPARE phase.

2.	RCU marks the CPU offline during the CPU_DEAD phase.

3.	Diagnostics regarding use of read-side RCU by offline CPUs use
	RCU's accounting rather than the cpu_online_map.  (Note that
	__call_rcu() still uses cpu_online_map to detect illegal
	invocations within CPU_DYING notifiers.)

4.	Offline CPUs are prevented from hanging the system by
	force_quiescent_state(), which pays attention to cpu_online_map.
	Some additional work (in a later commit) will be needed to
	guarantee that force_quiescent_state() waits a full jiffy before
	assuming that a CPU is offline, for example, when called from
	idle entry.  (This commit also makes the one-jiffy wait
	explicit, since the old-style implicit wait can now be defeated
	by RCU_FAST_NO_HZ and by rcutorture.)

This approach avoids the false positives encountered when attempting to
use more exact classification of CPU online/offline state.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Because newly offlined CPUs continue executing after completing the
CPU_DYING notifiers, they legitimately enter the scheduler and use
RCU while appearing to be offline.  This calls for a more sophisticated
approach as follows:

1.	RCU marks the CPU online during the CPU_UP_PREPARE phase.

2.	RCU marks the CPU offline during the CPU_DEAD phase.

3.	Diagnostics regarding use of read-side RCU by offline CPUs use
	RCU's accounting rather than the cpu_online_map.  (Note that
	__call_rcu() still uses cpu_online_map to detect illegal
	invocations within CPU_DYING notifiers.)

4.	Offline CPUs are prevented from hanging the system by
	force_quiescent_state(), which pays attention to cpu_online_map.
	Some additional work (in a later commit) will be needed to
	guarantee that force_quiescent_state() waits a full jiffy before
	assuming that a CPU is offline, for example, when called from
	idle entry.  (This commit also makes the one-jiffy wait
	explicit, since the old-style implicit wait can now be defeated
	by RCU_FAST_NO_HZ and by rcutorture.)

This approach avoids the false positives encountered when attempting to
use more exact classification of CPU online/offline state.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Track idleness independent of idle tasks</title>
<updated>2011-12-11T18:31:24+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paul.mckenney@linaro.org</email>
</author>
<published>2011-09-30T19:10:22+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=9b2e4f1880b789be1f24f9684f7a54b90310b5c0'/>
<id>9b2e4f1880b789be1f24f9684f7a54b90310b5c0</id>
<content type='text'>
Earlier versions of RCU used the scheduling-clock tick to detect idleness
by checking for the idle task, but handled idleness differently for
CONFIG_NO_HZ=y.  But there are now a number of uses of RCU read-side
critical sections in the idle task, for example, for tracing.  A more
fine-grained detection of idleness is therefore required.

This commit presses the old dyntick-idle code into full-time service,
so that rcu_idle_enter(), previously known as rcu_enter_nohz(), is
always invoked at the beginning of an idle loop iteration.  Similarly,
rcu_idle_exit(), previously known as rcu_exit_nohz(), is always invoked
at the end of an idle-loop iteration.  This allows the idle task to
use RCU everywhere except between consecutive rcu_idle_enter() and
rcu_idle_exit() calls, in turn allowing architecture maintainers to
specify exactly where in the idle loop that RCU may be used.

Because some of the userspace upcall uses can result in what looks
to RCU like half of an interrupt, it is not possible to expect that
the irq_enter() and irq_exit() hooks will give exact counts.  This
patch therefore expands the -&gt;dynticks_nesting counter to 64 bits
and uses two separate bitfields to count process/idle transitions
and interrupt entry/exit transitions.  It is presumed that userspace
upcalls do not happen in the idle loop or from usermode execution
(though usermode might do a system call that results in an upcall).
The counter is hard-reset on each process/idle transition, which
avoids the interrupt entry/exit error from accumulating.  Overflow
is avoided by the 64-bitness of the -&gt;dyntick_nesting counter.

This commit also adds warnings if a non-idle task asks RCU to enter
idle state (and these checks will need some adjustment before applying
Frederic's OS-jitter patches (http://lkml.org/lkml/2011/10/7/246).
In addition, validation of -&gt;dynticks and -&gt;dynticks_nesting is added.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Earlier versions of RCU used the scheduling-clock tick to detect idleness
by checking for the idle task, but handled idleness differently for
CONFIG_NO_HZ=y.  But there are now a number of uses of RCU read-side
critical sections in the idle task, for example, for tracing.  A more
fine-grained detection of idleness is therefore required.

This commit presses the old dyntick-idle code into full-time service,
so that rcu_idle_enter(), previously known as rcu_enter_nohz(), is
always invoked at the beginning of an idle loop iteration.  Similarly,
rcu_idle_exit(), previously known as rcu_exit_nohz(), is always invoked
at the end of an idle-loop iteration.  This allows the idle task to
use RCU everywhere except between consecutive rcu_idle_enter() and
rcu_idle_exit() calls, in turn allowing architecture maintainers to
specify exactly where in the idle loop that RCU may be used.

Because some of the userspace upcall uses can result in what looks
to RCU like half of an interrupt, it is not possible to expect that
the irq_enter() and irq_exit() hooks will give exact counts.  This
patch therefore expands the -&gt;dynticks_nesting counter to 64 bits
and uses two separate bitfields to count process/idle transitions
and interrupt entry/exit transitions.  It is presumed that userspace
upcalls do not happen in the idle loop or from usermode execution
(though usermode might do a system call that results in an upcall).
The counter is hard-reset on each process/idle transition, which
avoids the interrupt entry/exit error from accumulating.  Overflow
is avoided by the 64-bitness of the -&gt;dyntick_nesting counter.

This commit also adds warnings if a non-idle task asks RCU to enter
idle state (and these checks will need some adjustment before applying
Frederic's OS-jitter patches (http://lkml.org/lkml/2011/10/7/246).
In addition, validation of -&gt;dynticks and -&gt;dynticks_nesting is added.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
Reviewed-by: Josh Triplett &lt;josh@joshtriplett.org&gt;
</pre>
</div>
</content>
</entry>
<entry>
<title>rcu: Simplify quiescent-state accounting</title>
<updated>2011-09-29T04:38:22+00:00</updated>
<author>
<name>Paul E. McKenney</name>
<email>paul.mckenney@linaro.org</email>
</author>
<published>2011-06-27T07:17:43+00:00</published>
<link rel='alternate' type='text/html' href='https://git.toradex.cn/cgit/linux-toradex.git/commit/?id=e4cc1f22b2f4e9b0207a8cdb63e56dcf99e82d35'/>
<id>e4cc1f22b2f4e9b0207a8cdb63e56dcf99e82d35</id>
<content type='text'>
There is often a delay between the time that a CPU passes through a
quiescent state and the time that this quiescent state is reported to the
RCU core.  It is quite possible that the grace period ended before the
quiescent state could be reported, for example, some other CPU might have
deduced that this CPU passed through dyntick-idle mode.  It is critically
important that quiescent state be counted only against the grace period
that was in effect at the time that the quiescent state was detected.

Previously, this was handled by recording the number of the last grace
period to complete when passing through a quiescent state.  The RCU
core then checks this number against the current value, and rejects
the quiescent state if there is a mismatch.  However, one additional
possibility must be accounted for, namely that the quiescent state was
recorded after the prior grace period completed but before the current
grace period started.  In this case, the RCU core must reject the
quiescent state, but the recorded number will match.  This is handled
when the CPU becomes aware of a new grace period -- at that point,
it invalidates any prior quiescent state.

This works, but is a bit indirect.  The new approach records the current
grace period, and the RCU core checks to see (1) that this is still the
current grace period and (2) that this grace period has not yet ended.
This approach simplifies reasoning about correctness, and this commit
changes over to this new approach.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
There is often a delay between the time that a CPU passes through a
quiescent state and the time that this quiescent state is reported to the
RCU core.  It is quite possible that the grace period ended before the
quiescent state could be reported, for example, some other CPU might have
deduced that this CPU passed through dyntick-idle mode.  It is critically
important that quiescent state be counted only against the grace period
that was in effect at the time that the quiescent state was detected.

Previously, this was handled by recording the number of the last grace
period to complete when passing through a quiescent state.  The RCU
core then checks this number against the current value, and rejects
the quiescent state if there is a mismatch.  However, one additional
possibility must be accounted for, namely that the quiescent state was
recorded after the prior grace period completed but before the current
grace period started.  In this case, the RCU core must reject the
quiescent state, but the recorded number will match.  This is handled
when the CPU becomes aware of a new grace period -- at that point,
it invalidates any prior quiescent state.

This works, but is a bit indirect.  The new approach records the current
grace period, and the RCU core checks to see (1) that this is still the
current grace period and (2) that this grace period has not yet ended.
This approach simplifies reasoning about correctness, and this commit
changes over to this new approach.

Signed-off-by: Paul E. McKenney &lt;paul.mckenney@linaro.org&gt;
Signed-off-by: Paul E. McKenney &lt;paulmck@linux.vnet.ibm.com&gt;
</pre>
</div>
</content>
</entry>
</feed>
