linux-toradex.git/block, branch v2.6.33.5

block: ensure jiffies wrap is handled correctly in blk_rq_timed_out_timer

2010-05-12T22:02:43+00:00

commit a534dbe96e9929c7245924d8252d89048c23d569 upstream.

blk_rq_timed_out_timer() relied on blk_add_timer() never returning a
timer value of zero, but commit 7838c15b8dd18e78a523513749e5b54bda07b0cb
removed the code that bumped this value when it was zero.
Therefore when jiffies is near wrap we could get unlucky & not set the
timeout value correctly.

This patch uses a flag to indicate that the timeout value was set and so
handles jiffies wrap correctly, and it keeps all the logic in one
function so should be easier to maintain in the future.

Signed-off-by: Richard Kennedy 
Signed-off-by: Jens Axboe 
Signed-off-by: Greg Kroah-Hartman

Revert "block: improve queue_should_plug() by looking at IO depths"

2010-02-23T07:40:43+00:00

This reverts commit fb1e75389bd06fd5987e9cda1b4e0305c782f854.

"Benjamin S."  reports that the patch in question
causes a big drop in sequential throughput for him, dropping from
200MB/sec down to only 70MB/sec.

Needs to be investigated more fully, for now lets just revert the
offending commit.

Conflicts:

	include/linux/blkdev.h

Signed-off-by: Jens Axboe

cfq-iosched: split seeky coop queues after one slice

2010-02-05T12:11:45+00:00

Currently we split seeky coop queues after 1s, which is too big. Below patch
marks seeky coop queue split_coop flag after one slice. After that, if new
requests come in, the queues will be splitted. Patch is suggested by Corrado.

Signed-off-by: Shaohua Li 
Reviewed-by: Corrado Zoccolo 
Acked-by: Jeff Moyer 
Signed-off-by: Jens Axboe

cfq-iosched: Do not idle on async queues

2010-02-02T19:46:10+00:00

Few weeks back, Shaohua Li had posted similar patch. I am reposting it
with more test results.

This patch does two things.

- Do not idle on async queues.

- It also changes the write queue depth CFQ drives (cfq_may_dispatch()).
  Currently, we seem to driving queue depth of 1 always for WRITES. This is
  true even if there is only one write queue in the system and all the logic
  of infinite queue depth in case of single busy queue as well as slowly
  increasing queue depth based on last delayed sync request does not seem to
  be kicking in at all.

This patch will allow deeper WRITE queue depths (subjected to the other
WRITE queue depth contstraints like cfq_quantum and last delayed sync
request).

Shaohua Li had reported getting more out of his SSD. For me, I have got
one Lun exported from an HP EVA and when pure buffered writes are on, I
can get more out of the system. Following are test results of pure
buffered writes (with end_fsync=1) with vanilla and patched kernel. These
results are average of 3 sets of run with increasing number of threads.

AVERAGE[bufwfs][vanilla]
-------
job       Set NR  ReadBW(KB/s)   MaxClat(us)    WriteBW(KB/s)  MaxClat(us)
---       --- --  ------------   -----------    -------------  -----------
bufwfs    3   1   0              0              95349          474141
bufwfs    3   2   0              0              100282         806926
bufwfs    3   4   0              0              109989         2.7301e+06
bufwfs    3   8   0              0              116642         3762231
bufwfs    3   16  0              0              118230         6902970

AVERAGE[bufwfs] [patched kernel]
-------
bufwfs    3   1   0              0              270722         404352
bufwfs    3   2   0              0              206770         1.06552e+06
bufwfs    3   4   0              0              195277         1.62283e+06
bufwfs    3   8   0              0              260960         2.62979e+06
bufwfs    3   16  0              0              299260         1.70731e+06

I also ran buffered writes along with some sequential reads and some
buffered reads going on in the system on a SATA disk because the potential
risk could be that we should not be driving queue depth higher in presence
of sync IO going to keep the max clat low.

With some random and sequential reads going on in the system on one SATA
disk I did not see any significant increase in max clat. So it looks like
other WRITE queue depth control logic is doing its job. Here are the
results.

AVERAGE[brr, bsr, bufw together] [vanilla]
-------
job       Set NR  ReadBW(KB/s)   MaxClat(us)    WriteBW(KB/s)  MaxClat(us)
---       --- --  ------------   -----------    -------------  -----------
brr       3   1   850            546345         0              0
bsr       3   1   14650          729543         0              0
bufw      3   1   0              0              23908          8274517

brr       3   2   981.333        579395         0              0
bsr       3   2   14149.7        1175689        0              0
bufw      3   2   0              0              21921          1.28108e+07

brr       3   4   898.333        1.75527e+06    0              0
bsr       3   4   12230.7        1.40072e+06    0              0
bufw      3   4   0              0              19722.3        2.4901e+07

brr       3   8   900            3160594        0              0
bsr       3   8   9282.33        1.91314e+06    0              0
bufw      3   8   0              0              18789.3        23890622

AVERAGE[brr, bsr, bufw mixed] [patched kernel]
-------
job       Set NR  ReadBW(KB/s)   MaxClat(us)    WriteBW(KB/s)  MaxClat(us)
---       --- --  ------------   -----------    -------------  -----------
brr       3   1   837            417973         0              0
bsr       3   1   14357.7        591275         0              0
bufw      3   1   0              0              24869.7        8910662

brr       3   2   1038.33        543434         0              0
bsr       3   2   13351.3        1205858        0              0
bufw      3   2   0              0              18626.3        13280370

brr       3   4   913            1.86861e+06    0              0
bsr       3   4   12652.3        1430974        0              0
bufw      3   4   0              0              15343.3        2.81305e+07

brr       3   8   890            2.92695e+06    0              0
bsr       3   8   9635.33        1.90244e+06    0              0
bufw      3   8   0              0              17200.3        24424392

So looks like it might make sense to include this patch.

Thanks
Vivek

Signed-off-by: Vivek Goyal 
Signed-off-by: Jens Axboe

blk-cgroup: Fix potential deadlock in blk-cgroup

2010-02-01T08:58:54+00:00

I triggered a lockdep warning as following.

=======================================================
[ INFO: possible circular locking dependency detected ]
2.6.33-rc2 #1
-------------------------------------------------------
test_io_control/7357 is trying to acquire lock:
 (blkio_list_lock){+.+...}, at: [] blkiocg_weight_write+0x82/0x9e

but task is already holding lock:
 (&(&blkcg->lock)->rlock){......}, at: [] blkiocg_weight_write+0x3b/0x9e

which lock already depends on the new lock.

the existing dependency chain (in reverse order) is:

-> #2 (&(&blkcg->lock)->rlock){......}:
       [] validate_chain+0x8bc/0xb9c
       [] __lock_acquire+0x723/0x789
       [] lock_acquire+0x90/0xa7
       [] _raw_spin_lock_irqsave+0x27/0x5a
       [] blkiocg_add_blkio_group+0x1a/0x6d
       [] cfq_get_queue+0x225/0x3de
       [] cfq_set_request+0x217/0x42d
       [] elv_set_request+0x17/0x26
       [] get_request+0x203/0x2c5
       [] get_request_wait+0x18/0x10e
       [] __make_request+0x2ba/0x375
       [] generic_make_request+0x28d/0x30f
       [] submit_bio+0x8a/0x8f
       [] submit_bh+0xf0/0x10f
       [] ll_rw_block+0xc0/0xf9
       [] ext3_find_entry+0x319/0x544 [ext3]
       [] ext3_lookup+0x2c/0xb9 [ext3]
       [] do_lookup+0xd3/0x172
       [] link_path_walk+0x5fb/0x95c
       [] path_walk+0x3c/0x81
       [] do_path_lookup+0x21/0x8a
       [] do_filp_open+0xf0/0x978
       [] open_exec+0x1b/0xb7
       [] do_execve+0xbb/0x266
       [] sys_execve+0x24/0x4a
       [] ptregs_execve+0x12/0x18

-> #1 (&(&q->__queue_lock)->rlock){..-.-.}:
       [] validate_chain+0x8bc/0xb9c
       [] __lock_acquire+0x723/0x789
       [] lock_acquire+0x90/0xa7
       [] _raw_spin_lock_irqsave+0x27/0x5a
       [] cfq_unlink_blkio_group+0x17/0x41
       [] blkiocg_destroy+0x72/0xc7
       [] cgroup_diput+0x4a/0xb2
       [] dentry_iput+0x93/0xb7
       [] d_kill+0x1c/0x36
       [] dput+0xf5/0xfe
       [] do_rmdir+0x95/0xbe
       [] sys_rmdir+0x10/0x12
       [] sysenter_do_call+0x12/0x32

-> #0 (blkio_list_lock){+.+...}:
       [] validate_chain+0x61c/0xb9c
       [] __lock_acquire+0x723/0x789
       [] lock_acquire+0x90/0xa7
       [] _raw_spin_lock+0x1e/0x4e
       [] blkiocg_weight_write+0x82/0x9e
       [] cgroup_file_write+0xc6/0x1c0
       [] vfs_write+0x8c/0x116
       [] sys_write+0x3b/0x60
       [] sysenter_do_call+0x12/0x32

other info that might help us debug this:

1 lock held by test_io_control/7357:
 #0:  (&(&blkcg->lock)->rlock){......}, at: [] blkiocg_weight_write+0x3b/0x9e
stack backtrace:
Pid: 7357, comm: test_io_control Not tainted 2.6.33-rc2 #1
Call Trace:
 [] print_circular_bug+0x91/0x9d
 [] validate_chain+0x61c/0xb9c
 [] __lock_acquire+0x723/0x789
 [] lock_acquire+0x90/0xa7
 [] ? blkiocg_weight_write+0x82/0x9e
 [] _raw_spin_lock+0x1e/0x4e
 [] ? blkiocg_weight_write+0x82/0x9e
 [] blkiocg_weight_write+0x82/0x9e
 [] cgroup_file_write+0xc6/0x1c0
 [] ? trace_hardirqs_off+0xb/0xd
 [] ? cpu_clock+0x2e/0x44
 [] ? security_file_permission+0xf/0x11
 [] ? rw_verify_area+0x8a/0xad
 [] ? cgroup_file_write+0x0/0x1c0
 [] vfs_write+0x8c/0x116
 [] sys_write+0x3b/0x60
 [] sysenter_do_call+0x12/0x32

To prevent deadlock, we should take locks as following sequence:

blkio_list_lock -> queue_lock ->  blkcg_lock.

The following patch should fix this bug.

Signed-off-by: Gui Jianfeng 
Signed-off-by: Jens Axboe

cfq-iosched: Respect ioprio_class when preempting

2010-01-11T15:16:18+00:00

In cfq_should_preempt(), we currently allow some cases where a non-RT request
can preempt an ongoing RT cfqq timeslice. This should not happen.
Examples include:

o A sync_noidle wl type non-RT request pre-empting a sync_noidle wl type cfqq
  on which we are idling.
o Once we have per-cgroup async queues, a non-RT sync request pre-empting a RT
  async cfqq.

Signed-off-by: Divyesh Shah
Signed-off-by: Jens Axboe

block: removed unused as_io_context

2010-01-11T13:29:20+00:00

It isn't used anymore, since AS was deleted.

Signed-off-by: Jens Axboe

block: bdev_stack_limits wrapper

2010-01-11T13:29:20+00:00

DM does not want to know about partition offsets.  Add a partition-aware
wrapper that DM can use when stacking block devices.

Signed-off-by: Martin K. Petersen 
Acked-by: Mike Snitzer 
Reviewed-by: Alasdair G Kergon 
Signed-off-by: Jens Axboe

block: Fix discard alignment calculation and printing

2010-01-11T13:29:19+00:00

Discard alignment reporting for partitions was incorrect.  Update to
match the algorithm used elsewhere.

The alignment can be negative (misaligned).  Fix format string
accordingly.

Signed-off-by: Martin K. Petersen 
Signed-off-by: Jens Axboe

block: Correct handling of bottom device misaligment

2010-01-11T13:29:19+00:00

The top device misalignment flag would not be set if the added bottom
device was already misaligned as opposed to causing a stacking failure.

Also massage the reporting so that an error is only returned if adding
the bottom device caused the misalignment.  I.e. don't return an error
if the top is already flagged as misaligned.

Signed-off-by: Martin K. Petersen 
Signed-off-by: Jens Axboe