linux-toradex.git/drivers/md/raid1.c, branch T30_LinuxImageV2.0Beta2_20130626

md/raid1: perform bad-block tests for WriteMostly devices too.

2012-01-18T15:31:57+00:00

commit 307729c8bc5b5a41361af8af95906eee7552acb1 upstream.

We normally try to avoid reading from write-mostly devices, but when
we do we really have to check for bad blocks and be sure not to
try reading them.

With the current code, best_good_sectors might not get set and that
causes zero-length read requests to be send down which is very
confusing.

This bug was introduced in commit d2eb35acfdccbe2 and so the patch
is suitable for 3.1.x and 3.2.x

Reported-and-tested-by: Michał Mirosław 
Reported-and-tested-by: Art -kwaak- van Breemen 
Signed-off-by: NeilBrown 
Signed-off-by: Greg Kroah-Hartman

md: Avoid waking up a thread after it has been freed.

2011-09-21T05:30:20+00:00

Two related problems:

1/ some error paths call "md_unregister_thread(mddev->thread)"
   without subsequently clearing ->thread.  A subsequent call
   to mddev_unlock will try to wake the thread, and crash.

2/ Most calls to md_wakeup_thread are protected against the thread
   disappeared either by:
      - holding the ->mutex
      - having an active request, so something else must be keeping
        the array active.
   However mddev_unlock calls md_wakeup_thread after dropping the
   mutex and without any certainty of an active request, so the
   ->thread could theoretically disappear.
   So we need a spinlock to provide some protections.

So change md_unregister_thread to take a pointer to the thread
pointer, and ensure that it always does the required locking, and
clears the pointer properly.

Reported-by: "Moshe Melnikov" 
Signed-off-by: NeilBrown 
cc: stable@kernel.org

md/raid1,10: Remove use-after-free bug in make_request.

2011-09-10T07:21:23+00:00

A single request to RAID1 or RAID10 might result in multiple
requests if there are known bad blocks that need to be avoided.

To detect if we need to submit another write request we test:
 	if (sectors_handled < (bio->bi_size >> 9)) {

However this is after we call **_write_done() so the 'bio' no longer
belongs to us - the writes could have completed and the bio freed.

So move the **_write_done call until after the test against
bio->bi_size.

This addresses https://bugzilla.kernel.org/show_bug.cgi?id=41862

Reported-by: Bruno Wolff III 
Tested-by: Bruno Wolff III 
Signed-off-by: NeilBrown

md/raid1: factor several functions out or raid1d()

2011-07-28T01:38:13+00:00

raid1d is too big with several deep branches.
So separate them out into their own functions.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: improve handling of read failure during recovery.

2011-07-28T01:33:42+00:00

If we cannot read a block from anywhere during recovery, there is
now a better approach than just giving up.
We can record a bad block on each device and keep going - being
careful not to clear the bad block when a write succeeds as it might -
it will be a write of incorrect data.

We have now reached the state where - for raid1 - we only call
md_error if md_set_badblocks has failed.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: record badblocks found during resync etc.

2011-07-28T01:33:00+00:00

If we find a bad block while writing as part of resync/recovery we
need to report that back to raid1d which must record the bad block,
or fail the device.

Similarly when fixing a read error, a further error should just
record a bad block if possible rather than failing the device.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: Handle write errors by updating badblock log.

2011-07-28T01:32:41+00:00

When we get a write error (in the data area, not in metadata),
update the badblock log rather than failing the whole device.

As the write may well be many blocks, we trying writing each
block individually and only log the ones which fail.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: store behind-write pages in bi_vecs.

2011-07-28T01:32:10+00:00

When performing write-behind we allocate pages to store the data
during write.
Previously we just keep a list of pages.  Now we keep a list of
bi_vec which includes offset and size.
This means that the r1bio has complete information to create a new
bio which will be needed for retrying after write errors.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: clear bad-block record when write succeeds.

2011-07-28T01:31:49+00:00

If we succeed in writing to a block that was recorded as
being bad, we clear the bad-block record.

This requires some delayed handling as the bad-block-list update has
to happen in process-context.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim

md/raid1: avoid writing to known-bad blocks on known-bad drives.

2011-07-28T01:31:48+00:00

If we have seen any write error on a drive, then don't write to
any known-bad blocks on that drive.
If necessary, we divide the write request up into pieces just
like we do for reads, so each piece is either all written or
all not written to any given drive.

Signed-off-by: NeilBrown 
Reviewed-by: Namhyung Kim