This post is more than 5 years old

31 Posts

28071

September 14th, 2009 08:00

Array failure with new drives on CERC 1.5/6ch

Hi all. We have a Dell PE 1800 with a CERC 1.5/6ch  and 3 80GB SATA drives that has been in production for 3 years without problems.

We have some software upgrades to do and we need more space. 

 I removed the 3 80GB drives, and installed 3 500 GB Hitachi HDP725050GLA360 drives purchased from Dell.  I created a new RAID 5 array and restored a Ghost image of C: and it worked fine. While restoring D: (120 gb) after about 1.5 hours the controller alarms and  shows an array failure. All disks show online with no errors.  Nothing in the log to say what kind of failure, and all data is now inaccessable. I tried this a couple of times, and updated the CERC firmware to 4.1.0.7419 and still have the same issue.   I'm a bit limited as to what I can try, as the site is a 1.8 hour drive away. and it takes 2 hours for it to fail.

Any ideas what would be causing this?

 

2 Posts

April 12th, 2010 16:00

Hey MetroCS, thanks for writing your experiences.

 

I've pretty much given up on my drives and have moved them to workstation use.  Here is my theory on why these drives are failing in a RAID setup:

 

Hard drives have a feature that controls the length of timeout for raid controllers.  This is used for error handling.  Western Digital drives call it TLER, Hitachi drives call it CCTL, and samsung call it ERC.

Here's an wikipedia article that sums it up:

 http://en.wikipedia.org/wiki/Time-Limited_Error_Recovery

 

As for making these drives work:

I believe that the CCTL on these drives are set too either too low or too high, and basically the raid controller thinks that this drive has failed.  Since RAID controllers have much more tighter tolerances, these drives meant for desktop use just doesn't work well in a RAID configuration.

There are utilities to modify the TLER for western digital drives but I have not found anything for the Hitachi drives.  I've seen some generic tools, as these fuctions are supposed to be standard SMART functions and are able to be controlled by standard ATA commands, however, making these changes permanent, will require a special tool.

I've pretty much deemed this as too much work involved and have stopped trying.  Cheaper to buy proper working drives, such as the Western Digital RE drives than it is to work out a solution for these HITACHI drives, I dont like Hitachi drives anyways, their rate of failure has been pretty high in my experience.  Dont know if you know of the IBM click of death, but the HITACHI's are old IBM stuff.

 

Good Luck.

6 Operator

 • 

9.3K Posts

 • 

3 Points

September 14th, 2009 10:00

Assuming you're under warranty, I'd suggest to open a case with Dell Poweredge support and see if they have any ideas.

 

The only thing that comes to mind is to install another drive, connect it to a different controller (than the CERC), make this drive the boot device, recover the C-drive image to that, and then start the recovery of the D-drive to the CERC-based raid 5. This isn't for final production, but to see if after the raid 5 fails you can pull a DSET (support.dell.com/dset) as the OS isn't on the raid 5 that failed. Then work with support to check if the DSET report can shed any light on the problem.

31 Posts

September 14th, 2009 12:00

Good suggestion, but more work than I have time for. Plan B, I put back the original 80GB array, and all is well with it. I put 2 of the 3 500s in as a RAID 1 array, and will juggle data to it. So far it's failed building 3 times.

I'm formatting each drive individially to see if either one fails a format. I had tested all 3 drives running Ontrack full diags, and they all passed.

31 Posts

September 15th, 2009 12:00

Each drive formatted fine, but when I created a RAID 1 array it failed after about 25% with a "redundancy failure"

Out of time, so will use the drives seperately to offload some non-critical data, and get back to it later.

 

2 Posts

October 6th, 2009 00:00

Hi,

 

I'm having very similar problems.  Picked up some exact Hitachi 500GB drives (HDP7250GLA360) did you add an extra (50) in there?, I can't even get them to format, Verification Disk Media Fails as well.  I've checked them using Hitachi's drive diagnostics on another machine and they are fine.  Disk Verification works fine on other Seagate 75GB drives on the controller, tried changing the controller channels as well and still fails.  I'm leaning towards an incompatibility problem here with these Hitachi drives.  These came out of a warranty repair on a 1TB Lacie External drive enclosure,  Unfortunately the server is already out of warranty and I wont be able to get support from Dell. Also on the latest firmware.  They've been zero'd out using the Hitachi disk utility as well.  Anyone have an idea?

3 Posts

April 8th, 2010 06:00

QUARKNUX,

Did you ever get the Hitachi's working?  I been trying to get a RAID5 array up using three of these drives (brand new).  After many hours during the resyncing/scrub the virtual disk fails.  From reviewing the log, I suspected drive 5, replaced it with a new one same EXACT failure.  Drive 3 & 5 go offline Drive 4 shows online.  I have updated the firmware with the latest.

Likely suspects:

1. Drive 5 cable
2. Drive 3
3. Drive 3 cable
4. Controller
5. Incompatible Drives w/ controller

I can create a RAID 0 vitrual disk with them... but need redundancy

Here is a snippet of the log:

[1704]: Periodic Time Display: Wed - Apr  7 21:43:49 2010 (LowIO=0, RawBLK=0, RawIO=0)
[1705]: Periodic Time Display: Wed - Apr  7 21:58:50 2010 (LowIO=0, RawBLK=0, RawIO=0)
[1706]: Periodic Time Display: Wed - Apr  7 22:13:51 2010 (LowIO=0, RawBLK=0, RawIO=0)
[1707]: 
[1708]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1709]: 
[1710]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1711]: 
[1712]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1713]: 
[1714]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1715]: 
[1716]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1717]: 
[1718]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1719]: 
[1720]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1721]: 
[1722]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1723]: 
[1724]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1725]: 
[1726]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1727]: 
[1728]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1729]: 
[1730]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1731]: 
[1732]: HDMhpt370: id 3 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1733]: 
[1734]: HDMhpt370: id 5 lun 0 Received parity errorMpdName => Iob->taskStatus: HIM_IOB_PARITY_ERROR.
[1735]: (ID(0:03:0) Cmd[0x28] Fail: Block Range 268435456:268435583
[1736]: FSAPRINT: ID(0:03:0) - Cmd[0x28] failed (retries exhausted)
[1737]:            - Srb #c0d91024, SrbStat=0xf, SCSIStat=0x0, Flags=0x111b3
[1738]: (ID(0:05:0) Cmd[0x28] Fail: Block Range 268435456:268435583
[1739]: FSAPRINT: ID(0:05:0) - Cmd[0x28] failed (retries exhausted)
[1740]:            - Srb #c0ce7114, SrbStat=0xf, SCSIStat=0x0, Flags=0x111b3
[1741]: biowait: ERROR bp=0xC0D90E70
[1742]: biowait: ERROR bp=0xC0CE6F60
[1743]: ClearDiskLog: sector - 80, driveno 3
[1744]: ClearDiskLog: sector - 80, driveno 4
[1745]: ClearDiskLog: sector - 80, driveno 5
[1746]: FSAPRINT: !Container 1 failed SCRUB task: I/O error - drive 0:5:0 failed 
[1747]: CT_TerminateTask: CTArray[1] Match
[1748]: CT_TerminateTask: Free Memory=0xC0889CD0
[1749]: CT_TerminateTask: Terminate Task = R5Scrub_1
[1750]: FSAPRINT: !RAID5 Container 1 Drive 0:5:0 Failure 
[1751]: expevent 00020006 - Container 1 state change to 2
[1752]: CT_LogMissingEntry: Log missing entry, container 1, dev 5, signature 0x004DE1DA, nvEntry 115
[1753]: CtMarkDead: container 1, deadEntry 2, dev 5, signature 0x004DE1DA
[1754]: expevent 00020006 - Container 1 state change to 1
[1755]: expevent 00020006 - Container 1 state change to 2
[1756]: CtMarkDead: container 1, deadEntry 2, dev 5, signature 0x004DE1DA
[1757]: expevent 00020006 - Container 1 state change to 2
[1758]: CtMarkDead: container 1, deadEntry 2, dev 5, signature 0x004DE1DA
[1759]: ReadSliceMBR: can't read mbr dev_t:5
[1760]: ReadSliceMBR: can't read mbr dev_t:5
[1761]: ReadSliceMBR: can't read mbr dev_t:5

3 Posts

April 13th, 2010 07:00

Quarknux,

Thanks for responding and providing your insight.

Have you had success in creating a RAID5 1TB VD with other manufactures drives?

I did have success in creating a RAID5 250GB VD using the 500GB Hitachi's but  I am not sure that I want 4 VD's

I understand and do remember the IBM DeathStar or click of death, but to be fair I have seen all the manufactures have some models that are less then reliable

Metrocs

31 Posts

May 18th, 2010 10:00

What a pain. I assumed (big mistake) that since I purchased the drives from Dell, for this server, that they'd be "enterprise" drives, so hadn't thought about this.

This server is due for replacement this summer, but again we have a bit of a disk crunch, so decided to drop in a pair of WD drives, and set TLER to 7.

Bitten in the backside again though, the latest WD Caviar Black 1TB drives no longer support the TLER utility.

1 Message

September 14th, 2010 15:00

Hi,

 

For the record, I have a server running centos 5.4 with 4 x 500gb Seagate drives.

The drives are configured in raid-1 configuration and work fine, have not tried raid5, don't think i would trust this controller with a raid5.

The have a large cache, ie 32mb so dont know if that is helping.

It may be worthwhile running a fw update on the drives to see if that solves anything.

Hope this is of use to someone.

Regards

Joe.

October 18th, 2010 15:00

Count me in the list of "hurt by Hitachi" on this RAID controller. Same situation, except I went 4 drives and RAID 10. About an hour into the 150GB restore, 3 of the 4 drives failed.

The 3 failures? Hitachi 500GB.

The 1 success? Seagate Barracuda!

I wanted to mix up the drive brands just in case we got a bad lot. So much for that idea! 

The Hitachi Deskstar drives seems to work fine solo, or single volumes. But RAID them with this controller, and it's just a matter of time... boom.

No Events found!

Top