This post is more than 5 years old

2620

April 12th, 2018 08:00

T110 ii with H200 Abysmal Performance and Problems, need advice.

Hi Everyone, its my first post. go easy on me. Im sorry in advance for the long story but i think you need all the details to render an opinion on the topic. Thanks for reading...

I have a poweredge T110 ii with H200 RAID10. 
It shipped with 4x1TB but now has 4 x 4TB drives with a horrible track record but im not sure what to do to fix this. 

Server is 3.5 years old:

In year 1 & 2: I bought the server with 4 x 1TB drives but quickly outgrew the 2TB of hard drive space with our SQL software. after about a year we changed the drives to a larger size and thats when the problems began. I ordered 4 new 2TB drives to recreate the RAID10, did a full backup, created a new raid10 with the new drives and restored from backup. All good.

After a few months we were having slowness issues again, and when i went to look at the four drives, drive 0 and 2 were fine, but drives 1 & 3 had some issue causing them to be degraded. I didn't want to take any chances so i bought two new drives and installed them on those channels 1 & 3, the RAID10 rebuilt, but it took about a week. Yes a Week. During that time the network performance was really slow. 

After the sync was done, within a few days, now Drives 0 and 2 now showed a degraded status. so i did the same thing, replaced them with new drives. Another WEEK to sync. Now that i had all new drives in the server my stress level dropped. All good for the next year. I didn't have confidence in the H200 card that i had four bad hard drives. I have since used all of them in other computers and they all seem to work without issues even a couple of years later. 

The following year, Once again we were running out of space because the company was really growing at that time so i jumped up to 4 x 4TB to upgrade my Raid10 again and give us more storage. Growth has leveled off since. A couple of years later and we are still using about 5TB out of the 8TB of space. I dd the same thing; full backup, install new drives, recreate the RAID10 and restore. All good.

However, since i replaced the four drives, the server had been giving us an issue that if we didn't reboot it at least once every 7 days it would become completely unresponsive. The ONLY solution is a hard shutdown with the power button. Of course a hard shutdown is damaging and would trigger a resync of the raid10 but eventually things would normalize. 

Once i realized what pattern was happening i scheduled a reboot every few days to avoid this until i had time to get back to looking into it. OMSA showed no issues and booting into the BIOS showed no issues but i could never explain why this server would freeze up like this. Every once in a while the scheduled reboot would not occur and the server would totally lock up, so we would have to do a hard shut down. 

Two weeks ago: the scheduled reboot did not happen for some reason so the server locked up. I had no choice but to hard shutdown. However, after this, Drives 0 and 2 read fine, but 1 & 3 state as degraded. The first set, 0&1 resynced. the second set 2&3 did not. After a week of trying to rebuild drive 3 read as "REMOVED."

I bought two new 4TB drives. installed them in slots 1 & 3 and booted up. The H200 BIOS immediately showed that the first raid set (drives 0&1) is rebuilding/resyncing and the second set (2&3) still showed as DEGRADED and not as resyncing like the first set. I could not understand why. However I read somewhere that the rebuilding only happens in Windows, so i booted into windows and opened OMSA. Here it shows as rebuilding/resyncing/reconstructing on both raid sets. I should be relieved, right?

Not exactly, When i open up the detailed view, only one shows as having any progress. the other one says "not applicable." 

I am going to try to attach a screenshot.

Problem 1:  why the inconsistent status?

Problem 2: after 24 hours i am only at 2% rebuild on the OS drive and the IMPORTANT drive, where all my data is does not show ANY progress at all. Is it waiting for the first one to complete? I cannot find any documentation on this. Someone with knowlege and experience please let me know. The network has been basically crawling for two weeks and after telling my client that new drives will fix it, now im not so sure. 

Problem 3: Should I expect that the first two drives are going to fail when the sync completes, like it did previously? By the way the first four drives i replaced 2 years ago have continued to work flawlessly in other systems since i pulled them out of this server. So my confidence in this raid card is not great.

Lastly, I am considering buying a new server. just out of curiosity what would be recommended:
I need an OS drive for Server2016 and a Data drive for SQL software. 8TB of total storage. Is it best to do a single RAID10 and then put both partitions on the RAID10 and split the space between them?  for example 1TB for OS and 7TB or remaining space for the DATA drive? or will someone recommend something better for performance yet with redundancy? We cannot spend more than $4K. Thanks

Any light that can be shed on this topic will be greatly appreciated. 

Capture1.JPGCapture3.JPG

12 Elder

 • 

6.2K Posts

April 13th, 2018 10:00

Thank you for the service tag.

It is my understanding that you experienced no issues with the configuration that you purchased from us. You then changed the hard drives multiple times and have experienced issues with all of those drives. I show the current drives are not certified, I am assuming that all of the drives you have experienced issues with were not certified.

The firmware on our PERCs is designed to be used in conjunction with drives that are using our firmware. When we validate our firmware updates we do not validate disks that are not using our firmware. Because of this our PERCs do not work well with non-Dell drives.

I agree with your assumption that the issue is with the controller. I think it is a communication issue between the controller and the 3rd party drives, but I don't think anything is wrong with the controller.

Thanks

12 Elder

 • 

6.2K Posts

April 12th, 2018 10:00

Hello

Please send a private message with your service tag to ensure we have all appropriate information on your system.

Thanks

April 12th, 2018 17:00

Hi Daniel,

I sent you the requested information. Thanks for taking the time, its much appreciated!

April 13th, 2018 10:00

Hi Daniel,

I appreciate the time you took to look into it.

Swapping out Dell drives for non-Dell drives has worked for me in the past with several other Dell servers. The controller may complain by putting an occasional message on the screen, but it doesn't misbehave. I've also seen numerous posts on the subject. However, I guess I have my answer.

Thank you.

No Events found!

Top