We had good success with ScaleIO running on dedicated Centos 7 Bare Metal servers.
Environment is 4 nodes running Centos 7, dual 10Gb attached and local SSDs JBOD storage.
SDC client performance is great and systems are very stable using different workloads.
We have not been successful running SDS within VMware in hyper-converged fashion.
We take the exact same system as above, and re-install in this fashion
Environment is SAME 4 nodes this time running ESX 6.0u1a, dual 10Gb attached and local SSDs JBOD storage.
Hardware is the same, network is the same, we just replace the OS with ESX
Each node with LSI Controller Direct JBOD Mode (Controller cache is bypassed)
2 MDMs running as Centos 7 Guest OS on ESX local storage
1 TB running as Centos 7 Guest OS on ESX local storage
4 SDSs running as Centos 7 Guest OS on ESX local storage (we followed the SIO performance tuning guide)
8 SSDs dedicated to SIO in JBOD mode (2 per ESX server)
4 SDCs ESXi Kernel VIB driver mapping SIO Datastores from SSDs pool
Simple VM guest install on the SIO Datastores performance heavily impacted
Doing a simple linux DD bandwidth test on a guest VM using the SIO shared storage, the performance is not consistent some time results are somewhat ok 500MB/s, some times performance is very bad 20MB/s, yes we are purging the Linux FS cache before each run.
The system performance degrades tremendously and is inconsistent, when you do TOP or IOTOP you can see high wait time on the SDSs (yes we are using NOOP scheduler, all SIO datastores are thick-provisioned lazy zero).
It has been 3 days of trials...we are unable to pin-point where the problem is.
We use exactly the same hardware and config but running in Centos 7 bare metal, and performance is great!
It sounds like something related to how VMware addresses high throughput passing blocks through the kernel stack, we see high CPU wait time on the SDS VMs....
In contrast, running the same in Bare metal non-virtualized the CPU is hardly busy.
How long have you had this problem? have you been able to work it out or pin-point?
We been struggling with this problem for 4 days now, it all started when we imported VM templates into a freshly installed ESXi 6 4 x node cluster with ScaleIO datastores.
The copy/import performance is very slow 4GB VM took 4 to 5 minutes.
That trigger a set of additional tests, like installing vCenter Server 6 appliance directly on the ScaIeIO shared with all ESXi hosts datastore, the installation was never ever finished it had to be aborted.
Doing the same test on a local (not ScaleIO) SSD attached datastore completed successfully, performance was acceptable.
Yes you are correct, we understand about FIO and IOmeter, we used DD as a simple way to point sequential IO problems, the point is we ran DD multiple times on the ScaleIO datastores and the results were very inconsistent, not 10% or 20% delta but 500% to 600% delta between each run, that is obviously not good!
Lets talk about how the environment is architected in more detail:
ScaleIO version EMC-ScaleIO-XXX-1.32-2451.4.el7.x86_64
Node 1:
Xeon Dual 8 Core Hyper-Threaded 32GB RAM
LSI 3108 SAS 12Gb in JBOD Mode (cache is disabled)
2 x SAS SSDs JBOD (Dedicated for ScaleIO)
Intel X540 Dual Port 10Gb
ESXi 6.0u1a latest (SDC Driver instaled on ESXi kernel)
Inside ESXi Node 1 Guest VM Centos 7 latest for MDM01 8vCPU 8GB RAM VMXNET3 NIC
Inside ESXi Node 1 Guest VM Centos 7 latest for SDS01 8vCPU 8GB RAM VMXNET3 NIC (2 x SSDs JBOD as 2 x VMFS5 Datastores using Paravirtual Controller dedicated for ScaleIO)
1 x VMFS5 Datastore multi-mapped volume from ScaleIO SSD Pool to all 4 ESXi servers
Node 2:
Xeon Dual 8 Core Hyper-Threaded 32GB RAM
LSI 3108 SAS 12Gb in JBOD Mode (cache is disabled)
2 x SAS SSDs JBOD (Dedicated for ScaleIO)
Intel X540 Dual Port 10Gb
ESXi 6.0u1a latest (SDC Driver instaled on ESXi kernel)
Inside ESXi Node 2 Guest VM Centos 7 latest for MDM02 8vCPU 8GB RAM VMXNET3 NIC
Inside ESXi Node 2 Guest VM Centos 7 latest for SDS02 8vCPU 8GB RAM VMXNET3 NIC (2 x SSDs JBOD as 2 x VMFS5 Datastores using Paravirtual Controller dedicated for ScaleIO)
1 x VMFS5 Datastore multi-mapped volume from ScaleIO SSD Pool to all 4 ESXi servers
Node 3:
Xeon Dual 8 Core Hyper-Threaded 32GB RAM
LSI 3108 SAS 12Gb in JBOD Mode (cache is disabled)
2 x SAS SSDs JBOD (Dedicated for ScaleIO)
Intel X540 Dual Port 10Gb
ESXi 6.0u1a latest (SDC Driver instaled on ESXi kernel)
Inside ESXi Node 3 Guest VM Centos 7 latest for TB01 8vCPU 8GB RAM VMXNET3 NIC
Inside ESXi Node 3 Guest VM Centos 7 latest for SDS03 8vCPU 8GB RAM VMXNET3 NIC (2 x SSDs JBOD as 2 x VMFS5 Datastores using Paravirtual Controller dedicated for ScaleIO)
1 x VMFS5 Datastore multi-mapped volume from ScaleIO SSD Pool to all 4 ESXi servers
Node 4:
Xeon Dual 8 Core Hyper-Threaded 32GB RAM
LSI 3108 SAS 12Gb in JBOD Mode (cache is disabled)
2 x SAS SSDs JBOD (Dedicated for ScaleIO)
Intel X540 Dual Port 10Gb
ESXi 6.0u1a latest (SDC Driver instaled on ESXi kernel)
Inside ESXi Node 4 Guest VM Centos 7 latest for SDS04 8vCPU 8GB RAM VMXNET3 NIC(2 x SSDs JBOD as 2 x VMFS5 Datastores using Paravirtual Controller dedicated for ScaleIO)
1 x VMFS5 Datastore multi-mapped volume from ScaleIO SSD Pool to all 4 ESXi servers
Like we said before, when we ran simple copy test onto the shared ScaleIO datastore and performance is not stable.
We take the same configuration as above, and instead of installing ESXi bare metal we install Centos 7 bare metal, the performance is rock solid and stable, throughput is over 1.2GB/s on every test run.
We really want this ESXi ScaleIO configuration to work, as it will enable us to do additional benefits!
You can try set "VM Options > Advanced > Latency Sensitivity" to Hight (you need reserve all vm memory, as well this option will reserve cpu) for every sds vm.
I know, this is not the best solution, but we got 8Gb/s vm-to-vm speed after this.
PS. I'll be happy to find another ways to get good vm-to-vm speed in vmware...
Are there anybody? I have more news on this theme.
We have ScaleIO cluster on bare metal Centos7 (3 nodes with 1 SSD disk each), network - 1Gbit
On vmware vm with centos7 I setup two disks:
- the first is using scini sdc (native linux client)
- the second is using vmware sdc, then create VMFS datastore and place virtual disk there (I tried RDM this device without VMFS - got the same results)
After that I run fio: fio --name=testfile --readwrite=randread --time_based --runtime=10 --direct=1 --numjobs=1 --size=256M --bs=4k
Results:
- First variant 1800 IOPS
- Second variant 70 IOPS
As we see, linux sdc is 20 times faster than vmware sdc
PS. All recommendation on SDC from fine-tuning-scaleio-performance guide applied (except jumbo frames, but it is the same for both drives).
Just one thing guy, you are just trying sequential testing for what i saw.
When you do a sequential stuff on a bare metal host it will be sequential, but if you do a sequential stuff in a VM, vmware will transform the load into random patern.
So you can never match the same performance in a VM than a bare metal server and at the end you can't really compare both of them.
What you can compare is running a VM on KVM into a centos bare metal vs a Vm in VMWARE.
And by the way check the perf mon of scaleIO when you are doing the test.
bogdansmc
12 Posts
2521
0
Posted December 9th, 2015 05:00
We have the same problem with vmware installation.
And if we add direct mode for dd command IOPS are:
dd if=/dev/zero of=/mnt/ssd/t1 bs=4k oflag=direct
virtual disk: busy 97%, write iops 1200, write 5 MB/s
dd if=/dev/zero of=/mnt/ssd/t1 bs=64k oflag=direct
virtual disk: busy 97%, write iops 600, write 40 MB/s
dd if=/dev/zero of=/mnt/ssd/t1 bs=512k oflag=direct
virtual disk: busy 98%, write iops 150, write 70 MB/s
That's maximum we can get on 3 SSDs and 10 Gbit network