UNSOLVED

vearcut

updated

10 years ago

V

vearcut

16 Posts

1

4391

March 14th, 2017 09:00

ScaleIO host with IO Error

ScaleIO deployed in VMware environment. On one host operation with datastore (migrate VM, create vm and other) a long time.

All physical ssd devices working are fine. ScaleIO system without errors.

Smart from virtual device:

eui.7b850322485d5c3051e837c300000000

   Device: eui.7b850322485d5c3051e837c300000000

   Successful Commands: 82256182

   Blocks Read: 1484018027

   Blocks Written: 4402672521

   Read Operations: 12992279

   Write Operations: 68603358

   Reserve Operations: 0

   Reservation Conflicts: 0

   Failed Commands: 1902

   Failed Blocks Read: 0

   Failed Blocks Written: 790144

   Failed Read Operations: 0

   Failed Write Operations: 1376

   Failed Reserve Operations: 0

VMkernel.log:

2017-03-14T13:59:42.752Z cpu2:858185)J3: 3302: Aborting txn (0x43070883c2b0) callerID: 0xc1d00006 due to failure pre-committing: I/O error

2017-03-14T14:00:02.871Z cpu25:33353)WARNING: [972927942] IO-ERROR comb: 2da880000381. offsetInComb 2324480. SizeInLB 2048. SDS_ID 9a2ad19900000002. Comb Gen 28. Head Gen 5374.

2017-03-14T14:00:02.871Z cpu25:33353)WARNING: Vol ID 0x51e837c300000000. Last fault Status IO_HARD_ERROR(20).Last error Status NOT_CONN(4) Reason (ERROR) Retry count (20) chan (1)

2017-03-14T14:00:02.871Z cpu25:33353)scini: blkScsi_PrintIOInfo:3274: ScaleIO R2_01:hCmd 0x439d98caaa40, OpCode 0x93, rc 1 scsiStat 2, senseCode 4, asc 0, ascq 0

2017-03-14T14:00:02.871Z cpu7:33507)ScsiDeviceIO: 2651: Cmd(0x439d98caaa40) 0x93, CmdSN 0x7afd7d from world 858816 to dev "eui.7b850322485d5c3051e837c300000000" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0.

2017-03-14T14:00:04.935Z cpu1:33406)NMP: nmp_ResetDeviceLogThrottling:3349: last error status from device eui.7b850322485d5c3051e837c300000000 repeated 2 times

2017-03-14T14:00:23.064Z cpu16:33353)WARNING: [972948135] IO-ERROR comb: 2da880000381. offsetInComb 2324480. SizeInLB 2048. SDS_ID 9a2ad19900000002. Comb Gen 28. Head Gen 5374.

2017-03-14T14:00:23.064Z cpu16:33353)WARNING: Vol ID 0x51e837c300000000. Last fault Status IO_HARD_ERROR(20).Last error Status NOT_CONN(4) Reason (ERROR) Retry count (20) chan (1)

2017-03-14T14:00:23.064Z cpu16:33353)scini: blkScsi_PrintIOInfo:3274: ScaleIO R2_01:hCmd 0x439d98caaa40, OpCode 0x2a, rc 1 scsiStat 2, senseCode 4, asc 0, ascq 0

2017-03-14T14:00:23.064Z cpu3:33507)NMP: nmp_ThrottleLogForDevice:3298: Cmd 0x2a (0x439d98caaa40, 858816) to dev "eui.7b850322485d5c3051e837c300000000" on path "vmhba64:C0:T0:L0" Failed: H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0. Act:NONE

2017-03-14T14:00:23.064Z cpu3:33507)ScsiDeviceIO: 2651: Cmd(0x439d98caaa40) 0x2a, CmdSN 0x7afd9b from world 858816 to dev "eui.7b850322485d5c3051e837c300000000" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0.

2017-03-14T14:00:23.064Z cpu8:858816)FS3DM: 2767: status I/O error zeroing 1 extents (1048576 each)

How to fix this problem or how to localize this problem?

Full log: http://pastebin.com/raw/tGqVghZq

After start backup proccess with Veeam Backup & Replication i see this error. This problem appears only on one node, the rest are all working well.