PowerFlex: 모든 경로가 종료된 상태에서 vCenter에서 ESXi SDC가 응답을 중지함

요약: 하나 이상의 PowerFlex 볼륨에서 APD(All Paths Down) 상태로 인해 ESXi server가 vCenter에서 응답을 중지합니다.

이 문서는 다음에 적용됩니다. 이 문서는 다음에 적용되지 않습니다. 이 문서는 특정 제품과 관련이 없습니다. 모든 제품 버전이 이 문서에 나와 있는 것은 아닙니다.

증상

PowerFlex 볼륨에서 ESXi SDC에 지속적으로 I/O 오류가 발생하면 하나 이상의 PowerFlex 볼륨에 대해 APD(All-Path-Down) 상태가 될 수 있습니다. 이 상태로 인해 vCenter에서 응답이 중지될 수 있습니다.

일반적 으로:

  • 일부 ESXi 호스트가 vSphere Client에서 연결 끊김으로 표시됨
  • I/O 오류 누적: vmkernel.log파일로 교체합니다.
2018-01-10T22:30:08.321Z cpu29:33684)ScsiDeviceIO: 2651: Cmd(0x439e41930500) 0x28, CmdSN 0x8819c3 from world 34407 to dev "eui.<mdmId+volId>" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0.
  • hostd.log APD 상태 시작 시 다음 오류가 포함될 수 있습니다.
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10313 : Lost connectivity to storage device eui.<mdmId+volId>. Path vmhba64:C0:T27:L91 is down. Affected datastores: <Datastore Name>.
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 778923875,
-->    chainId = 1635216758,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "PHSVCESQL1018.partners.org",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.start",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10314 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down state.
  • hostd.log 메시지가 있습니다 esx.problem.storage.apd.timeout 때 hostd 서비스가 응답하지 않습니다.
2017-10-24T17:06:58.277Z info hostd[29A40B70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 690973144,
-->    chainId = 1635216641,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "ESXi.host.local",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.timeout",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       },
-->       (vmodl.KeyAnyValue) {
-->          key = "2",
-->          value = "140"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:58.278Z info hostd[29A40B70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10336 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down Timeout state after being in the All Paths Down state for 140 seconds. I/Os will now be fast failed.
  • 의 스니펫 vmkwarning 위와 일치합니다. hostd파일로 교체합니다.
커널이 제거를 시도하는 것을 볼 수 있습니다. vmhba64:C0:T30:L62하지만 때문에 할 수 없습니다. hostd 재검색 및 응답 중지 상태 중에 사용 중 상태로 유지:
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: NMP: nmpUnclaimPath:1516: NMP device "eui.<mdmId+volId>" quiesce state change failed: Busy
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4507: Path vmhba64:C0:T30:L62 is being removed
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4737: Failed to issue command 0x0 (cmdSN 0x0) on path vmhba64:C0:T30:L62: No connection
2017-10-24T17:04:38.268Z cpu8:33147)WARNING: ScsiScan: 2007: Could not delete path vmhba64:C0:T30:L62
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_IssueCommandToDevice:4553: I/O could not be issued to device "eui.<mdmId+volId>" due to Not found
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceRetryCommand:133: Device "eui.<mdmId+volId>": awaiting fast path state update for failover with I/O blocked. No prior reservation exists on the device.
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceStartLoop:725: NMP Device "eui.<mdmId+volId>" is blocked. Not starting I/O from device.
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:603: Retry world failover device "eui.<mdmId+volId>" - issuing command 0x43a6402bbac0
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:678: Retry world failover device "eui.<mdmId+volId>" - failed to issue command due to Not found (APD), try again...
  • 에서 볼 수 있듯이 다음이 발생하고 있습니다. storagerm 로그: 
2017-10-24T17:05:49.274Z: Write 0xffcda788[512] -> 68 failed. 38:Function not implemented, offset=0, bufLen=512
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> Write error to fd 68, error: Function not implemented
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> I/Os from datastore eui.207d160928aa82202102c97700000060 took 62.962148(>= 30.000000) seconds to complete stats computation. Reducing its polling frequency.
2017-10-24T17:06:58.277Z: Write 0xffcda788[512] -> 58 failed. 6:No such device or address, offset=0, bufLen=512
2017-10-24T17:07:04.484Z: <Datastore Name, 0> Some host is down, need to reset the slot allocation
2017-10-24T17:07:08.554Z: Skipping device eui.207d160928aa82202102c9530000003e either due to VSI read error or abnormal state
2017-10-24T17:07:08.580Z: open /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile(0x202, 0x0) failed: Input/output error
2017-10-24T17:07:08.580Z: Input/output error Error -1 opening/truncating file /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile
  • VM이 다음과 같이 표시될 수 있습니다. /vmfs/volumes/.../...vmx 표시 이름 대신 files를 사용합니다.
  • VPXA 및 vCenter와의 연결 끊김으로 인해 DVS 포트에 장애가 발생할 수 있습니다.
2017-10-24T17:06:55.704Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-14505 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/14505
2017-10-24T17:06:55.943Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-8339 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/8339
2017-10-24T17:06:55.994Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-15520 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/15520
2017-10-24T17:06:56.017Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-7684 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/7684
  • vCenter에 연결된 VM 및 호스트의 연결 끊김

또한 가능:

  • ESXi 또는 VM에 대한 SSH 연결을 설정할 수 없음(관리 네트워크가 Distributed vSwitch에 있는 경우)
  • 사용할 수 없습니다 esxcli 콘솔 세션에서. (응답이 중지됩니다. localcli 옵션을 지정합니다. 해결 방법 섹션을 참조하십시오.)
  • SVM을 포함한 VM은 프로세스를 종료하지 않고는 전원을 켜거나 끌 수 없습니다. 
  • 재부팅 또는 부팅 시 ESXi 호스트가 응답을 중지할 수 있습니다.
  • 부팅 프로세스는 일반적으로 반드시 그런 것은 아니지만 부팅 프로세스 후 응답을 멈춥니다. nfs41client 모듈이 로드되었습니다. 호스트의 콘솔(DCUI)에 다음 메시지가 표시됩니다.
nfs41client loaded successfully

영향

  • vCenter를 통해 ESXi 호스트를 관리하거나 SSH 연결을 설정할 수 없습니다.
  • vMotion 기능 없음

원인

APD 상태에서 ESXi 사용자(hostd agent) 또는 게스트 OS의 시간 초과로 인해 중단되지 않은 게스트 OS의 모든 I/O가 무기한 재시도되어 시스템 리소스가 고갈되고 vCenter에서 ESXi가 응답하지 않는 상태가 됩니다.

해결

  • 호스트가 다시 APD로 전환될 수 있으므로 기본 APD 조건을 수정하지 않고 ESXi 호스트를 재부팅하면 도움이 되지 않습니다.
  • APD가 발생한 ESXi 호스트에서 명령을 실행해야 하는 경우 "localcli" 대신 "esxcli," 후자가 응답을 멈춥니다.
예:
  • 다음을 사용하여 데이터 저장소가 마운트된 것으로 표시되는지 확인합니다.
[root@92U-16:~] localcli storage filesystem list
Mount Point                                        Volume Name  UUID                                 Mounted  Type    Size           Free
-----------------------------------------------------------------------------------------------------------------------------------------
/vmfs/volumes/5975cf1e-9306f9bc-0dbc-a0369fdaccbc  SATADOM17    5975cf1e-9306f9bc-0dbc-a0369fdaccbc  true     VMFS-5    55834574848   53979643904
/vmfs/volumes/59916bcd-22a730ae-db91-a0369fdaccbc  LocalDS17    59916bcd-22a730ae-db91-a0369fdaccbc  true     VMFS-6  1920118816768  986341965824
/vmfs/volumes/5975cf15-c44cea1b-de13-a0369fdaccbc               5975cf15-c44cea1b-de13-a0369fdaccbc  true     vfat        299712512      83927040
/vmfs/volumes/16a83277-c690cda2-9723-26fe2e41d0c3               16a83277-c690cda2-9723-26fe2e41d0c3  true     vfat        261853184      97923072
/vmfs/volumes/5975cf1f-17e61cfc-a0ae-a0369fdaccbc               5975cf1f-17e61cfc-a0ae-a0369fdaccbc  true     vfat       4293591040    4260626432
/vmfs/volumes/79e9c87d-f55f1864-b3ce-6e24607afc68               79e9c87d-f55f1864-b3ce-6e24607afc68  true     vfat        261853184      99840000
  • 호스트 수준에서 재검색을 시도하려면 다음을 사용합니다.
localcli storage filesystem rescan
  • 이미 APD 상태인 ESXi 호스트를 재부팅해야 하는 경우 매핑된 볼륨을 기록해 두고 일시적으로 매핑을 해제합니다. 문제가 해결된 후 호스트에 다시 매핑합니다.
 
참고: ESXi 호스트가 여러 MDM 또는 PowerFlex 시스템에 연결되어 있는 경우 영향을 받는 시스템의 볼륨만 매핑 해제해야 합니다.
 
  • 만약 unmap_volume 복구 중에 작업이 필요하며, 볼륨을 다시 매핑하고 데이터 저장소를 다시 마운트한 후 일부 VM을 다시 등록해야 할 수 있습니다.
솔루션
버전 2.0.1.3에서는 기본적으로 비활성화된 PDL( Permanent Device Loss) 기능이 도입되었습니다. 이 기능을 활성화하면 SDC가 60초 후에 볼륨에 I/O를 보낼 수 없을 때 APD를 PDL로 전환할 수 있습니다. 이 시간 초과 값은 일부 환경에서 영향을 확인하지 않고 견딜 수 있는 것보다 더 길 수 있으며 추가 조정이 필요할 수 있습니다.

추가 정보

해당 제품

PowerFlex rack, ScaleIO
문서 속성
문서 번호: 000437810
문서 유형: Solution
마지막 수정 시간: 27 3월 2026
버전:  3
다른 Dell 사용자에게 질문에 대한 답변 찾기
지원 서비스
디바이스에 지원 서비스가 적용되는지 확인하십시오.