PowerFlex: El SDC de ESXi deja de responder en vCenter con todas las rutas inactivas

Resumen: Los servidores ESXi dejan de responder en vCenter debido al estado All Paths Down (APD) en uno o más de los volúmenes PowerFlex.

Este artículo se aplica a Este artículo no se aplica a Este artículo no está vinculado a ningún producto específico. No se identifican todas las versiones del producto en este artículo.

Síntomas

Cuando un SDC de ESXi experimenta continuamente errores de I/O en volúmenes de PowerFlex, puede ingresar al estado All-Path-Down (APD) en uno o más volúmenes de PowerFlex. Este estado puede hacer que deje de responder en vCenter.

Por lo general:

  • Algunos hosts ESXi se muestran como Disconnected en vSphere Clients
  • Los errores de I/O se acumulan en vmkernel.log:
2018-01-10T22:30:08.321Z cpu29:33684)ScsiDeviceIO: 2651: Cmd(0x439e41930500) 0x28, CmdSN 0x8819c3 from world 34407 to dev "eui.<mdmId+volId>" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0.
  • hostd.log puede contener los siguientes errores al comienzo del estado APD:
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10313 : Lost connectivity to storage device eui.<mdmId+volId>. Path vmhba64:C0:T27:L91 is down. Affected datastores: <Datastore Name>.
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 778923875,
-->    chainId = 1635216758,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "PHSVCESQL1018.partners.org",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.start",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10314 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down state.
  • hostd.log tiene el mensaje esx.problem.storage.apd.timeout Cuando el hostd El servicio deja de responder:
2017-10-24T17:06:58.277Z info hostd[29A40B70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 690973144,
-->    chainId = 1635216641,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "ESXi.host.local",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.timeout",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       },
-->       (vmodl.KeyAnyValue) {
-->          key = "2",
-->          value = "140"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:58.278Z info hostd[29A40B70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10336 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down Timeout state after being in the All Paths Down state for 140 seconds. I/Os will now be fast failed.
  • Fragmento de vmkwarning que coincide con lo anterior hostd:
Observe que el kernel está intentando eliminar vmhba64:C0:T30:L62, pero no puede debido a hostd Mantenerlo en un estado ocupado durante la reexaminación y el estado de Deja de responder:
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: NMP: nmpUnclaimPath:1516: NMP device "eui.<mdmId+volId>" quiesce state change failed: Busy
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4507: Path vmhba64:C0:T30:L62 is being removed
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4737: Failed to issue command 0x0 (cmdSN 0x0) on path vmhba64:C0:T30:L62: No connection
2017-10-24T17:04:38.268Z cpu8:33147)WARNING: ScsiScan: 2007: Could not delete path vmhba64:C0:T30:L62
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_IssueCommandToDevice:4553: I/O could not be issued to device "eui.<mdmId+volId>" due to Not found
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceRetryCommand:133: Device "eui.<mdmId+volId>": awaiting fast path state update for failover with I/O blocked. No prior reservation exists on the device.
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceStartLoop:725: NMP Device "eui.<mdmId+volId>" is blocked. Not starting I/O from device.
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:603: Retry world failover device "eui.<mdmId+volId>" - issuing command 0x43a6402bbac0
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:678: Retry world failover device "eui.<mdmId+volId>" - failed to issue command due to Not found (APD), try again...
  • Lo siguiente está ocurriendo, como se ve en el storagerm Registros: 
2017-10-24T17:05:49.274Z: Write 0xffcda788[512] -> 68 failed. 38:Function not implemented, offset=0, bufLen=512
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> Write error to fd 68, error: Function not implemented
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> I/Os from datastore eui.207d160928aa82202102c97700000060 took 62.962148(>= 30.000000) seconds to complete stats computation. Reducing its polling frequency.
2017-10-24T17:06:58.277Z: Write 0xffcda788[512] -> 58 failed. 6:No such device or address, offset=0, bufLen=512
2017-10-24T17:07:04.484Z: <Datastore Name, 0> Some host is down, need to reset the slot allocation
2017-10-24T17:07:08.554Z: Skipping device eui.207d160928aa82202102c9530000003e either due to VSI read error or abnormal state
2017-10-24T17:07:08.580Z: open /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile(0x202, 0x0) failed: Input/output error
2017-10-24T17:07:08.580Z: Input/output error Error -1 opening/truncating file /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile
  • Las VM pueden mostrarse como /vmfs/volumes/.../...vmx en lugar del nombre para mostrar.
  • Los puertos DVS pueden comenzar a fallar debido a la pérdida de conectividad con vpxa y vCenter:
2017-10-24T17:06:55.704Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-14505 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/14505
2017-10-24T17:06:55.943Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-8339 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/8339
2017-10-24T17:06:55.994Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-15520 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/15520
2017-10-24T17:06:56.017Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-7684 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/7684
  • Pérdida de conectividad de VM y hosts conectados a vCenter

También es posible:

  • No se puede establecer una conexión SSH a ESXi o a las VM (si la red de administración está en un vSwitch distribuido)
  • No se puede usar esxcli en la sesión de la consola. (Deja de responder, use localcli en su lugar. Consulte la sección Solución alternativa).
  • Es posible que las VM, incluida la SVM, no puedan encenderse o apagarse sin detener los procesos. 
  • Es posible que el host ESXi deje de responder durante el reinicio o el arranque.
  • El proceso de arranque por lo general, pero no necesariamente, deja de responder después de nfs41client Se carga el módulo. Los siguientes mensajes se muestran en la consola del host (DCUI).
nfs41client loaded successfully

Impacto

  • Incapacidad de administrar hosts ESXi a través de vCenter o de establecer una conexión SSH.
  • Sin funcionalidades de vMotion

Causa

En el estado APD, las I/O del usuario ESXi (hostd agente) o cualquier I/O del SO huésped que no se anule debido al tiempo de espera agotado del SO huésped, se vuelven a intentar indefinidamente, lo que agota los recursos del sistema y conduce al estado de falta de respuesta de ESXi en vCenter.

Resolución

  • Reiniciar los hosts ESXi sin corregir la condición subyacente de APD no ayuda, ya que el host puede volver a ingresar APD.
  • Si es necesario ejecutar comandos en un host ESXi que experimenta APD, utilice "localcli" en lugar de "esxcli", ya que este último deja de responder.
Por ejemplo:
  • Utilice lo siguiente para comprobar si los almacenes de datos se muestran como montados:
[root@92U-16:~] localcli storage filesystem list
Mount Point                                        Volume Name  UUID                                 Mounted  Type    Size           Free
-----------------------------------------------------------------------------------------------------------------------------------------
/vmfs/volumes/5975cf1e-9306f9bc-0dbc-a0369fdaccbc  SATADOM17    5975cf1e-9306f9bc-0dbc-a0369fdaccbc  true     VMFS-5    55834574848   53979643904
/vmfs/volumes/59916bcd-22a730ae-db91-a0369fdaccbc  LocalDS17    59916bcd-22a730ae-db91-a0369fdaccbc  true     VMFS-6  1920118816768  986341965824
/vmfs/volumes/5975cf15-c44cea1b-de13-a0369fdaccbc               5975cf15-c44cea1b-de13-a0369fdaccbc  true     vfat        299712512      83927040
/vmfs/volumes/16a83277-c690cda2-9723-26fe2e41d0c3               16a83277-c690cda2-9723-26fe2e41d0c3  true     vfat        261853184      97923072
/vmfs/volumes/5975cf1f-17e61cfc-a0ae-a0369fdaccbc               5975cf1f-17e61cfc-a0ae-a0369fdaccbc  true     vfat       4293591040    4260626432
/vmfs/volumes/79e9c87d-f55f1864-b3ce-6e24607afc68               79e9c87d-f55f1864-b3ce-6e24607afc68  true     vfat        261853184      99840000
  • Utilice lo siguiente para intentar volver a examinar en el nivel del host:
localcli storage filesystem rescan
  • Si es necesario reiniciar un host ESXi que ya está en el estado APD, anote los volúmenes asignados a él y desasigne temporalmente; Asígnelos nuevamente al host después de solucionar el problema.
 
Nota: Si el host ESXi está conectado a varios sistemas MDM o PowerFlex, solo se deben desasignar los volúmenes del sistema afectado.
 
  • Si la solicitud en unmap_volume Durante la recuperación, es posible que algunas de las VM deban volver a registrarse después de que se vuelvan a asignar los volúmenes y los almacenes de datos se vuelvan a montar.
Solución
En la versión 2.0.1.3, se introdujo la función de pérdida permanente del dispositivo (PDL), que está deshabilitada de manera predeterminada. Cuando esta función está habilitada, puede convertir APD en PDL después de que el SDC no pueda enviar I/O a un volumen después de 60 segundos. Este valor de tiempo de espera agotado aún puede ser más largo de lo que algunos entornos pueden soportar sin ver el impacto y puede requerir más ajustes.

Información adicional

Productos afectados

PowerFlex rack, ScaleIO
Propiedades del artículo
Número del artículo: 000437810
Tipo de artículo: Solution
Última modificación: 27 mar 2026
Versión:  3
Encuentre respuestas a sus preguntas de otros usuarios de Dell
Servicios de soporte
Compruebe si el dispositivo está cubierto por los servicios de soporte.