PowerFlex: ESXi SDC, vCenter da Tüm Yollar Çalışmıyor ve Yanıt Vermiyor

Summary: ESXi sunucuları, bir veya daha fazla PowerFlex disk bölümündeki All Paths Down (APD) durumu nedeniyle vCenter da yanıt vermeyi durduruyor.

Bu makale şunlar için geçerlidir: Bu makale şunlar için geçerli değildir: Bu makale, belirli bir ürüne bağlı değildir. Bu makalede tüm ürün sürümleri tanımlanmamıştır.

Symptoms

ESXi SDC, PowerFlex disk bölümlerinde sürekli olarak G/Ç hatalarıyla karşılaştığında, bir veya daha fazla PowerFlex disk bölümü için Tümüne Giden Yol Aşağı (APD) durumuna girebilir. Bu durum, vCenter da yanıt vermeyi durdurmasına neden olabilir.

Genellikle:

  • Bazı ESXi ana bilgisayarları vSphere istemcilerinde bağlantısı kesildi olarak görünüyor
  • G/Ç hataları vmkernel.log:
2018-01-10T22:30:08.321Z cpu29:33684)ScsiDeviceIO: 2651: Cmd(0x439e41930500) 0x28, CmdSN 0x8819c3 from world 34407 to dev "eui.<mdmId+volId>" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x4 0x0 0x0.
  • hostd.log APD durumunun başında aşağıdaki hataları içerebilir:
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10313 : Lost connectivity to storage device eui.<mdmId+volId>. Path vmhba64:C0:T27:L91 is down. Affected datastores: <Datastore Name>.
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 778923875,
-->    chainId = 1635216758,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "PHSVCESQL1018.partners.org",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.start",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:08.144Z info hostd[2AE0BB70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10314 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down state.
  • hostd.log mesajını içerir esx.problem.storage.apd.timeout Ne zaman hostd Hizmet yanıt vermiyor:
2017-10-24T17:06:58.277Z info hostd[29A40B70] [Originator@6876 sub=Hostsvc.VmkVprobSource] VmkVprobSource::Post event: (vim.event.EventEx) {
-->    key = 690973144,
-->    chainId = 1635216641,
-->    createdTime = "1970-01-01T00:00:00Z",
-->    userName = "",
-->    datacenter = (vim.event.DatacenterEventArgument) null,
-->    computeResource = (vim.event.ComputeResourceEventArgument) null,
-->    host = (vim.event.HostEventArgument) {
-->       name = "ESXi.host.local",
-->       host = 'vim.HostSystem:ha-host'
-->    },
-->    vm = (vim.event.VmEventArgument) null,
-->    ds = (vim.event.DatastoreEventArgument) null,
-->    net = (vim.event.NetworkEventArgument) null,
-->    dvs = (vim.event.DvsEventArgument) null,
-->    fullFormattedMessage = <unset>,
-->    changeTag = <unset>,
-->    eventTypeId = "esx.problem.storage.apd.timeout",
-->    severity = <unset>,
-->    message = <unset>,
-->    arguments = (vmodl.KeyAnyValue) [
-->       (vmodl.KeyAnyValue) {
-->          key = "1",
-->          value = "eui.<mdmId+volId>"
-->       },
-->       (vmodl.KeyAnyValue) {
-->          key = "2",
-->          value = "140"
-->       }
-->    ],
-->    objectId = "ha-eventmgr",
-->    objectType = "vim.HostSystem",
-->    objectName = <unset>,
-->    fault = (vmodl.MethodFault) null
--> }
2017-10-24T17:06:58.278Z info hostd[29A40B70] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 10336 : Device or filesystem with identifier eui.<mdmId+volId> has entered the All Paths Down Timeout state after being in the All Paths Down state for 140 seconds. I/Os will now be fast failed.
  • Kod parçacığı vmkwarning yukarıdakilerle örtüşen hostd:
Çekirdeğin kaldırmaya çalıştığına dikkat edin vmhba64:C0:T30:L62, ancak nedeniyle olamaz hostd Yeniden tarama ve yanıt vermeyi durdurma durumu sırasında meşgul durumda tutma:
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: NMP: nmpUnclaimPath:1516: NMP device "eui.<mdmId+volId>" quiesce state change failed: Busy
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4507: Path vmhba64:C0:T30:L62 is being removed
2017-10-24T17:04:38.267Z cpu8:33147)WARNING: ScsiPath: 4737: Failed to issue command 0x0 (cmdSN 0x0) on path vmhba64:C0:T30:L62: No connection
2017-10-24T17:04:38.268Z cpu8:33147)WARNING: ScsiScan: 2007: Could not delete path vmhba64:C0:T30:L62
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_IssueCommandToDevice:4553: I/O could not be issued to device "eui.<mdmId+volId>" due to Not found
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceRetryCommand:133: Device "eui.<mdmId+volId>": awaiting fast path state update for failover with I/O blocked. No prior reservation exists on the device.
2017-10-24T17:04:38.337Z cpu42:34088)WARNING: NMP: nmp_DeviceStartLoop:725: NMP Device "eui.<mdmId+volId>" is blocked. Not starting I/O from device.
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:603: Retry world failover device "eui.<mdmId+volId>" - issuing command 0x43a6402bbac0
2017-10-24T17:04:38.560Z cpu32:33507)WARNING: NMP: nmpDeviceAttemptFailover:678: Retry world failover device "eui.<mdmId+volId>" - failed to issue command due to Not found (APD), try again...
  • Aşağıda görülen olay storagerm Günlük: 
2017-10-24T17:05:49.274Z: Write 0xffcda788[512] -> 68 failed. 38:Function not implemented, offset=0, bufLen=512
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> Write error to fd 68, error: Function not implemented
2017-10-24T17:05:49.274Z: <Datastore Nam, 0> I/Os from datastore eui.207d160928aa82202102c97700000060 took 62.962148(>= 30.000000) seconds to complete stats computation. Reducing its polling frequency.
2017-10-24T17:06:58.277Z: Write 0xffcda788[512] -> 58 failed. 6:No such device or address, offset=0, bufLen=512
2017-10-24T17:07:04.484Z: <Datastore Name, 0> Some host is down, need to reset the slot allocation
2017-10-24T17:07:08.554Z: Skipping device eui.207d160928aa82202102c9530000003e either due to VSI read error or abnormal state
2017-10-24T17:07:08.580Z: open /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile(0x202, 0x0) failed: Input/output error
2017-10-24T17:07:08.580Z: Input/output error Error -1 opening/truncating file /vmfs/volumes//<Datastore Name>/.eui.<mdmId+volId>/slotsfile
  • VM'ler şu şekilde görünebilir: /vmfs/volumes/.../...vmx görünen ad yerine dosyalar.
  • DVS bağlantı noktaları, vpxa ve vCenter ile bağlantı kaybı nedeniyle arızalanmaya başlayabilir:
2017-10-24T17:06:55.704Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-14505 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/14505
2017-10-24T17:06:55.943Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-8339 to file /vmfs/volumes/59553df0-a1c109ac-b164-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/8339
2017-10-24T17:06:55.994Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-15520 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/15520
2017-10-24T17:06:56.017Z warning hostd[29C81B70] [Originator@6876 sub=Hostsvc.NetworkProvider] Error saving dvport 38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c-7684 to file /vmfs/volumes/59553bbc-77b8edaa-15da-54ab3a16bf9d/.dvsData/38 c1 36 50 b6 92 e4 32-1f 16 2d 37 80 dd 7b 2c/7684
  • vCenter'a bağlı VM'lerin ve ana bilgisayarların bağlantı kaybı

Ayrıca mümkün:

  • ESXi veya VM'lere SSH bağlantısı kurulamıyor (yönetim ağı Dağıtılmış bir vSwitch'te ise)
  • Kullanılamaz esxcli konsol oturumunda. (Yanıt vermiyorsa, localcli belirtirdik. Geçici Çözüm bölümüne bakın.)
  • SVM de dahil olmak üzere VM'ler, işlemleri sonlandırmadan açılıp kapanamayabilir. 
  • ESXi ana bilgisayarı, yeniden önyükleme veya önyükleme sırasında yanıt vermeyi durdurabilir.
  • Önyükleme işlemi genellikle, ancak zorunlu olmamakla birlikte, nfs41client modül yüklendi. Ana bilgisayarın konsolunda (DCUI) aşağıdaki mesajlar görüntülenir.
nfs41client loaded successfully

Etki

  • ESXi ana bilgisayarlarını vCenter üzerinden yönetememe veya SSH bağlantısı kuramama.
  • vMotion özelliği yok

Cause

APD durumunda, ESXi kullanıcısından (hostd agent) veya konuk işletim sisteminin zaman aşımı nedeniyle iptal edilmeyen konuk işletim sistemindeki tüm G/Ç'ler süresiz olarak yeniden denenerek sistem kaynaklarını tüketir ve vCenter'da ESXi'nin yanıt vermeme durumuna yol açar.

Resolution

  • Ana bilgisayar APD'ye tekrar girebileceğinden, temeldeki APD koşulunu düzeltmeden ESXi ana bilgisayarlarını yeniden başlatmak yardımcı olmaz.
  • APD kullanan bir ESXi ana bilgisayarında komut çalıştırma gereksinimi varsa "localcli" yerine "esxcli", ikincisi yanıt vermeyi bıraktığı için.
Örneğin:
  • Veri depolarının bağlı olarak görünüp görünmediğini kontrol etmek için aşağıdakileri kullanın:
[root@92U-16:~] localcli storage filesystem list
Mount Point                                        Volume Name  UUID                                 Mounted  Type    Size           Free
-----------------------------------------------------------------------------------------------------------------------------------------
/vmfs/volumes/5975cf1e-9306f9bc-0dbc-a0369fdaccbc  SATADOM17    5975cf1e-9306f9bc-0dbc-a0369fdaccbc  true     VMFS-5    55834574848   53979643904
/vmfs/volumes/59916bcd-22a730ae-db91-a0369fdaccbc  LocalDS17    59916bcd-22a730ae-db91-a0369fdaccbc  true     VMFS-6  1920118816768  986341965824
/vmfs/volumes/5975cf15-c44cea1b-de13-a0369fdaccbc               5975cf15-c44cea1b-de13-a0369fdaccbc  true     vfat        299712512      83927040
/vmfs/volumes/16a83277-c690cda2-9723-26fe2e41d0c3               16a83277-c690cda2-9723-26fe2e41d0c3  true     vfat        261853184      97923072
/vmfs/volumes/5975cf1f-17e61cfc-a0ae-a0369fdaccbc               5975cf1f-17e61cfc-a0ae-a0369fdaccbc  true     vfat       4293591040    4260626432
/vmfs/volumes/79e9c87d-f55f1864-b3ce-6e24607afc68               79e9c87d-f55f1864-b3ce-6e24607afc68  true     vfat        261853184      99840000
  • Ana bilgisayar düzeyinde yeniden tarama yapmayı denemek için aşağıdakileri kullanın:
localcli storage filesystem rescan
  • Halihazırda APD durumunda olan bir ESXi ana bilgisayarının yeniden başlatılması gerekiyorsa bu ana bilgisayarla eşlenen disk bölümlerini not edin ve geçici olarak bunların eşlemesini kaldırın; Sorun çözüldükten sonra bunları tekrar ana bilgisayara eşleyin.
 
Not: ESXi ana bilgisayarı birden fazla MDM veya PowerFlex sistemine bağlıysa, yalnızca etkilenen sistemdeki disk bölümlerinin eşlemesi kaldırılmalıdır.
 
  • Eğer unmap_volume Kurtarma sırasında işlem yapılması gerekir. Disk bölümleri yeniden eşlendikten ve veri depoları yeniden bağlandıktan sonra bazı VM'lerin yeniden kaydedilmesi gerekebilir.
Çözüm
2.0.1.3 sürümünde, varsayılan olarak devre dışı bırakılan Kalıcı Cihaz Kaybı (PDL) özelliği kullanıma sunulmuştur. Bu özellik etkinleştirildiğinde, SDC 60 saniye sonra bir disk bölümüne G/Ç gönderemediğinde APD'yi PDL'ye dönüştürebilir. Bu zaman aşımı değeri, bazı ortamların etkiyi görmeden dayanabileceğinden daha uzun olabilir ve daha fazla ayarlama gerektirebilir.

Additional Information

Etkilenen Ürünler

PowerFlex rack, ScaleIO
Makale Özellikleri
Article Number: 000437810
Article Type: Solution
Son Değiştirme: 27 Mar 2026
Version:  3
Sorularınıza diğer Dell kullanıcılarından yanıtlar bulun
Destek Hizmetleri
Aygıtınızın Destek Hizmetleri kapsamında olup olmadığını kontrol edin.