Dell Unity: SPA and SPB Panic Due to Write Request Pool Exhaustion
Summary: A DAE expansion can generate a burst of I/O activity that, under specific conditions, may exhaust the RMD global write request pool. If this occurs, both Storage Processors (SPA and SPB) can encounter the same panic condition simultaneously, resulting in a temporary loss of data access. Upgrading to Unity OE 5.1.x or later increases the RMD write request pool size and reduces the likelihood of this issue occurring during periods of high I/O activity. ...
This article applies to
This article does not apply to
This article is not tied to any specific product.
Not all product versions are identified in this article.
Symptoms
The user expanded a DAE, resulting in an I/O burst. Both SPA and SPB subsequently experienced a panic condition simultaneously, causing a period of data unavailability.
SPA Panicked two times:
Fri Mar 19 05:57:34 UTC 2021 system-state: set sp-critical-error
Fri Mar 19 06:38:35 UTC 2021 system-state: set sp-critical-error
SPB Panicked three times:
Fri Mar 19 06:13:43 UTC 2021 system-state: set sp-critical-error
Fri Mar 19 06:39:35 UTC 2021 system-state: set sp-critical-error
Fri Mar 19 06:51:05 UTC 2021 system-state: set sp-critical-error
/spa/EMC/C4Core/log> zgrep -A1 0x81254002 c4_safe_ktrace*
c4_safe_ktrace.log.10.gz:2021/03/19-06:50:11.935709 11 7FA30677870B std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.10.gz:2021/03/19-06:50:11.935710 ~~~~ 7FA30677870B std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.15.gz:2021/03/19-06:38:29.041315 ~~~~ 7FA9A02CF709 std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.15.gz:2021/03/19-06:38:29.041316 ~~~~ 7FA9A0FF4701 std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.23.gz:2021/03/19-05:57:33.014928 101K 7F4A6A827706 std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.23.gz:2021/03/19-05:57:33.014932 0 7F4A6A827706 std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.7.gz:2021/03/19-06:59:52.097333 272 7FC21798570C std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.7.gz:2021/03/19-06:59:52.097334 ~~~~ 7FC21798570C std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
/spb/EMC/C4Core/log> zgrep -A1 0x81254002 c4_safe_ktrace*
c4_safe_ktrace.log.11.gz:2021/03/19-06:50:59.806541 3880 7FBDE30B570E std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.11.gz:2021/03/19-06:50:59.806545 ~~~~ 7FBDE30B570E std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.21.gz:2021/03/19-06:39:29.781207 13 7F1582EEF70D std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.21.gz:2021/03/19-06:39:29.781208 ~~~~ 7F1582EEF70D std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.24.gz:2021/03/19-06:10:44.513508 1105 7FDFD279970C std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.24.gz:2021/03/19-06:10:44.513511 ~~~~ 7FDFD279970C std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)c4_safe_ktrace.log.10.gz:2021/03/19-06:50:11.935709 11 7FA30677870B std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.10.gz:2021/03/19-06:50:11.935710 ~~~~ 7FA30677870B std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.15.gz:2021/03/19-06:38:29.041315 ~~~~ 7FA9A02CF709 std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.15.gz:2021/03/19-06:38:29.041316 ~~~~ 7FA9A0FF4701 std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.23.gz:2021/03/19-05:57:33.014928 101K 7F4A6A827706 std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.23.gz:2021/03/19-05:57:33.014932 0 7F4A6A827706 std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.7.gz:2021/03/19-06:59:52.097333 272 7FC21798570C std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.7.gz:2021/03/19-06:59:52.097334 ~~~~ 7FC21798570C std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
/spb/EMC/C4Core/log> zgrep -A1 0x81254002 c4_safe_ktrace*
c4_safe_ktrace.log.11.gz:2021/03/19-06:50:59.806541 3880 7FBDE30B570E std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.11.gz:2021/03/19-06:50:59.806545 ~~~~ 7FBDE30B570E std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.21.gz:2021/03/19-06:39:29.781207 13 7F1582EEF70D std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.21.gz:2021/03/19-06:39:29.781208 ~~~~ 7F1582EEF70D std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
c4_safe_ktrace.log.24.gz:2021/03/19-06:10:44.513508 1105 7FDFD279970C std:RMD: Allocating write rqst failed! Status is 0x81254002
c4_safe_ktrace.log.24.gz:2021/03/19-06:10:44.513511 ~~~~ 7FDFD279970C std:Cmipci1: Notify peer that we are going down NOW. We can't wait!! (0x1)
SPA:
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 139957593749248 aka 139957593749248) [PID:30127 TID:24677 CORE:11 [csx_ic_std.x] [asyncFlush185] [03/19/2021 05:57:30 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 140366523385600 aka 140366523385600) [PID:29836 TID:31282 CORE:7 [csx_ic_std.x] [asyncFlush116] [03/19/2021 06:38:30 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 140338173945600 aka 140338173945600) [PID:29848 TID:7080 CORE:3 [csx_ic_std.x] [asyncFlush46] [03/19/2021 06:50:11 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 140471598331648 aka 140471598331648) [PID:29945 TID:25809 CORE:7 [csx_ic_std.x] [asyncFlush100] [03/19/2021 06:59:52 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
SPB:
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 140599284365056 aka 140599284365056) [PID:29864 TID:25385 CORE:13 [csx_ic_std.x] [asyncFlush214] [03/19/2021 06:09:18 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]/WriteRequestPoolSize
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 139730383755008 aka 139730383755008) [PID:30143 TID:25017 CORE:8 [csx_ic_std.x] [asyncFlush65] [03/19/2021 06:39:31 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
c4_safe_native.log:CSX RT: panic requested at: KLogBugCheck.c:57 (thread: 140453508536064 aka 140453508536064) [PID:29844 TID:27770 CORE:9 [csx_ic_std.x] [asyncFlush202] [03/19/2021 06:51:01 UTC]] (panic action:DEFAULT expr:<no-expr> flags:-) [info:0]
Cause
All observed panics are consistent with a known issue. The panic occurred in RMD when it was unable to allocate write requests from the global write request pool. This condition was triggered by a burst of I/O activity, which exhausted the available write request resources.
Resolution
Permanent Fix
Upgrade to Unity OE 5.1.x. In this release, the RMD global write request pool size has been increased to better handle bursts of I/O activity.
Workaround
Follow the below plan to increase the write request global pool size to avoid the panic. This occupies additional memory for each SP.
To increase the write request global pool size to 32768 (as an example, or any other value), follow below steps on both SPs:
- Check current value:
-
reg_tool get /SYSTEM/CurrentControlSet/Services/RemoteMirroring/Parameters/WriteRequestPoolSize
-
- Set value to 32768 (as an example)
-
reg_tool set /SYSTEM/CurrentControlSet/Services/RemoteMirroring/Parameters/WriteRequestPoolSize=REG_DWORD@0x00008000
-
- Reboot both the SPs one by one.
Dell Unity: How to Reboot a Storage Processor (User Correctable).
Dell Unity: How to monitor the peer SP boot sequence (Dell Correctable)
Caution:
This parameter should be reset back to the default setting after the NDU completes.
Additional Information
For more detailed procedures, see SolVe Online: Self Service Procedures for Products.
Affected Products
Dell EMC Unity XT 380, Dell EMC Unity XT 380F, Dell EMC Unity XT 480, Dell EMC Unity XT 480F, Dell EMC Unity 650F, Dell EMC Unity XT 680, Dell EMC Unity XT 680F, Dell EMC Unity Family |Dell EMC Unity All Flash, Dell EMC Unity Family
, Dell EMC Unity Hybrid
...
Products
Dell Unity 450F DC, Dell Unity 300 DC, Dell Unity 350F DC, Dell Unity 400 DC, Dell EMC Unity XT 880, Dell EMC Unity XT 880F, Dell Unity Operating Environment (OE)Article Properties
Article Number: 000184814
Article Type: Solution
Last Modified: 12 ربيع الآخر 1448
Version: 4
Find answers to your questions from other Dell users
Support Services
Check if your device is covered by Support Services.