I have a working Unity VSA running in VM Ware Workstation (test platform). It is successfully providing storage to pods running in CoreOS based Kubernetes cluster.
I have download and installed the CSI drivers from the git repo (https://github.com/dell/csi-unity/tree/v1.0.0) following the instructions and using Helm to create the necessary unity-controller-0 and unity-node- pods.
The unity-controller-0 pod is reliably running, authenticates with the VSA and the logs contain no errors.
The unity-node pods initally come up as Running, but within seconds their state changes to Error and then CrashLoopBackOff. This cycle repeats continually.
Looking at the logs with 'kubectl logs unity-node-k5k5h -n unity -c driver' I see the following:
Endpoint unix:///var/lib/kubelet/plugins/unity.emc.dell.com/csi_sock ls: cannot access /var/lib/kubelet/plugins/unity.emc.dell.com/csi_sock: No such file or directory gounity logger initiated. This should be called only once. csi-unity logger initiated. This should be called only once. time="2020-03-02T23:41:18Z" level=info msg="X_CSI_MODE:node" time="2020-03-02T23:41:18Z" level=info msg="X_CSI_UNITY_NODENAME:node05" time="2020-03-02T23:41:18Z" level=info msg="configured csi-unity.dellemc.com" autoprobe=false endpoint="https://10.0.1.50" insecure=true mode=node nodename=node05 password="******" pvtMountDir=/var/lib/kubelet/plugins/unity.emc.dell.com/disks systemname= user=admin time="2020-03-02T23:41:18Z" level=info msg="identity service registered" time="2020-03-02T23:41:18Z" level=info msg="node service registered" time="2020-03-02T23:41:18Z" level=info msg=serving endpoint="unix:///var/lib/kubelet/plugins/unity.emc.dell.com/csi_sock" time="2020-03-02T23:41:18Z" level=info msg="Executing NodeGetInfo. Host name:node05" time="2020-03-02T23:41:18Z" level=info msg="AutoProbe has not been called. Executing manual probe" time="2020-03-02T23:41:18Z" level=info msg="Executing nodeProbe" time="2020-03-02T23:41:18Z" level=info msg="Executing Authenticate REST client" time="2020-03-02T23:41:21Z" level=info msg="Response code:200 for url: /api/types/loginSessionInfo" time="2020-03-02T23:41:21Z" level=info msg="Authentication response code:200" time="2020-03-02T23:41:21Z" level=info msg="Authentication successful" time="2020-03-02T23:41:21Z" level=error msg="Cannot read directory: /sys/class/fc_host Error: open /sys/class/fc_host: no such file or directory" time="2020-03-02T23:41:22Z" level=info msg="Executing NodeGetInfo. Host name:node05"
The last two lines repeats many times.
So, it looks like it may be something to do with the lack of fibre channel adapter on my worker nodes, but as I'm using the virtual storage appliance I should not need (and can not configure) FC adapters.
Is this currently a limitation of the CSI drivers? Is there a way round this? I'd appreciate any help / guidance anyone can give.
@ankur.patel Excellent video! A welcomed sanity check for myself. I was following all the same steps you show in your video. Unfortunately, my problem persists. I made a post that provides more details on my problem here
I did a watch on kubectl get pods while the deployment process was running. The controllers come up fine (I specified 2 as I have a 4 node cluster, but I also tried using 1).
The nodes will then deploy, come up as running, then error out causing a 1/2 node status.
No matter what I try the driver log show "ls: cannot access /var/lib/kubelet/plugins/unity.emc.dell.com/csi_sock: No such file or directory" and "Cannot read directory: /sys/class/fc_host Error: open /sys/class/fc_host: no such file or directory"
Its not clear where that directory resides. Is the pod looking at the the actual EMC array? I checked the EMC with the service account, that file does not exist. Also why is the pod even checking that? I disabled FC, I only want iSCSI.
For the registrar pod I see the following "Received NotifyRegistrationStatus call: &RegistrationStatus{PluginRegistered:false,Error:RegisterPlugin error -- plugin registration failed with err: rpc error: code = Unavailable desc = runid=7 The node [taxmd-k8n03-v] is not added to any of the arrays,}" "E0402 13:44:07.602419 1 main.go:92] Registration process failed with error: RegisterPlugin error -- plugin registration failed with err: rpc error: code = Unavailable desc = runid=7 The node [taxmd-k8n03-v] is not added to any of the arrays, restarting registration container."
I cant tell which error is actually causing the pods to fail.
Couple of notes on some things I had to do to actually get the installer to work. - We are using an ubuntu env, the controller appears to want to use root when logging into nodes. That will not work on ubuntu. I had to adjust the sudoers file to allow a local user rights to do things, then edit the script to perform the cat command it uses with sshpass as sudo. - The verify script would fail every time when checking the pods unless I had each pod's iscsi initiator logged into a target and a drive mapped, is this expected?
Digging deeper in the code, I can see the go code is trying to check for fiber hba. The pods breaks because of the code behind them. Why are the pods checking for hba's when I dont have any? Not sure how anyone else was able to get this installed but there are no hba drivers on our nodes. This will fail every time. How do I disable this?
I got it working! The driver pod was not failing at all (can ignore the fc_host errors), it was the registrar pod. Problem was the Unity still had old hosts built in it from the failed attempts. Cleared those out and it finally worked. There were other issues as well, I will make a new post detailing tips and tricks for others.
bmcfeeters
1 Rookie
•
72 Posts
3002
0
Posted March 5th, 2020 07:00
Hi Carl,
Unfortunately, we don't have iSCSI support in the Unity CSI driver yet. That functionality is targetting for release in mid-April.
Thanks
Bryan