drb45

updated

5 years ago

D

drb45

1 Rookie

•

15 Posts

0

3088

April 9th, 2021 11:00

csi-powerscale 1.5.0 NFS client entry changes

In csi-powerscale up through 1.4.0, NFS clients would be entered using the node FQDN if resolvable. It seems in 1.5.0, this has changed again and now entries are added using just the IP.

This is fine, but it caused an issue with existing volumes, especially single-writer volumes. Volumes attached at the time of upgrade stay attached, but when they are detatched/un-mounted, the existing FQDN client entry is not removed from the export. Then, the next mount attempt will fail:

[Warning] AttachVolume.Attach failed for volume "k8s-abc1234zyx" : rpc error: code = FailedPrecondition desc = runid=41 export '76' in access zone 'MYAZ' already has other clients added to it, and the access mode is SINGLE_NODE_WRITER, thus the request fails

My guess is the 1.5.0 version is now only looking for the IP to remove, and so doesn't clean up the FQDN entry that was originally added by 1.4.0?

Another issue I've noticed is that "localhost" is also being added to the client lists on the export. And localhost is left behind on the export when the volume un-mounts.

  • Thar_J

    42 Posts

    1919

    0

    Posted July 9th, 2021 01:00

    Hi @drb45 

    I apologize for the delayed response. 

    Are you using the current csi-powerscale 1.6 version.

    If you are using 1.5 driver, kindly try this:

    Example : kubectl patch daemonset isilon-node -n isilon -p ‘{“spec”: {“template”: {“spec”:{“dnsPolicy”: “ClusterFirst”}}}}’

    If  using 1.6 driver, then the dnsPolicy value is configurable via myvalues.yaml file, in myvalues.yaml kindly change the  dnsPolicy”: “ClusterFirst"  on all nodes.
     
    Please check to change dnsPolicy : 'ClusterFirst' and see if the problem persists. 
     
    Regards
    Thar_J
  • Thar_J

    42 Posts

    1997

    0

    Posted April 12th, 2021 00:00

    Hi 

    Regarding localhost :

    This was done explicitly to enhance the security. An empty export is accessible to all the hosts on the network -- so we populate it with a dummy 'localhost' entry so as to restrict this access.

     

  • Thar_J

    42 Posts

    1989

    0

    Posted April 12th, 2021 01:00

    Would like to know is it a Helm based installation or Operator based installation. 

  • drb45

    1 Rookie

    •

    15 Posts

    1955

    0

    Posted April 12th, 2021 05:00

    It's a Helm install.

    On the localhost thing - I was pretty sure but just tested it to verify, an empty NFS client list doesn't make an export wide open, it can't be mounted by any clients. But I can see the logic in the dummy localhost.

  • Thar_J

    42 Posts

    1842

    0

    Posted April 15th, 2021 02:00

    Hi 

    Just wanted to understand was there a possible change in the network config (dns change) any recently.

    Kindly attach the kubectl logs for below , we would like to analyze the logs.

    kubectl get pods -n isilon

    kubectl logs -n isilon -c driver

    kubectl logs -n isilon -c driver

     

  • drb45

    1 Rookie

    •

    15 Posts

    1816

    0

    Posted April 15th, 2021 05:00

    There were no network or DNS changes. I had to roll back to version 1.4.0 because of the change in behavior. After rolling back, it resumed the previous behavior of using FQDN on the export client lists.

  • Thar_J

    42 Posts

    1778

    0

    Posted April 16th, 2021 00:00

    Hi @drb45

    I understand after rolling back the csi-powerscale driver from 1.5 to 1.4 the previous behaviour of FQDN is working fine. We have tried the upgrade in various environments at ease , We would like to isolate the underlying cause which has caused the FQDN issue.  

    Would like to know the steps you followed to upgrade the driver from 1.4 to 1.5.0 previously.

    Kindly let us know to upgrade the current powerscale driver from 1.4.0 to 1.5.0. We would like to help you on that. 

    Regards

    Thar_J

     

     

     

     

     

     

  • drb45

    1 Rookie

    •

    15 Posts

    1760

    0

    Posted April 16th, 2021 07:00

    I'll have to build a test cluster to try and reproduce this, because I can't keep upgrading/downgrading in our live environment since it's an actively used application. I'll report back when I've had a chance to do that.

    If it helps, our cluster is running RKE 1.18.15 (being upgraded to 1.19.9 later today) on RHEL 7 nodes.

    My process to upgrade was:

    • Updated my config file to align with the changes from 1.4 -> 1.5
    • Created the cluster credential secret using the new secrets.json file
    • Ran the csi-install.sh script with the --upgrade flag
  • drb45

    1 Rookie

    •

    15 Posts

    1406

    0

    Posted June 10th, 2021 12:00

    Hi @Thar_J

    Sorry for the long radio silence on this, I've been slammed with other projects and this hasn't been a high priority since 1.4.0 was working fine.

    Today I had some time so I built a new K8s (RKE) 1.19.10 cluster on RHEL 7 nodes to test this on, which matches our production cluster. I started with csi-powerscale v1.3.0.1 because I was also testing something else which I'll put in another post (fsGroup behavior). I then upgraded it to 1.4.0, and 1.5.0. Below are my findings:

    • v1.3.0.1: NFS client(s) = worker node FQDN
    • v1.4.0: NFS client(s) = worker node FQDN
    • v1.5.0: NFS client(s) = worker node IP

    Between versions there were no changes at all to the OS or networking config on the nodes, in fact they weren't even rebooted.

     

  • Thar_J

    42 Posts

    1378

    0

    Posted June 13th, 2021 23:00

    Hi @drb45 

    The CSI Driver first tries with FQDN, if it doesn't work, it tries via IP. But the priority is given to FQDN.

    Regards

    Thar_J