Announcement Banner
UNSOLVED

sluetze

updated

11 years ago

S

sluetze

2 Intern

•

300 Posts

0

1955

July 6th, 2015 09:00

single threaded API

Hello community,

I already have a SR open with support, but would like to get some opinions / information here.

We have an environment in a size, which makes it necessary to automate use-cases like Create folders, Modify folders and Delete folders. Since the Isilon offers an API we used it to realize this business need.

We just ran into the fact, that the Isilons API works single threaded, which means each API-call has to wait until all API-calls which were issued before are finished. In cases of calls which do not take a lot of time (like quota modification, share creation, creation and modification of groupdrives and so on) this is no real problem. But there is at least one call which takes it's time. DELETE of a file structure. In this case the success and duration of the DELETE-call depends on the amount of inodes in the structure. we did some tests and can say the following:

Size of Filestructure Number of Files / Folders Duration of DELETE call
25GB ~40.000 (mixed files and folders) ~5 min
25GB 25 instant
some KB 100.000 (mostly files) ~10 min
some KB 300.000 (mostly folders) ~30 min

so we can see, we have the dependencies mostly on the amount of files and folders in this structure. I also guess the overall duration is dependent on the snapshot policies, metadata policies, hardware (disks, nodes, cpu), load and some other configurations. But i am quite satisfied with the duration of the delete-calls. my problem is, that each call i send after such an delete is delayed by that time. and that these times add. so if i send 50 5min-Delete-Calls all other tasks are delayed by 250 minutes - which results in some screwed up applications which want to use the API i.e. for desaster recovery....

Does anyone has made the same observation? Any ideas or solutions?



First reply from support was "works as designed". I leave to the humble reader if a scaling up to petabytes Filesystem / BigData Storage System should have problems with delete-calls of folders with many inodes via API.

IMHO the API should either

- be multithreaded to allow several commands parallel

- have a fastlane for "fast" or "more important" calls

- make the DELETE-Call "asynchronous" so it schedules a treedelete-job in the background (or something like that) but does not block the API while doing that DELETE

Message was edited by: sluetze

  • scott_owens

    60 Posts

    1107

    0

    Posted July 6th, 2015 15:00

    Instead of doing a delete, could you do a move to a temporary folder such as /ifs/delete_me_later.

    At the end of the day, you could schedule a treedelete job to delete the /ifs/delete_me_later folder

  • Newday3000

    57 Posts

    1107

    0

    Posted July 6th, 2015 15:00

    The API is multithreaded, for cluster configuration changes.

    Are you making parallel requests from your application ? if it's a loop

    from a scripting language it's likely single threaded requests and they're

    all queued up.

    Regards

    Andrew

    On Mon, Jul 6, 2015 at 6:42 PM scott_owens

  • sluetze

    2 Intern

    •

    300 Posts

    1107

    0

    Posted July 6th, 2015 22:00

    @scott_owens:

    Another workaround is to delete the contents over smb or nfs before scheduling any deletes. But this workarounds would need me to either schedule a additional script /job on the isilon which deletes the folders. which I would like not to do.

    @Andrew: Yes the Application fires the requests parallel. For OneFS 7.0. there are two APIs documented. platform API and filesystem API - maybe platform is multithreaded (didn't test that) but filesystem definitively is singlethreaded. Or I am doing something significant wrong.

  • Peter_Sero

    6 Operator

    •

    1169 Posts

    1107

    0

    Posted July 7th, 2015 00:00

    If you have been using a single session with cookie auth, you might try instead sending the DELETEs as isolated, with per request auth.

    -- Peter 

  • sluetze

    2 Intern

    •

    300 Posts

    1107

    0

    Posted July 8th, 2015 00:00

    Hi Peter,

    we used Basic Authentification, so we authenticate for each API-call.

    since every node has it's own apache-server you are only able to do a cookie-authentication if you do the authentication per node. Maybe I did not get the concept, but this means I have to "bind" my automation onto one node. Because by using smartconnectzones I can only authenticate to the apache-server of the node that was involved in the cookie-creation. Also I have to implement functionalities to handle nodefailures - either by dynamic IP configuration (not recommended for SMB environments) or by external functionalities in my automation.

    So a clusterwide authentication would be great.

    And central logging (since I have automation-procedures which consist of more than one API-call I have to search each call on all nodes... could be anywhere...)

  • Peter_Sero

    6 Operator

    •

    1169 Posts

    1107

    1

    Posted July 8th, 2015 02:00

    ... did a quick test with curl,  something like:

    curl -X DELETE ... url-dir1 &

    curl -X DELETE ... url-dir2 &

    Found that each run takes about 25s for 10000 files,

    and both absolutely run in parallel, one can see

    both directories draining files simultaneously,

    top -H on the Isilon shows two busy threads:


      PID USERNAME      PRI NICE   SIZE    RES STATE   C   TIME   WCPU COMMAND

    35745 root            8    0   152M 24560K CPU7    7   0:14 56.98% isi_object_d

    35745 root            8    0   152M 24560K ref     0   0:14 56.98% isi_object_d

    This can be starting point for you. My next guess would be that your application

    fails to run the threads in parallel. E.g. treads do get created but then become

    serialized when attempting to use some pooled resource for the https requests.

    hth

    -- Peter

  • sluetze

    2 Intern

    •

    300 Posts

    1107

    0

    Posted July 8th, 2015 02:00

    thanks peter - I will try some additional testing outside my automation application.

  • Peter_Sero

    6 Operator

    •

    1169 Posts

    779

    0

    Posted July 8th, 2015 02:00

    Yep, the cookies don't work with smartconnect load balancing... need to stick with the initial node.

  • sluetze

    2 Intern

    •

    300 Posts

    1107

    0

    Posted July 10th, 2015 05:00

    Hi Peter,

    did some Tests via powershell and was able to get DELETE-calls running in parallel on two different nodes and also running in parallel on one node (using two sessions from the same client) on OneFS7.1.1 branch. Will do some further tests on 7.0.2.x branch to see if I can observe a different behavior.

    Regards

    -- sluetze

    Edit: Finished testing. Parallel processing also works on the namespace API in the 7.0.2 branch.

    I have to search the error in my application. Fun Fact: support is still investigating with engineering how to get the api away from singlethreading

  • Peter_Sero

    6 Operator

    •

    1169 Posts

    519

    0

    Posted July 10th, 2015 07:00

    Good luck! And give support a relief...