Storytime: Tales of HPC Disasters, Recovery, and Resilience webinar poster

Storytime: Tales of HPC Disasters, Recovery, and Resilience

Watch Now

Ever wondered what it's like to face High Performance Computing disasters head-on and emerge stronger than ever? Join our exclusive panel of HPC experts as they reminisce on their unforgettable experiences, share their epic tales of recovery, and impart invaluable lessons learned along the way.

Whether you're a seasoned pro or just starting your HPC journey, this webinar is a must-attend.

Speakers

  • Zane Hamilton, Vice President of Sales Engineering, CIQ: LinkedIn

  • Rose Stein, Sales Operations Administrator, CIQ: LinkedIn

  • Forrest Burt, HPC Systems Engineer, CIQ: LinkedIn

  • Allan Sill, Managing Director of HPCC, Texas Tech University: LinkedIn

  • Jonathon Anderson, Solutions Architect Manager, CIQ: LinkedIn

  • Fernanda Foertter, Director of HPC, Voltron Data: LinkedIn

  • Gary Jung, ScienceIT Department Head, Lawrence Berkeley National Laboratory: LinkedIn

Transcript

[Music] oh [Music] [Applause] [Music] oh [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on powering the next generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and aper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source Yay good afternoon happy Thursday yes thank you it is a wonderful Thursday I agree happy Thursday to you as well can you believe that we are like three weeks out from SC

four that's totally amazing are you going to SC I I am going and I'm also excited there's people on here that I haven't seen since SE last year that I'll actually get to see in person so pretty excited about that very exciting so if any of you out there are going and you would like to connect with ciq we are at Booth 1855 so definitely come by and if you want to actually schedule a meeting so that you can grab one of our amazing experts and talk to them oneon-one just let me know reach out on the website and that email will come to me

and we will get you all set up so that is very exciting but today right now I'm sorry are you still going off on that no I'm not my headphones cut out though so that was fun did so I was just I was like you were talking I could see that you were talking I was just hoping they came back by the time you finished well you faked it really well man you did like the traditional like the man nod oh yeah yep yeah whatever she's saying the man nod man I'll give you that one so today I so we had do research um Computing

roundtables every now and then I think probably like once a month about is but today is cool and I don't know maybe it's just like the title story time I love story time like that is like one of the the best maybe it's like from like being a little kid you know Mom came in or the dad or somebody or maybe you made it up in your own mind but like you know the teacher sat you down and told you a story and it's it's engaging and it's interesting and there's always some kind of drama we're definitely talking about the drama today because we want

View full transcriptHide full transcript

to know from our amazing guests who are about to pop on here in one second HPC disasters and how to recover and learn from mishaps because they happen don't matter how smart you are no matter how much you planned they're going to happen so bring them on let's bring on our amazing HPC Engineers there we go I tried to invite some other people that that have had some disasters that wanted to talk but they weren't able to come today so they were having disasters they were having their own dister not funny Gary it's not funny it's good to see all you guys how is everyone

excellent doing well yeah good well I know we don't have a lot of time with you Alan so we'll start off and kind of and do quick introductions and we'll start off with you if you don't mind Alan tell us who you are uh Alan S I run the high performance Computing Center at Texas Tech University uh we just had a case study uh studied published by ciq actually and uh I co-direct multi-university um cloud and autonomic Computing uh Cooperative Research Center indust University Cooperative Research Center thanks and you will be at SC and you will be doing a talk I believe uh possibly we

we will have an announcement there um uh of an event happening at Texas Tech uh that will be the entire Community but I can't say what it is until then no that's fine thank you Alan Gary good to see you again it's been a while yeah it's been a little bit happy to be back uh yeah my name is Gary I'm the science it department head at Berkeley lab where I manage the HPC program for the institution and I also manage the HPC program for UC Berkeley excellent Jonathan hey thanks everyone um Jonathan Anderson I'm a Solutions architect here with ciq and I've been doing

this HBC thing for quite a while looking forward to sharing some of the the stories of the past excellent and finally Forest good morning everyone my name is Forest Bert I'm a Solutions architect here at CQ on the same team as Jonathan and uh my background in HBC is kind of from the academic National Lab space I was previously a CIS admin before I was a Solutions architect here um CIS admin on an HPC cluster while I was a student uh so that was quite a time we got some interesting stories from that so I'm excited for today sounds like a disaster in the making

just does just kidding Forest just kidding it was always kind of in motion you know true it's academic HBC at small institutions it's it can be a battle at times that's awesome yes because you're always trying to push right like you just want to like how much can this handle so you then you find out so I think you did you want to go first and like talk do you have something on your mind specifically a story that you want to share with us I know you've probably got a few oh welcome Fernando you got to introduce yourself yeah perfect timing uh I'm Fernanda forer

and uh I don't know I don't know what to say anymore I work now uh I've been in HBC since grad school 2006 so it's coming up to 20 years soon God I'm old okay crazy how fast it's gone isn't it so fast it's awesome all right Alan you have stories and you don't have a lot of time so we want we want to hear your disaster story sorry something came up this afternoon uh yes so um I want to take you to um a few years ago was just searching for some pictures to illustrate but I can't give them up in time uh when

we had one of our snowstorms in here in West Texas the jet stream either blows North and you go running or it blows South and you get snow and uh we got a lot of snow um it was just after the New Year University was shut down and we got feet of of of drifts in front of the um my data center I got a notice that uh things had gone offline this was before we had our generator um part of the moral of the story is how to convince your boss to give you that stuff you've been asking for to keep the disasters away

okay so um I've been telling them I went a generator for a long time this is one how I got it uh so all the power is off um it takes me uh you know uh 20 minutes to dig out of my driveway I drive on relatively okay roads get there takes me a good hour to dig with the shovel that I brought with me into the only room only door I had a a card key for the next step is actually getting more access to the building uh get the door open uh the uh everything is dark but there's this really bad state in

the machine room we have a UPS but I go in there and all of the batteries have been drained out of the ups and including the um cach battery for the storage system and the power is coming on and off every few minutes uh the building generator tries to reset uh and there was some of our equipment that was on the building generator anyway um basically for all that time we had been turning on and off the power to a um storage system with a drained cash battery um you know jumping forward for interest of time to the part where we brought in uh vendor

consultants and the expert said I have never seen a file system come back from this level of damage um it was luster uh we uh got to the point where we were doing bit level I kid you not bit level editing on Lisk FS configuration files because that system was not under support contract it was a homebuilt luster system on Commercial Hardware so there's lots of lessons to take away from all of this but the story actually isn't over because it comes in two parts uh so miraculously um after sort of managing this process and doing a lot of the stuff myself uh we got

like 99.9% of the home system back and and high 99 Point somethings of the work system I forget what it was about 95% of the scratch system which is enough to hose most of your scratch but it's scratch um and that's one reason I have this current job right now uh was resolving all that um we got the storage system put under support contract we got the process started which a couple years later got it generator but a year later power outage right about the same time not nearly as severe weather uh the system goes down it it comes back up and the file system

is once again trashed and um well I'll have to jump ahead just in interest of time again but uh the end of that part of the story is the vendor Who Sold us the hardware initially admitted it was a bug in their controller that caused the initial problem in the first place um and they credited us um uh several months worth of support contract costs not quite the full year and um and that I got flown into the next um companywide meeting to tell this story so morals of this story um you know support contracts can be good things they're no substitute for for you

know the you and the vendor actually understanding what's going on um keeping the power on is by and large a good good thing if you can manage that uh don't believe the initial diagnosis uh even if it says you're hosed um you know and don't believe that the fix is a fix unless you've tested it so that's as quickly as I can tell the story uh uh with time and a beverage I can make a much longer version but uh I'll stop there it's fantastic I mean it's not fantastic but it's a great story know there's a lot of pain involved in that thank you

for sharing Alan yeah all right Gary I think there was a request for a story from Gary I don't know if you saw that in the chat Gary there is the jgi flooding so this is actually Mother Nature pushing the limits on you guys this is this is different than what I actually actually thought but um Gary before before you go on to it I just want to say congratulations if you guys know yes but he um received the 2023 Berkeley lab Lifetime Achievement Award so we are in the presence of greatness awesome congratulations y that's awesome thank you very much abut I'm very honored

by that award fantastic so um so yes so the uh the jgi I have to say that I only know about this secondhand and uh but it it there was an interesting thing where um and it goes to show you uh what happens when you don't have the right environmentals but there was a large user facility who ran their own systems um and uh at one point uh something happened to the environmental uh Cooling and the uh room began to overheat and so it uh didn't set off the smoke alarm but the the heat did set set off the sprinkler system so it it actually

sprayed the whole data center with water and um uh my understanding is is that they uh ended up being able to recover from that they they took it and they they dried it all out and uh and then plugged it back in and uh they were able to make it work again but um but it it was one of those things where uh they were like well there's no smoke alarm there's like what's going on like why and then it was just because the level of heat had had risen to the point that set off the alarm um but I I don't you know that's

not a firsthand experience but um Greg had said oh we should talk about the time that happened and uh he was there at the same time and uh we had heard about that Greg you have any Greg are you online you have anything you can chime in on that no so anyways uh yeah that that was that was a pretty good disaster where we had done that but um I could share another one but I'll I'll let somebody else go uh I could talk about like a a complicated complicated one I have another natural disaster one after Fernanda so at the lab at o National

Lab uh this one time they were changing some Transformers for the summit project and so they had to had more capacity for Summit and um I think it was Summit you know memory now starts to get a little fuzzy when you get older um and so they had left one of the doors open from one of the big Transformers that were out and back and a um I think it was a raccoon got uh cold and um instantly became ground and um everything went off everything went offline and um there was um there was a an animal I think it was a raccoon that was

simply kind of charred and still connected when somebody went back there to to look what happened um the entire the entire Center was brought down because of that that's awful so bad wow I've had a home got shut down because of a a rat in the wall chewing all the chewing all the that was like the most delicate way I've ever heard anyone say it and the raccoon got ground and we're all yep we get what just happened oh that was awesome sad for the little guy but kind of funny still I'm pretty sure I've heard of like squirrels doing that to sections of the

power grid and I seem to remember like a site or something that tracks the rather large number of instances of large sections of the grid going down yeah that's fantastic so this mine's not HPC related but I used to work and I sat in a Data Center and we had a freak thunderstorm that came up in Dallas and lightning hit the exact place where three different powers come in to a dat dat Center so two different utility providers providing from two different directions a generator and a battery backup it hit the exact place like a one in 10 million shot and took out all of

them at the same time and we were down for hours it was very quiet and a little odd in the data center for a very long time until they could figure out how to restore power of any kind generator or anything else there was no way to take power from a feed and give it to the data center it was wild nothing happening hours no oh man so if any of you guys watching have a little Horror Story that you would like to share with us please do in the chat we would love to read them and maybe pop some of them on the screen

as well Jonathan yeah you ever broken anything never so i' I'm I'm a note taker so I'm taking notes on the side of stories that I might relay and I'm up to 12 now oh start I'll start with this one which I think is is the most successfully provocative for for ciqs so um we were a relatively early adopter at the site of uh at the time Intel's omnip paath Fabric and I never really felt confident in my understanding of what was going on but my understanding is that um early versions of the the omnip paath kind of software stack there there's the the native

like verbs transport for how it communicates and how most MPI would would use the the fabric um but then there's also like an IP layer for for normal network communication so we were running both uh both compute and our storage over this network and it kept having trouble uh there was something wrong with the IP stack and every now and again the file system would freak out and start expelling nodes and and we'd have problems and we've been through a few um we've been through a few rounds of applying updates and and trying bug fixes and things and we thought we had it all sorted

out and then of course I think it was literally the day before Thanksgiving uh the entire cluster started freaking out and dumping nodes again and I remember sitting there just sometimes I get outside myself and I'm watching the event take place and thinking how absurd it is that this thing that we've been working on for months is coming back up and and causing problems just as we're trying to go on vacation and and as I was looking across the cluster uh just double-checking my work I realized that all of the nodes Had A random assortment of packages from the Intel Opa stack at different versions

that didn't match each other or themselves all over the cluster and uh at the time we were running the whole cluster stateful we were provisioning things with foran and just applying these patches manually to the the nodes you know with a big uh parallel shell or something like that and that moment of realizing that the the cluster was not in the state that any of us thought that it was wasn't in any good consistent State at all not only were all of the nodes different from each other but they weren't even consistent within themselves because the the installer just hadn't done a good job of

being consistently applied to all of the nodes that's what really got me on board with the idea that we we should probably give that stateless provisioning uh concept a second uh a second look and we built something ourselves at the time uh but that's it's probably why I'm such a DieHard werewolf fanatic today because yeah you want to know what state your nodes are in if you want to go on Thanksgiving break surprising that that wasn't that long ago that that wasn't an uncommon problem with the foran line of things that you would try to use the template and they would not always be the

same I this wasn't even really a for an issue uh this was we we had the nodes and then we were just running the installer for the OPA software on all of the nodes and it was early software and you know it's trying to build some packages and install them man you know manually by a script um but if you're deploying to several hundred or over a thousand nodes and you're not noticing that some parts of it fail on some nodes then things just end up in inconsistent State and we weren't being careful and you have to be careful if you're managing all this discreet

state had to build a like 111 machines at one time and I think we had nine different variants using the same exact K Kickstart file it was odd anyway thank you Jonathan yes Forest I know you've broken stuff uh I definitely have but I think we might save that for later in the program I'll start with the kind of the most institutional type um uh kind of issue that I've been aware of I I am I will say I'm much more interested to hear all the all stories mine are very you know I haven't been in this industry as long as yall so I will

um these are the the few kind of interesting experiences I've had one amusing thing was we were putting together a cluster at my last institution uh kind of like the big one that we'd been planning for the couple of years that I was there um and we were getting like procurement and ordering and doing all that type of stuff and like trying to Route all the different uh pieces and parts around and all that we had uh kind of an odd situation where the cluster that we were building was actually like 250 miles away at one of the national lab sites data centers um so

we did not have so this like major cluster that we had a ton of people on was not we could just walk up to the data center or anything like that it was you know Hotel a trip you had to go drive clear out there on a weekend that type of thing um um and we got this cluster ordered and immediately started getting it was supposed to go to um basically directly to the site and it wasn't supposed to go through like Central receiving at the University or anything but because of kind of you know one of those logistical hiccups in these old Cobalt based

systems I imagine somewhere along the line uh what we had start happening was all of the parts for the cluster began to arrive individually packaged in their own boxes at Central receiving and when I say individually packaged I mean quite literally all individually packaged uh every Cable in its own box every hard drive in its own box uh basically every single like Nick component everything we had all we were going going through an integrator to have this all built like like basically secondhand they were going to put the thing together for us and then hand it back to us um because we it was just

kind of a logistical headache to have to have the whole team out there for the extended amount of time it took to put it together um and so like I said it was supposed to start arriving out at the site but instead we started getting all the pieces and parts that were supposed to be put together each individually wrapped and it ended up being like 5 to 700 boxes individually in the end that we had to th and somehow like basically get thrown onto a van and driven clear out to the site and it was just uh it was it was quite an interesting Oddball

like we had we had spent two years planning the whole thing and it was just this odd like logistical headache at the very end of the whole thing like Logistics to mess up two years of planning I know right that happened to us and we actually ended up naming the cluster Ikea there you go perfect there you go I haven't seen it as uh itemized as that before but instead I've seen it arrive at someone's house that's an interesting one that's real so what about what happens to liability and warranty if it goes to something like that and it gets moved you just hope it

doesn't hope great Gary did you have another one um yeah you know I I have a couple of things where they're where things went wrong I I do have um yeah and I'm kind of thinking about which one to talk about uh you know in terms of disasters for us usually disasters are for us is when we're out of production and we don't really have a timeline of of when we're going to be able to get the system back online and um you you know nothing is more um nothing produces more anxiety than when you have a new system and it's not performing and and

that's happened to us a few times and so we always learn how to write better rfps after After experiencing something where you said oh gee if we just wrote that in RFP or they could have tested for this but um yeah I we had one large system from a tier one vendor so I'm not going to name any vendors here but um it it was it was a a a pretty large procurement was like about $5.6 million and it was and the and there were two identical systems going to two sites and so one was going to one site and one was coming to our

site and uh the other site got their system up and running and then and then we didn't and we're like well you know why why is it that they can get this running and why we can't and uh there was a number of problems but um we finally did get the system going and uh but then it wouldn't run linpac it it was it was getting like 50% efficiency and we couldn't figure out why that is and we're how is it that we're only getting 50% efficiency like they they have to have sold these before like has anybody run linpack on these things and um

it turns out no nobody nobody ran a linpack on them and then so they ended up like having to do a firmware fix but essentially this was like the early days of like when uh the Intel processors would go into turbo mode and the the basic problem was that they they would go in the processor in turbo mode but then at different times so some would decide to go in and other ones that decide to go out and you know and you know we tried things with the BIOS trying to get to lock it into full performance really didn't help um and so uh I

I you know I think there's things just can't take for granted uh even even if you're buying from a tier one vendor that that it's all going to work and um so we we we ended up getting that fixed it it was it was there's a big problem you know figuring that out then then we once we got that to fix work then uh okay it would run fine the first time and then when you rerun The Benchmark it wouldn't run and then like why is it that it would only run the first time and and not the second time you reboot the system and

run it and you do it again and you know linpack had run the first time and not the second time and then it turned out that there was a a a memory problem with the infiniband card and so again another firmware fix and uh to fix that but you know so all this is going on and this is like eight weeks from from the time we took the delivery and so everybody's like looking at us like why can't we get the system going um and then finally uh then then once we got the thing running then we did started to do our MP test and

then we're like hey uh the latency we're measuring on these on the uh pingpong test the MPI pingpong tests are a lot longer than they specify on online and then so we met again with the tier one vendor and the infin ban vendor we showed how we're doing the measurements and they all took notes we're like meeting like every day at this point and then finally what happened is they they came back and they said okay you guys are right it doesn't the latency isn't as good as we advertise and your numbers are the right ones and so when they did that then they then

you look online and all the latency numbers that they had advertise for that infinance which came off came off the web and uh and and they actually ended up giving us something else instead but um I I think for us it was a disaster just because you know if from beginning to end we probably spent 10 12 weeks trying to get this system going but in the end it ran really well but I just cannot believe that you know uh something is out there you would assume that somebody has tested it and and uh you can't take that for granted absolutely and I I have

not thought about turbo mode in so long when you say that it makes me feel old again yeah so we we write our our you know our you know must run in turbo mode consistently you know all the time but you everybody's um turbo mode uh differs and then there's also a very fine specs in the Bios in terms of like you know how much power is feeding the memory and stuff like that all of it actually affects uh latency down to the Fine detail in terms of how fast it spins up so even something that sounds the same between two different vendors May perform

very differently very interesting thanks Gary hey Adam thanks for watching one of our vendors got hard bit by AVX UNT Turbo decocked on our Haswell base cluster during acceptance testing it was something else happened to everyone that that was pretty close yeah thank you good thing you guys can read it so Fern the da system that I saw come up around that time just because no one anticipated the different clock speeds for ABX 512 sorry brose no totally you feel his pain it's good it's good Fernand I think you had a few um uh stories that you wanted to tell but if you wanted to

like throw in there a little bit you guys I'm very curious because Gary you mentioned anxiety coming up and like we all kind of deal with stress and dramas and traumas in the work and something taking 10 weeks that should have taken a day we deal with that differently so if there's any like fun little anecdotal you know stories that you can tell about either yourself or other people freaking out what they do when they freak out I want to know it's it's going to get resolved eventually I think that's the thing is after you do these a few times then then then you just

have to tell yourself okay it's it's G to work out it's you know something will we'll come up with something and uh and uh yeah I I think there's a lot less of that now because a lot more technology has cons converged so that people are using a lot of the same things but you know at one time people are using different interconnects they're using different processors uh they're mixing and matching stuff and so I you know I I've heard someone say that every time you build a cluster is like serial number zero um because it's it's just a brand new combination of things with

including once you include the software that's true and so um so you just have to explain to people as best you can uh that uh breathe you have to remind them to breathe right yeah yeah and then warn people that that you know that this there's something of an art to doing this um you know what I was going to talk a little bit about is is that um if you're a one if you're a small company and you buy a cluster in his TurnKey you know that's one thing is probably going to work fairly well for you but you know for anybody running HPC

yeah disagreement here anybody running HPC where they're getting systems over time they have to integrate it into their environment then they're running their own stack you know they're you know how they do things are going to be custom so um there's there's going to be a a lot of unknowns once you integrating new equipment into an existing en environment so that's that's where it's really important to understand that and uh uh write your RFP carefully so that this way you have some recourse and pay for support yeah yeah something I've seen is that like nothing makes me more anxious during a disaster or when when

I'm anticipating something going wrong than when I'm trying to do it myself or when I think that making it good or fixing the problem is all on me and think about this the most with a time where you know we we were doing a storage procurement and it didn't go well uh the vendor wasn't delivering in the way that we had had written in the RFP so that was good um but at the same time the the previous system was coming off of support and it wasn't really possible to renew it um while the new system wasn't available to migrate data to and uh I

for a long period you know a week or or two um was really anxious about what are we going to do we we I felt trapped into this contract that wasn't delivering uh cuz we needed somewhere to put data but that also didn't seem like a good thing to move to and then they just realized you know we've got this whole team here not even just the sist admins but this whole team of application Specialists and other HPC people when we just got into a room and we had a conversation and uh the solution we came up with was you know we've got the HBC

scratch system and maybe it's scratch but at least it's under support and it will be for several years so we carved off a part of that and life booed the data onto that while we sorted out the the actual campaign storage and that's something I never would have considered on my own and uh making it a problem that we were all there to solve really helped alleviate the anxiety around that whole situation it's good one Jonathan just all right Fernand I know you have several stories about the days the early days of uh turbo and hyperthreading was also a very controversial talk topic because we

had uh software level hyperthreading and then Hardware you know uh hyperthreading and that was many many discussions uh about whether or not those things should be left on um but the first the first job I had we had to do a an upgrade so I was my company was one of the first adopters of um bright computing cluster management uh tool I forget what it's called but um what was great is you could create these images and then you can deploy these images to the nodes and that was like the easiest thing ever um coming off of what I was doing in college which is

using the old rocks cluster distri to to update our clusters um so we get to the weekend and I have an IT counterpart and um he was originally like a Windows admin and then had recently uh you know learned Linux and so was I was kind of helping him learn Linux but also he had become quite self-sufficient in it but he had a lot of experience in just a data center so this is my first time walking into a proper data center you know with real cabinets and real like you know procedures we drive and we had moved the entire cluster to like Minnesota somewhere

I forget where it was and so we're there on the weekend and we're trying to just essentially install the FL the first commercial cluster I bought straight off of Dell um we had bought 20 nodes and we were putting it all together and I created um two two networks so one was like a gigabit Network and the other one was an infinite band Network so one for compute and one for backend plus some file uh communication and the the infiniband one the infin ban cars this is the first time I'm seen so remember I just come out of college right so like college was straight

up ethernet in a closet somewhere um this was like fancy stuff for me and uh and I had no idea what I was doing but we were installing the stuff we're putting it together and every time we would you know try to bring up the system half of the nodes would drop and we just could not figure it out so we spent two days I barely slept uh trying to figure out why and we finally figured out like we ended up just I said you know what let's just get all the serial numbers together and I like maybe this is a batch thing and a

good job young Fernanda figure this out because I was like maybe it's a batch thing so I put all the serial numbers in a in a you know Exel Excel spreadsheet and it turned out exactly that that there was like half of the batch was a set of serial numbers the other half was another set of serial numbers and what we figured out is if you don't know about infin and the cables themselves have intelligence at the end of the cable and those had a different firmware from the actual cards that had been shipped with the system and so they were not communicating and we

would just swap a cable out and then put another one in and we only had a limited number of cables and then as soon as we move this cable to another place that node would you know drop off the the cluster so that was my first like wow this is a really fragile system uh to put together and then when I moved to idge you know we're we're not doing 20 nodes we're doing 20,000 nodes right so this is to put that machine together what took us two days to figure out two people in a room imagine that when you're trying to do acceptance on

a system like that but and then the personal story that I had the other one that was kind of short is uh I was in charge for a little bit about a year and a half at oid to uh do the um the the software modules in the system so you know when you get to a supercomputing center HBC Center you can do module ad right whatever your your uh piece of software you want to run and um it was one of the first times where I had to recompile everything because we had an outage on purpose to do the upgrade and then all the

drivway were updated and all the compilers were updated everything else uh system kernel Etc so I had to go and recompile the application so I went to work and I stayed up until probably like I think I I stayed there all night and then the next day people were coming in and that day I was on call um and I was really clever I had t-s I had open you know eight different windows we had like eight different I think login nodes at the time time for balancing and so I was logged into eight different nodes and I was just building application each of the

nodes um as as I was going along so make J dash I think it was -8 that back then application one's running I've got like lamps building here and something else building there the next day uh they bring the machine and the Machine has now been up since you know they release it to me the next day we say okay we're going to release the cues everything's going to come online and not even 30 minutes in and people are like I can't log in I can't the whole thing is down and and we're like we're trying to figure out what it was and it turned

out uh that finally uh somebody figured out or one of the other people in the the place um the support team figured out that I had filled up all of the uh RAM disk temp uh shares or you know sltm with all of my compiling and all of the artifacts because I would stop in the the middle of of compiling go like oh I forgot to you know make sure that this flag is on or and so I had filled up all of these sort of artifacts of Compilation All of the temp drives and all of the login noes have been filled up with all

of my crap and then when people were starting to compile their own they would run out of you know space and then it would stop people from logging in so that was uh my first I really screwed that up uh stopped an entire HPC Center until we could clean those up so but it turned out okay it's fantastic that's I can't believe that it would let you there are no um limits or cleaning up after you that's F well there were limits and I and I was cleaning up some things but but not everything I clean up because I was stopping the make in the

middle of it making nice Forest came off mute well I'll tell this story we were doing a little Cable Management cleanup this is my story of you know messing up something at the HBC Center uh we were doing a little Cable Management maintenance one time on some of the systems uh that we had for this cluster like the original one that I was working on and uh basically there was just a few different like cables I had found that like seemed like they they weren like going anywhere um so I was able to kind of start tracing these back I found a few of them

I got them unplugged like that one of them basically just went into a gigantic like ball of cables and I thought I traced it out there was no tag everything else had like tags on both end of it like you know this is you know whatever like the host name that type of stuff there was just one that had nothing on it when I traced it back to the ball of cables it seemed to just kind of like terminate in this empty like just open dangling connection so I was like okay okay you know I asked sudden can I just cut this one so I

can pull it out from this end take the other line pull it out from that and she's like yeah no problem I cut that and about 10 minutes later I noticed there's a GPU zero doesn't have any like connectivity light on it so like so I trace back the other end of the cable as I'm you know trying to figure out well hold on maybe I need to go you know look at this a little closer and eventually I figured out that um the other end of it actually LED back into GPU zero so it was an easy fix we were on ethernet at the

time for those we didn't have like those on any type of infin ban backbone this I'll qualify this and say this is when I was very new wandering around the data center um but uh like I said that was a pretty easy fix because we were just on ethernet so swap one cable for another still pretty funny it was it was I just basically had to go I was like go I went to the chief side and was like okay so that was actually the cable for GPU zero and she was like oh a bummer thanks for letting me know go swap it out all

righty so nice yeah Grand times were those the actual words that she used oh bummer yeah actually yeah she was very understanding it it was a good yeah we that was a lot of fun that was it was just her and I working for a long time we also had like the incident where very early on when I was was there I went to go SSH into our head node and I got you know basically no route to host I couldn't get in on SSH and I didn't realize that you know at the time because this is I'd only been there for like a couple

of months that that was like instant alarm and about five minutes later the chief say adman Peaks her head in the room was like can you SSH R too she was like oh I said oh I was just about to come tell you I I couldn't get in just a second ago and she's like oh well that's not good we need to go figure out what's up and we found out in very very short order that most of the downtown had lost power uh but the problem was that because of how the data center was set up something along the lines of like the UPS's

could potentially stay on for a certain amount of time even if the cooling went off and so you could like overheat the center now thankfully after we raced down there and kind of did the fire drill to see what was up the other sis admin who was in the office next door to the data center had already shut the whole down um so thankful he was there but it was uh could turn into a Gary overheating the data center and spraying water everywhere yeah exactly we um yeah we ended up there for like 10 hours afterward because when it brought down the cluster it some

of it wasn't on UPS's so it brought down all the switches and some of the stuff like that um this was like I said a very small University cluster so it kind of you know we're doing what we could couldn't you know keep things online uh we lost like the switches and all that and one of the switches uh had not had kind of the configuration saved in the way that persists after it gets power cycled so the switch when it got cycled lost the configuration and we had to go call the engineer who had previously worked for us and was working somewhere else to

come be like what did you do with this switch that made it work because it was him surpris for us how often that happens on switches the configs don't get made persistent it's crazy how often it happens yeah so we lost that and that took quite a while to get back running but like 9 10 o'clock that evening we had it back go and then outage resolved uh we yeah that that was that's my most interesting like natural power outage disaster type story Jonathan you had 12 you have to be down to 11 what was this about a solars yeah story is entitled when a

solar eclipse took down our HPC Center um so it I I looked it up I think that this was in 2017 but there was like a a a full solar eclipse and everyone was excited and um my wife was bringing my kids to the park that was across the street from my office and it was going to happen like right over a lunch hour so I was going to take off and and head over there and watch the eclipse with my family and just as I'm about to leave naos blows up it's saying not only that like the login no one could log in and

we were getting alerts from within the cluster that parts of it couldn't talk to each other either like what on Earth is going on of course the same way of course it would happen right before Thanksgiving or right before I'm trying it's always breaking right when you're trying to go somewhere um but I trace it down to DNS something's wrong with DNS and uh we were pretty proud of our DNS infrastructure at that point because we had managed to convince the campus it Department to delegate actual DNS to us we didn't have some fake internal DNS domain that we were based off of which a

lot of of HBC clusters even at that scale are um so I went and I I got in touch with uh the campus DNS guy like hey you know something's wrong with DNS what's happening and I find out they're also you know logistically on fire as well the only message I get back from them is it's because of the eclipse and at that moment I just I gave up I'm like whatever there's nothing I can do about this it's clearly not my problem I'm going to go watch the eclipse with my family sure enough we do that I come back when I get back everything

is fine and yeah exactly what on Earth could possibly be doing this find out that because we were using not only the campus dnf DNS like name servers you know we were delegated from them we had our own name servers but we were using the campus resolvers to get back to ourselves uh so we would use the the real campus recursor and then that would hit our own uh name servers uh people live streaming the eclipse on campus had DDOS the campus uh name resolvers and we had not considered that perhaps we would be even more able to stay up than the campus DNS infrastructure

so they became a single point of failure for us and DNS went down for us and we couldn't resolve any names so nothing could talk to each other so our takeaway from that was even if you integrate with the campus infrastructure still maintain your ability to operate without it that's so it didn't run on solar wind because I was hoping that would be the answer that's a different [Laughter] story I have another overheating story let's have it yeah we there there was a times um actually wasn't that long ago in the future where in the past where um we had it was towards the end

of the day and there was a leak a water leak in the parking lot and so our facilities people came it was at the end of the day they're thinking okay what we'll do is we'll shut off the water to the to the building and then we'll deal with it tomorrow and so they did that and then they left and then it wasn't until about 1 o' in the morning that all of a sudden all of our alarms went off and what had happened was that when they shut off the water it also cut off the water to our cooling towers and that was about

the time 1:00 in the morning was when all the water had started had evaporated out of the cooling towers and then it affected our cooling so uh nobody thought about that when they shut off the water and it didn't occur till like several hours later it was that almost sounds like the old uh why do today what you can put off until tomorrow yeah just turn the water off we fix it tomorrow exactly love it I have one that's a a college one so in the early days of sun grid engine this is not so much as a disaster for us it was a disaster

for the entire campus so campuses have uh sometimes a condo model where a professor will get some grant money and then put in um you know nodes into this uh Center and this cluster and it generally is a you know it's sort of a homogeneous uh you know they'll have homogeneous partitions uh we figured out that if we used Sun grid engines um it was like a some sort of like grid that would then spawn different jobs we could beat the system that kept us from running too many jobs based on how much we had contributed to the condo model and and so one time

once we figured this out all of the students in the group which is about 12 15 had essentially flooded the entire cluster spawned all of the jobs and beat the system because of the way that it was set up uh to keep us or prevent us from running all of the things all of the time and it took them a little bit be you know before they could figure out and all of these jobs were like two week jobs that were run and nobody really noticed because a lot of them were named different things and nobody really paid attention to what was spawning all of

these different jobs but um yeah set up your cues right was a lesson there college kids to game the system always college kids gotta finish our degrees man it's true it's very true I have one more water one here and uh theme Gary it's all about water yeah the uh because we're because the theme is been water since the the fire sprinkler one but uh down at UC Berkeley the the data Center on the third floor and um and they did that purposely and because the previous one used to be in the basement of Evans Hall and um which seems like a safe place to

put a data center I'm sure a lot of companies would put their data centers in a basement and uh it seems like a good place until somebody sets off the fire alarm and the sprinklers go off on all the upper floors and then where do you think all the water goes Rains Down wow there's no like you would almost think they would have thought about that and put some sort of barrier so that if that did happen it's a good idea I have flooded a three-story home this has nothing to do with HPC it was a toilet and a cat but that's for another day

high performance grounded a performance cat It's just sometimes you don't know what you don't know and of course looking back you're like oh like all all of you have these stories where it's like oh you know this happened and and then we figured it out and it was this you know and you're like oh okay yeah like you know okay that that makes sense similar type of don't flush a cat down the toilet do not flush the cat down the toilet also do not flush the kitty litter down the toilet oh no I didn't know and then I sounds was so bad it was so

bad and my my aunt and my over my was my cousin and her whole family they were out of town for like two weeks and I was just supposed to like feed the cat and clean the litter box and third floor up I'm putting the kitty litter poop in the toilet and flushing it it backs up and it starts to link well you don't have to go every single day so two days days later I come back I walk in her house it's literally raining inside of her house all the way down three stories out the garage did you get invited back no no no

I been disowned that was scary thanks for listening guys I wanted to participate absolutely Jonathan we don't have time for 10 more stories I'm sorry actually up on time we're going to have to do another one of these because I feel like there are there are a lot of stories to be told so we're going to have to recycle this topic again too lots of hopefully by then this group doesn't have new stories to add I think every I hope everything stays smooth fand is laughing again lucky for everyone I'm no longer touching the keyboard so wow I would say congratulations but I don't that

I miss it uh I know I'm in that like really midcareer having that existential like do I want to do the up do I want to stay here do I want to go back to the keyboard uh are people gonna believe me that I even know what I'm doing do I even believe myself that I know what I'm doing uh I'm I'm right in that right in in the throws of my mental health breakdown going back to the mental health issues we believe that you know what you're doing so don't ever question you're appreciate we know you keep you keep allowing my video to be

shown so I think you do absolutely absolutely Rose I support you getting out of management anyway you guys I am so grateful love I love my people I love my staff that's fantastic okay whatever you do it's gonna be amazing you are val you are valued wherever you go there you go there you you go so you guys thanks for coming and sharing your stories with us we have had absolutely so much fun thank you everybody for watching wherever you are at thank you for listening if you are not listening to our new podcast that's right and we would love for you to like And

subscribe and share and do all the amazing things we'll be back here next week same time same place we will be at sc23 in Denver Colorado we would love to meet with you so if you would like to meet with us either come by boo 1855 or reach out to us on our website um even just like the info atci iq.com that will come to me and I will get you all scheduled up you are amazing we love you so much have a great day thank you [Music] everyone

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert