
Part 2 - Storytime: Tales of HPC Disasters, Recovery, and Resilience
Ever wondered what it's like to face High Performance Computing disasters head-on and emerge stronger than ever? Join our exclusive panel of HPC experts as they reminisce on their unforgettable experiences, share their epic tales of recovery, and impart invaluable lessons learned along the way.
Whether you're a seasoned pro or just starting your HPC journey, this webinar is a must-attend.
Speakers
-
Zane Hamilton, Vice President of Sales Engineering, CIQ: LinkedIn
-
Rose Stein, Sales Operations Administrator, CIQ: LinkedIn
-
Gregory Kurtzer, Founder and CEO, CIQ: LinkedIn
-
Forrest Burt, HPC Systems Engineer, CIQ: LinkedIn
-
Allan Sill, Managing Director of HPCC, Texas Tech University: LinkedIn
-
Jonathon Anderson, Solutions Architect Manager, CIQ: LinkedIn
-
Chris Stackpole, Advanced Clustering Technologies: LinkedIn
Transcript
[Music] oh [Music] [Applause] [Music] a [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on powering the Next Generation eration of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and Apper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source Zane came off automatically this time I was kind of waiting too I was like wait I I was hovering I was waiting how are you Rose hello sir how are you
doing well I see that you've got a little coat over your shirt as well today and I'm scarfing it I hav looked I know it's um oh it's warmer than I thought it was it's 58 it's 65 and I'm so how is this even possible how are you warmer than I am oh well all right oh man it's good to be here so ah wait disasters caution Zane caution we are here for our round two of the disasters in HPC and of course we like to bring it around right because like talking about drama and you know commiserating about like all the things that happen
all the things that go wrong super super fun haaha we can all then like Bond and relate about that but then we added another part to this little piece today where we're going to actually talk about how to recover from the mishaps as well and so I think this is going to be a very wonderful well-rounded experience I might just keep this maybe I I think you should definitely not that's actually very distracting can't like covering half your eyes you can't see it's like kind of funny right I was like all right I got I got to do something something I gotta do something I
just like the drama statement as you fall out of your chair did you good I like that okay I might keep my day job and not go into acting not going into acting okay well just know there is a backup plan right it's a backup plan awesome well good I think we have um some friends that are GNA join us today so let's see oh Forest 2 and Greg curzer Jonathan Anderson Alan sill this is Magic IAL this is magical thanks for being here guys so I don't know do we need introductions I feel like everybody's been here enough we probably do not I did
View full transcriptHide full transcript
twist Greg's arm I don't know if anybody remembers who that guy is it's been a while vaguely vague vague I don't know I'm kidding Allan you'll be pleased to know I have my have my cup hey Texas Tech cup yep right come to Super and get another one awesome stack welcome back man it's been a while what's up man it has thank you how are you busy I hope not with disasters no fortunately no that's good that's good all right well it's been a while since we've seen you so let's just do a quick introductions all right why don't you introduce and let us know
who yah okay sure uh yeah I'm stack uh I am an HBC admin now for a long time since we were just talking about Texas Tech that's where I cut my teeth uh I got pulled into the student program under Allen so uh and uh have been doing HBC ever since I've now worked for several different organizations and currently am uh working for a company where we build uh hpcs for people uh and I help with everything from design to uh troubleshooting and maintenance and things of that nature um also on the rocky uh Linux testing team absolutely thank you for that I think that's
a nice little segue into Mr Allen well stack hold on I did not know you worked for Allen yeah long time ago okay so at some point we need stories story time yeah very very proud to know that actually to be reminded it's um so alen s u managing director at high performance Computing Center Texas Tech uh co-director of the NSF industry University Cooperative Research Center in cloud and Aon Computing tilter at windmills um namely U trying to get them powering data centers uh and many other um projects um I'm just here to see if I can get some side time helping build the Raspberry
Pi versions of werewolf and uh and kernel modules I like it I do like it Jonathan you're the next one on my screen so I'm just going to Point At You hey everyone um Jonathan Anderson Solutions architect here at ciq uh and I have a list of disasters to talk about if we get to them that's awesome all right we'll do four that's my main credential right now is i' I've participated in many disasters have you caused any of them now I have at least a couple that are my fault on this list okay good Forest BT good morning all my name is Forest ber
and he froze that's it that's that's all that and he's on mute now you're on mute you're does this have anything to do with the coffee that got spilled maybe today in our live demo of disasters R this morning sorry my um my my internet connection everything is kept online by this thing being plugged in and I accidentally dled it so I reset everything my name is force bird I'm Solutions architect here at CQ I work with Jonathan on our uh Solutions architect team and uh I have seen a few HBC disasters in my time in my uh early days of being in HBC I
almost caused a couple of disasters thankfully averted normally by the uh the more experienced CIS admins at hand but uh yeah I've I've seen some stuff but I'm excited to hear from everyone else awesome Greg since you are on a a short time frame here I guess we we'll start with you if you have a story if not we're GNA go to Jonathan's list um yeah I've broken a number of things um yeah in my tenure as being a HPC uh administrator um the question is really just which stories do I want to embarrass myself the most with um the the one of my early
stories is and this is something that I I really hope other people have done um remove minus RF on a production system um with a misplaced space uh so trying to clean sltm slash another directory and I do not I swear I do not know how that space got in there after the slash by the time I noticed what I typed I was control cing as fast as I possibly could over and over and over and somehow amazingly it was fast enough that it didn't take out the whole system but I did lose almost all of Slash bin you know it starts off alphabetically now
this is luckily before uh somebody in in some position somewhere decided to sim link bin to user bin and SLS bin to user sbin and lib to user user lib and Etc um so Ben was actually a separate directory um I think on this system it may have even been a separate Mount point I don't remember but um I was able to clug it now it was really hard to get it back because uh you can probably imagine that a lot of kind of core important programs are in bin so um but I was it but rsync fortunately was not and rsync doesn't rely on
shelling out to any commands so I AR synced from another system running a similar version and then reinstalled all kind of the core RPMs on that system luckily even though there was well over a hundred people logged into that system using it in production there was not many people who even noticed so I kind of got away with that I'm glad Gary isn't on this call because he was on V1 of this call and he was my boss at the time so I'm glad he isn't he isn't here now um but I have one more quick thing then and then uh I'll hang around as
long as I can maybe I'll have an opportunity to say to say another one but um a lot of disasters that I'm familiar with didn't actually happened during production they happened pre-production um specifically on ordering massive millions of dollars worth of Hardware only to find out that uh the operating system SL kernel SL pixie boot stack that you wanted to run on this doesn't work now a while ago before pixie was standard on nxs some of you um as gray or grayer than me may recall ether boot which was the successor or the the prese pror predecessor thank you I need more coffee I really
need more coffee was um the predecessor to uh to pixie and it used to not even come on the Nyx you could choose to flash a ROM that sits on the Nyx but it was very temperamental um couple that with another project that I had percus a while ago that did K exex into the operating system um you end up being very Hardware dependent so to spend millions and millions of dollars on you know hundreds if not thousands of compute nodes all at once only to unbox and try to set up the first to realize that oh crud it doesn't work and then you try
the second this doesn't work either and then you look down the row of systems that you just bought kind of with your your heart like beating and stomach feeling a little off and it's like oh no um it taught us some really valuable lessons that I hope people here listening on um learn from specify in your RFP in your RFQ or whatever you're putting out there from procurement make sure you're specifying exactly the software infrastructure and the operating system and tools that you're going to be using and make sure that that is part of the acceptance criteria for whatever systems that you're buying so um
this way if it doesn't work it's not your butt completely on the line the V has to come in and help fix it or something um but that has been a massively helpful thing that that um that I learned early on after dealing with again systems that just didn't boot or didn't operate right um Etc is make sure you specify the entire software stack uh as much as you can that you're going to be running uh it also helps open- Source projects like werewolf and Rocky and other things things for vendors to hear that this is the software stack that needs to be um needs
to be compatible with whatever Hardware that they end up specking out and shipping so it it solves a number of problems kind of all at the same time but uh the biggest one is just to make sure that you know as the person responsible for integrating a lot of these systems you don't want to be caught you know with millions of dollars worth of worth of Hardware that just doesn't work with the stack that works everywhere else I like how stack is laughing the whole time yeah I was I was about ready to say as well um you know stack and stacks companies specialize in
a lot of this integration and whatnot so they just you just know everything works um so there's a good reason to go with with what they're doing um but yeah there there's you know so stack maybe maybe you want to plug a little bit of what your you know your company does and how how it works and how it solves these sorts of issues uh it solves these sorts of issues by that being me and a couple of the other guys that figure all that out before we sell anything we end up with all the new hardware and we figure out oh well that's not
going to work um but yeah that's basically what we do is we uh only sell stuff that we know Works in HPC we dedicate to HPC so all the software stack and everything uh by the time it arrives you've already got it you know it works yep very help even more than just the hardware like you can have a firmware version that just throws everything off like yes I mean there there's a lot to these systems and again to put if you're doing a a procurement through a through a um you know a more like say typical ihv you know just basically selling your boxes
put as much in there as you possibly can um or use a company um uh that that does this for you and and stack works at a great one so there you go yeah firmware bites a lot uh just actually had that two weeks ago with I won't name them but a very large motherboard vendor um where their pixie stack was just busted um had to get them to fix it yep I was talking to someone else I was trying to get them to come on and that was their story uh couldn't make it to this but they had done a firmware update on some
storage and it actually caused the storage to start overriding itself so oh I can't even imagine the horror stories and disasters around storage um does anybody else call luster scratch for a reason I call it we had a whole story about that last time yes yeah I gave a whole story about that last time um including bit level editing of Elder skes it's bad so since we're on this storage one uh wasn't fullon disaster but Mand did it caused problems for a while um but we were early adopters of Seth and by early adopters I mean we were running sefs in production on Firefly when
even the lead developer was saying this is a bad idea um but it was the only storage that was really trying to do what we were needing it to do um and so um yeah we had a user who discovered that a particular applic that they used um would write its database to disk and uh it would write hundreds of thousands of less than 1K files uh and that did not go over so well with a distributed storage system like stff trying to keep up with all those files and replicating them and um it was uh that was the first time that uh stff really
got hammered on our environment and caused a lot of problems uh we had to redesign quite a bit of stuff for that one application until we were able to actually work with stuff and be like hey we need some help on this and even then um that many tiny files coming from a parallel whole lot of different computers all at once I mean a lot of file systems croak under that one but um we almost lost data because Stu was croaking so hard uh under that in the early days that was rough sound like a load test to me it's just done in production I've
got another load test story actually um we got a new system in uh we did our uh for various reasons the way that our power was at this particular location um we didn't have rack level power what we did was we had those little oneu bricks throughout the cluster um and we had done testing for those to make sure that they should in theory be able to handle it um what we failed to do was test what happens if one side completely drops off um while Under full load um and The Power Team forgot to do their load calculations for when all of those power
bricks are in use and so we had done a bunch of load testing beforehand um but we had one user who came in that was uh the only user that was using pretty much every node they get their hands on at one time under heavy CPU usage for long periods of time and what was happening was uh so the power would pull from both sides A and B at the same time but occasionally it would spike on one and when it spiked on one with enough of them that whole system would overload and flip which means all the load shifts to the other side which
would overload and flip and to have multiple groups of these nodes do that all at once would cause all the main Breakers to flip and so uh we get the right user running uh the right job and all the power the whole area just gone um so that was fun to figure out what exactly was going on because again when we would do individual testing it was fine it was just on the larger scale size is where it would break and be like oh yep okay we've MISD done our calculations we need to have more power uh both in our lower groups and on the
backside too so um yeah that was a fun one to uh come in multiple times in a week and be like why is everything down it's awesome cascading power failures by load test yeah Power failures enter a lot in many disaster stories uh uh including some we covered last time I I feel like many of the examples have been giving are experts disasters like um you have to know what you're doing to get into this kind of trouble or at least you have to know a little enough to be danger enough to be dangerous but um I think we should cover like more Elementary level
things that go wrong um things that are just um big Goofs from the uh start um oh and by the way if you know of a goof that you've been affected by or that you've carried out um the the HBC docial uh uh group is running a their virtual noodles award um uh inspired by Vanessa wanting to dump a bowl of noodles on um a certain unnamed developer's head when she encounter a problem uh so she decided to formalize this and um and she allowed people to nominate um certain vendors or certain um uh or themselves for getting a a virtual bowl of noodles dumped
over their head and and voting is going on now so uh look look around in the HBC social uh uh slack or um Discord or masteron you'll find it um so with that aside some of the the things are just big you know Simpsons dull mement moments um that um that don't require such sophistication to get into um you know I'll give an example that um we took over a the a chemistry cluster that they had bought with their own NSF grant money but then when we my my Center opened a um stack you may remember this when when we opened one of our new
clusters about six years ago and made it available for free um then that just sucked all the uh the support out of this other cluster because the chemists thought why should we be paying our research money to maintain our own cluster when we can get it for free from the center assistant so long and story short we we inherited this cluster uh but it had not been designed well from the start I mean the cluster was fine but it was in a room under a lecture hall it's more of a mezzanine and the the UPS system was in the room with the cluster uh which
you know you can get away with if you plan your cooling well that's a big if they had not the cooling system was uh based on refrigerant by the way which rather than chilled water or a handlers or anything actually was a fancy air conditioner that um sucked U heat out of the room and dumped it into the chilled water system 40 feet away through Refrigeration pipes and of course it was not connected to a a backup power or a generator at all so the sequence goes like this you lose power that's the beginning of many of these disaster stories um the um UPS kicks
on so the cluster keeps running um you see where this is headed uh and the cluster keeps running so it makes heat along with the UPS which is in the room and discharging and making heat and you know a couple minutes later the fire suppression system goes off and because this was near an office space that there was an expensive breathable uh fm200 system so there's 10,000 bucks right there to refill that um you know so it all comes from just not designing it right from the beginning right so sort of like Greg's story Greg sort of was an oop story but it was a
sophisticated one this is a very basic one uh if you're going to put heat into a room you have to take heat out of the room even when the power is off so putting a UPS system to keep your computers running is not the right approach put the UPS system and to run the cooling system if you have to have one all right some of that's an afterthought like you buy the equipment you put it in a room and then it's oh and we have to cool it oh absolutely I'm a particle physicist and there's a story you know that the the engineers view of
how particle physicist design detectors is that they just float in the air several hundred tons of steel and electronics and the the the physicist never draw any sort of support structure when they when they sketch the the detector right so this is the the sort of computing equivalent the cluster will just so to speak float in the air and all the power and cooling will just magically appear when needed the spherical HPC cluster in a vacuum basically Jonathan hit the list all right so uh since Alan was calling for like newbie level mistakes or or early things where maybe you're you're doing something you shouldn't
have been doing in the first place I'll tell a really early one I I don't think I was still an intern at this time I think I was actually an employee but for some reason I had been tasked with setting up um it was an email Alias of some kind either for for a monitoring Alias or maybe a new mailing list or something like that and the way things worked at this lab uh we were kind of a splinter group uh within the the m mathematics Department of the lab but we had to integrate with uh an upstream it Department system and so we we
maintained a list of email aliases that we wanted and these were all in a big subversion repository and you define your Alias and then you run a script that goes off and combines it with other aliases and uh and you know pushes it into the Upstream email system so I had done this and as part of doing that I I had an observation and I thought I'm going to be helpful I noticed that there was extraneous new lines at the end of this file it's like that doesn't need to be there that's just like editing Cru so I cleaned up the file and saved and
committed and everything was fine and uh I I remember someone remarking that they had had a really productive day because email had been really quiet all day and it was about you know 3: in the afternoon I had done this in the morning about 3:00 in the afternoon the people ized no one was getting any email and maybe that is wrong maybe someone should be getting email and so of course the way that this Alias file was being combined with other Alias files was assuming the presence of some new lines at the end of this and instead when it had concatenated them together had created
a syntax error concatenating the end of our file with the beginning of someone else's and just crashed the email server and no one had noticed now I will not take responsibility for the fact that no one had noticed but uh it did mean we have a more had a a more productive day uh without email that day uh I will complain part of this is recovery uh that is in my mind the worst case of the way CIS ADM sometimes gets done is the solution was to add new lines back to the file which is fine in the short term but uh kind of going
back to what Allan was saying earlier about problems happen when things aren't designed correctly in the first place like you shouldn't just assume that your input data is correctly formatted and it there should be tests in a script like that and it should we should have when having experienced this problem uh corrected the script and made it more robust uh so that is that is my lessons learned from that kind of thing but I was months into my first job so uh not in a position to push for it at that time just little Bobby tables Jonathan little Bob tables sanitize your input anybody I
think this predates that comic I was gonna say if anybody else gets that reference um post it let's see who gets that first all right Forest I believe you've got one that you've want to share you have been so patiently waiting let's hear it I I just have a few like kind of small elements mistake stories not all mine but all hopefully some amusing I'll start with the one that was mine I was a freshly minted uh rude admin on my first uh HPC cluster a little while after um joining the site that I was at and uh we were using bright cluster manager at
the time and bright cluster manager at this point you uh basically it did all like the Yum updates and it had like all these package repositories and stuff like that that have been add and andall software from do all that it essentially completely kind of takes control of the Yum side of things so I was going to do something like install something on the cluster before I had been before it had been explained to me how to do that and I had queued up a pseudo yum install something or another or pseudo yum update something along the lines of that and as I'm sitting there
about to like you know pull the trigger on this as it's literally like is this okay yes no I hear a voice speak up from behind me and it's just as quiet don't do that I'm like what and uh the Like Chief uh not the chief Su admin but the other HPC engineer who I was working with at the time just kind of gently says you never yum update a bright cluster manager server or cluster that will completely mess everything up and we'll have to like completely take it down roll it back do not press yes to what you're about to do there I'm like
all right thanks for the heads up so that that when I was very very new that was my that was kind of my quintessential I almost nuked the entire system mistake but thankfully it was a burden um rock was like that too Rock was Rock's cluster system was exactly that way you could not y updated it's a it puts puts one in mind of the I think there's a Linux magazine cartoon with the older assisted men talking to the younger assisted men the younger assistant is saying it's a virtual file system so and the older assistant said what did you do yeah yeah we almost
had something along those lines going on there it was uh yes that was pretty funny we also had um another small amusing incident that we had we had a 10 gigabit pipe back into uh this big optical Network that was kind of serve the whole region that we were in um it kind of like links all the universities the National Labs stuff like that all together via like this as we found out 10 gig total pipe um we thought that our leg of it was basically like 10 gigs and then like the backbone was much more it turned out the entire like Regional Optical Network
that we were on was like their Max bandwidth was 10 gigabit uh and so we started up our test and we're you know getting like a full 10 GBE throughput on it we're real impressed we're like wow this is we're really connected to the network this is working great and suddenly the phones start ringing and it's of course the regional Optical Network people like what what's going on over there the whole we're getting calls from everyone else saying it's saturated what where's why is the whole network being used by your guys' place and we're like oh we're we're running a bandwidth Jack we'll go ahead
and cancel that um and so we after that we knew that our you know the limits of our Optical network but it was kind of funny to get the call from them uh about it I also heard kind of a users's gone wrong story uh a little bit after the fact like the end result of a user that uh we'd always kind of had problems with one of those people who's very smart very like competent about HBC you know he's on like our HBC Advisory Board stuff like that but is always kind of pushing the limits on the storage trying to use up as much
of the cluster as possible you know Etc kind of pushing the limit well eventually uh I kind of heard that he started like using scratch storage filling it up it got to the point where the rest of the cluster was starting to get brought down and stuff um and it was just kind of interesting after like years of dealing with this one user suddenly the whole thing the relationship kind of devolved and um started like filling up scratch and things got canceled and he got kicked off in the end we found uh you know people running vs code servers on our head node at one
point and that really enraged our Chiefs sadman um so there a lot of small you know we think of these massive disasters and stuff in HBC but the every day you know just kind kind of static and churn of things is often particularly interesting as well especially when you have people that are trying to do trying to do things just to push the limit you know breaking out of their GPU jobs things like that you know every once in a while you find that like a research group has kind of figured out a little hack into your cluster and like I said you'll suddenly see
that like a whole group of them from one lab have all got VSS code servers running somewhere you'll see that one group of them from one lab have all got like jobs that have kind of overran some of their uh um you know assignments you know they're on all the GPU nodes but it's you know sometimes it can kind of be a fight versus you know the pen testers among your users that's awesome stack I know you had another one yeah um this kind of goes to the earlier comments Allan was making about just good design uh I worked for uh um you know I'm
not going to say I'm just going to say a very very large entity everybody here knows um and uh the group that was in charge of Designing the new data center um did not know what they were doing and they were taking lowest bid offered advice and uh the area that we were working in the um building is huge it is so big that there are actually buildings inside of buildings like it's just this monstrosity of a uh of a network of systems and build and um anyway um when they built this new data center I've got so many stories about it but this one
we go into tour the first time and as we're touring with uh all the engineers and everybody who been designing and building it of course I'm questioning all the design decisions that they didn't take my recommendation on and um I asked the the the fire chief Marshall um hey we were supposed to have um a special coolant for the fire suppression and I see water and he was like oh yeah they made that decision kind of last minute to save money okay as long as our uppers are aware of that and they're willing to sign a risk saying we're going to put millions of dollars
worth of equipment in here and if we have a fire or even a false alarm that sets those things off uh you're taking that risk so as long as you're wri willing to put that in writing I don't care but what I do care about is where does that water come from and who does it share with and he was like oh well that's off the main so it serves not only this building but also the ones all around it and I was like you're splitting it the the server room is going to be on its own water oh well that's GNA add this all
extra cost and I really pushed hard on this even with the push back of no the server room has to have its own water supply and uh eventually I got it well server room I think was in active for eight or nine months it wasn't that long and they were renovating another uh one of those buildings to make it into office space and the contractor swung a ladder clipped the water and all the other buildings flooded they lost all that office space just completely drenched but because they had set the uh cluster off on its own water it wasn't touched so the whole server room
was saved and I walk in and one of the guys who had been fighting so hard to keep it under a certain budget and had been you know a very vocal opponent about splitting the water just looked at me and nodded and that was it that was all I needed like okay we're good because that was just a very simple thing but if you're not watching for it you're not thinking about those things uh that really could have ruined all that new equipment and I'm sure the equipment is worth more than the office uh oh yeah desks and chairs and carpet that's that's a good
story um I can think of several like it but I think again if we try to go a metal level up we ought to maybe pass on tips to people on how to talk how to do what you did talk people into the right thing before disaster happens I have a quick example of a my own research group that had its own uh data storage uh uh in its in its lab in addition to the central storage and and it didn't have a battery backup unit and um I went and got a quote sent it you know sent it to them went and met the
department shair who was a member of my team and he said I don't know if we can afford it I said you got a piece of paper he said yeah so I wrote something on it and I turned it over on left on the corner of his desk and said uh you know next time a power outage happens turn this over and of course he couldn't resist right so he turns it over it says I told you so they bought the battery backup unit so we need we need to figure out how to get people to you know game the system in the right way
ahead of time before these things happen head off disasters well I think a lot of that has to do with uh the critical thinking skills instead of just looking at uh price and or um just trying to focus on well let's just get the job done but actually have that uh time to just step back and think about um what can things go wrong um How do I actively utilize this and some of that is experience um that same data center uh story um they were arguing for new tiles that had to require center cut tiles for all power cables you couldn't do side cut
tiles and when we had this big meeting about it and they no no no price is going to win out and I kept trying to emphasize that the side tiles were really important to make sure the systems keep running and um it finally took one of the the upper level managers to just be like all right stop and explain this to me as if I'm a child why is this Center piece uh cut tile better for an environment that's running and I just took a piece of paper and I took a pencil and I said this is your network connection on a system that can't
move but you have to get underneath the system how are you moving that and he was like uh you can't move it I'm like no and I was like all right now take a tie side and that same pencil and I can remove a tire uh the tile and he was just likeoh that makes so much more sense right like that's something that kind of comes from experience but also thinking about like when things are going wrong how do I still get the job done that I need to without disrupting everything else that almost looked like a a crayon drawing to me like here let
me do this in crayon so that you can understand the problem it's fantastic heral in the yeah and we've been talking about hardware and software uh or or RM minus F uh typos but um you know I think what you just said stack is is really uh U best Illustrated in the area of security you where the the Mantra is you know it's much less expensive to deal with a problem before it happens than after it happens so you know sec disasters could be another whole program I'm not sure we'd all want to confess to them though no probably not multiple users trying to thwart
the security protocols that the system administration team has put into place in order to keep the system secure the the the one that comes up for me is uh we always ran our systems with multiactor authentication you like onetime passwords little tokens and whatnot so you can get basically one SSH connection per per login um but users didn't didn't like having to type in their OTP passphrases on every login so they came up with all sorts of actually quite ingenious ways of circumventing this so they can just always just get a passwordless shell anytime that they want um so having to explain I'm not going
to go into too much specifics on this because I don't want to give any users out there any good ideas how to do this um but just don't um don't but I I do have a personal story that's a little embarrassing um a lot of people here know I've got you know a long amount of experience with werewolf which is a cluster provisioning toolkit um at some point you You' think that since I was developing this you know that I would kind of understand how it works um at some point I was setting up um uh a system at Berkeley lab and uh wherewolf well
let me the network architecture for Berkeley lab as well as many organizations have Network segments in many cases dedicated to entire buildings or entire you know floors with buildings and whatnot but these Network segments typically do not have a huge amount of um control over the broadcast domain like it's usually it's active switches and whatnot and and and what you know going to all all rooms all you know floors and whatnot so if you spin up a DHCP server on one of these networks um you wreak havoc on every single workstation and every system that is doing DHCP requests so I inadvertently brought down the
entire building of workstations not servers because servers had static IP addresses but in testing werewolf I or setting up werewolf I set it up to run on its public network card which again put a DHCP resource uh or service on our production workstation and server Network and literally every every workstation um that did a DHCP renew or DHCP request got a bogus IP address which was not really on the appropriate network with a bogus Gateway that didn't work and um yeah I took so I had I had a number of people show up at my door once they finally figured out who was responsible for
this and what was going on uh it was a little embarrassing it was almost as embarrassing as when I tried to teach a an intern uh a lesson about not leaving uh the workstation logged in and uh at some point after this going on over and over and over again uh when the this intern he stepped away from the computer I go up to the computer and I see he's not only logged in but he left a root terminal open on that system which means I can change his password and and do other things uh to that system without ever being authenticated uh so I
did I changed his password and whatnot then I logged him out and then I went back to my office and he came back to his desk and he's trying to you know get back into the system and i' let him struggle for a little while but after a bit I come over and um and I said well here let me try and of course you know he's like what are you going to do what are you and boom I just log right in and he said how did you do that I said well this is running sentos I put a back door on every sentos
system in the world you told him that yeah I did now of course to me he laughs and goes as soon as I leave he runs down to the cyber security group to go ratp me out and um next thing I know is again I've got a bunch of people at my door um interrogating me and uh yeah that was I I had to convince them that I was just joking and I really don't have but it's really hard to prove that you don't have a back door into a system like it's just hard start reviewing code yeah you just I mean how many hundreds
of millions of lines of code is in the entire operating system now granted you can limit it to a few kind of core Services um but yeah that that was a fun one as well you are Forever on a watch list Greg if if you weren't already you is now it's true that's true oh man that's funny do we have do we have time for more we how are we doing on time yeah we still got time time for a couple more Jonathan's list is still going I know he's got a list all right so I have a cluster Design Story uh so this was
the first cluster that I was involved in kind of from the groundup designing and uh you know nothing and it was omnip paath at the time uh but that's maybe the most unusual uh thing about it was we were pretty earlier adopter of omnip paath but you know from an architecture standpoint pretty standard fat tree infiniband or omnip paath whatever interconnect Network and then you have this problem of well we we aren't going to run our our administrative traffic our our backend traffic over this Omni path it's difficult to boot over that kind of thing so you need another Network and the compute nodes that
we had selected um were in this weird inconvenient time where they decided that the built-in interfaces shouldn't just be gigabit Ethernet anymore they need needed to be faster than that but we didn't have uh prevalent 10 gig ethernet over RJ45 connectors so the builtin network on those systems were SFP which if you if you're not familiar it's usually for Optics right and we're like we we aren't going to build a whole fiber Network for or we don't want to you know buy 10 gig switches for all of our backend admin traffic so uh our our our vendor I'll try not to to call them out
too much our vendor uh said okay well the the cheaper thing than than building out a 10 gig Network for your admin network is just to buy these little SFP to RJ45 transceivers and everything we're like okay that's great we'll just uh get put this little SFP module in there yeah yeah yeah yeah uh so the other cost-saving thing you run them one gab per second is what happens then you run them at one gigabit per second is what happens right yes yes it's meant to to step it down from 10 gig to 1 gig um and I also am aent even now sorry you
know why I'm sorry i' I've done exactly what you say yes the next thing the next thing that was the issue uh I I remain a strong proponent uh in compute nodes of running your baseboard management controller over that shared port with your management like your onboard management you don't I don't think there's a need to on a whole separate physical Network just for out of band on your compute noes there too many cables it's not worth it that's remains my perspective but so now we have not only the host management connection but the outof band uh BMC connection over this transceiver uh or you
know step down transceiver from SFP to RJ45 well we get the cluster deployed and the the vendor brings it all up and it's all working and everything is great and then uh we shut a node down um to do like a some maintenance or something and then realize we can't access The BMC over the network it doesn't work and we realize that any node that is shut down you can't talk to over the network again these bmc's you're meant to they're meant to stay on all the time so that you can turn the nodes on and off uh over the network well uh this functionality
is completely broken now and we we learn that the issue is that the BMC in these nodes uh because it assumes that that shared network conect ction is a 10 gig SFP has no capability of negotiating the port to 1 gbit so not only when it's been powered off can it not until the OS is booted negotiate that physical ethernet port to one gig to talk to our 1 gig network uh but the Intel driver for that network connection when it does just a complete normal software shutdown is issuing a reset that resets it to its default State and renegotiates Ates it to 10 gig
by default so even if you just reboot the node while it's rebooting uh you lose all out of band connectivity to it so we never did end up getting it completely fixed the uh the vendor was unwilling to change the BMC the the uh the baseboard controller such that if a node had been physically turned off that when the power was restored you could turn it back on we so every time we had like a power maintenance which we had a a free air cooling tower uh that had to be cleaned two or three times a year which meant we had to do a full
data center data center power outage a couple times a year and every time we would do that you'd have to go in and physically press every power button on every compute node in the cluster and of course these aren't meant to be pressed very often so they're Tiny Terrible like membrane buttons that then wear out after maybe 20 presses and now what do you do so that's a whole issue but we at least did get I think it was Intel a adjusted the driver such that when the kernel was shutting down it didn't reset the uh the interface so as long as the node had
ever been turned on it would stay negotiated to one gigabit and you could turn it off and on and The BMC would be fine uh but that that remain a problem for that cluster for its lifetime we did almost what you did but we actually had bought the 10 gig uh RJ45 switches if you do use cat 6 right 6 cables or 6A cables you can you can do it and we had all that but these transceivers wouldn't run at 10 gigs even with 10 gig rated cables and 10 gig switches and the reason is power uh the trans the adapter it consumes enough power
that the port would run too hot and so uh the thing got limited to one gig anyway uh I I took our vendors Network special IST on a tour of our cluster after we built it and I just opened the back of the racks and pointed and he looked at it he thought for while he says yep and then the next maintenance we switched everything out for SFP uh plus uh connectors uh that cluster has run six years just fine with those All Copper 10 gig our most recent cluster we have 25 gigs to the node with which are actually done out of 100 gig
uh ethernet switch with splitter cables so you know I'm I'm a big fan of fast uh control networks um and I don't know why anyone makes a one gigabit per second connection anymore I I these days and for the cluster after that I advocated for a faster ethernet and then used that for our storage uh separate from um yeah it opens up a lot of things monitoring we have uh we've talked about it an earlier calls we have a really fancy extensive uh Advanced uh uh scheduler integrated monitoring Network that pulls redf fish Telemetry and correlates it with user jobs and everything and you know
with the 25 gig control Network I don't even think about that um and like you say storage you can do lots of things um I think a lot of people and may you guys are experts in this right isn't a trend towards booting just enough to get the fast fabric up whatever it is infinite band on path and then booting over that uh transferring doing the pixie booting over that I've never tried that but I yeah there are definitely people who do it um in my experience the biggest issue is just what a moving Target those those Fabrics tend to be and maybe you you
work really hard and you get it working with the current generation and then when the next Generation comes around it doesn't work work and you have to do all that work over again the next update doesn't work and then you have to deal with it all over again yes those are has been my experience they've been very finicky you can get Port Parts in time where it works and then everything else is in it's it's a great idea it should be what we should be able to do but it's it's one of those uh reality meeting practice and I shake my Angry yeah it's another
category of um of disaster story right which is just outsmarting yourself right actually that's a point that I wanted to make about about the previous story of um of the transceivers is I from that and experiences like it um I've tried to develop a sense of when when I'm deviating from what's expected and trying to assess whether I really know what I'm doing and having a one gig Fabric or a one gig Network going all of the nodes going through an adapter to get to an SFP port in in retrospect like my my older self would think that that is a smell of something gone
wrong before I encounter the issue with it you know people say that you have to to know the rules before you can break them and I think this is one of those things where do do we really want to make this this decision that deviates from what the people who built the node expected us to be able to do and what kind of issues are we going to have as a result so we shouldn't miss the opportunity to point out and Greg I'm surprised you didn't jump in and do this that a lot of the design of werewolf 4 was to uh allow you to
separate those layers so you can uh you know reliably get uh the nodes up quickly and then do more clever things with more certainty um if I could only penetrate the mysteries of how Raspberry Pi handles its fr whereare I keep bringing this up so that Jonathan won't forget to answer his email it's a mystery to everyone no we I yeah I don't want to get too much into into the Weeds on that but I I I I'm uh I'm supportive I I found some information but we'll we'll talk after I feel like I have a reason uh yeah no it it works just I
can't get the PCI to run um with the Rocky and send off images okay I have a reason for bringing it up I think one way to head off disasters is to train um people on you know less than production systems everybody has a test cluster but you know that test cluster is often used by the smartest and your development team right and nobody dares touch it so I'm really on a campaign to try to get really cheap cluster Hardware into people's hands now I don't mean you know $1 thousand used d noes or something or you know HB one I don't mean PCS that
get recycled I mean little little tiny single board computers and I'm that close to being able to get uh the full werewolf uh image shooting and everything working on Raspberry pies with a little cluster board called the touring Pi 2 uh and uh it say it has a built-in switch it can support um uh four nodes right in it uh and if I can get just get control over that pcie so I can put a second Network on it I can make one of those the head node the whole thing costs a few hundred bucks I would like to get these into High School classrooms
into uh I let's see if I can get ciq to sell a a cluster Administration training kit you know batteries included ships with with Rocky and werewolf installed you can train your sis on that and then uh then you can turn them loose on your cluster it comes with werewolf on a little business card CD just to make Greg's money no it it just come these things are really cheap and fun and um yeah and anyway so that's why I I have a side project in trying to get uh training for assistments on a better basis both the academic HPC um field and the data
center field as a whole desperately need people with this kind of train and nobody's stepping up to do it so yeah there was a project from um Oklahoma University I believe years ago called the little Fe that was using um little Intel Nooks for for some of it and there have been a couple of different projects where they've used some of these little tiny systems to build clusters um and it's great for a very basic introduction where I've struggled with a lot of those before especially when I'm seriously trying to train interns or or uh junior staff is to have all the tools that I
would have and so um Lenovo is making these little tiny PC things um and they've been making them now for a couple years um there's a couple generations of them and if you look for the right ones they actually have onboard ipmi and they have the support for multiple networks and they're only like six to 800 bucks if you find them like online um and those are really great because you can set up and say a few hundred bucks each as opposed to yeah it's a few hundred bucks each node as opposed to a few bucks for the whole kit right so and that's the
hard part but like that's the part where I've really struggled before is I want it to be close enough that people can say all right you've got multiple networks here so you're actually dealing with how multiple networks function you're dealing with ipmi and all of its weird intricacies um but it gives you a little bit more realistic view of how all that works versus some of these boards and projects that have done it before it works really only on the top level software layer so if you're a new user or basic admin that's fine but it really doesn't give you the details of like how
to deal with the hardware problems itself but yeah those are the cheapest ones I've known so if somebody could come up with a better cheaper way of having a little test cluster that has ipmi multiple networks that would be amazing yep I'll buy one well t p 2 board does have does have a BMC it's very crude and it doesn't yet support redf fish which I'm trying to push into it cuz that's that's uh that's the replacement for ipmi Lally like power on power off that kind of thing yeah there's an API and that's what I'm working with uh it's the funky way you have
to shoot an image into the uh the raspberry pies whe whether it's Rocky or raspian that it gets in the way so actually stack I think you if you really want to do this and and you do don't mind spending a uh you know two 3,000 you can probably build such a thing out of the nukes and uh and actually if people want to look into this serve the home has a uh the guy Patrick Kennedy has a his own little private week spot for these little uh computers and he periodically runs reviews of the little Nook uh things so uh yeah I think that's
a perfectly viable approach and and you could get closer to the real Hardware okay so I think we're going off on tangent here but I like your style a lot so what this means is that we're going to have to have a whole another webinar for exactly what you guys are talking about because Solutions are at hand and you start putting stuff into the universe with guys like you and GS and you start getting ideas of of what is wanted and what is needed and things just start to pop so it it's very exciting to be on The Cutting Edge with you everybody I'm so
so grateful for you showing up thank you for your stories and your expertise and your laughter and your embarrassment and all the good things um thank you for listening everybody on our podcast flops and threads so cool we now have a podcast there and we're gonna have to go into like where that name came from because I am very curious but again that's for another time thank you for watching on YouTube go to our website we've got so much going on the HPC stack we got the fuzzball we're going to be at sc23 in I think a week and a half yeah it's getting it's
get close you go on stack be there what no company will be there I will not be I've got just a lot going on right now so all right well if you change your mind I I I can get you in I can get you a PA pass we'll see Allen there though Allan's going yes yes we're gonna miss you there stack yeah yeah we'll drop by and see my company they'll be there there'll be plenty of people there so we'll do all right thank thank you everybody thank you so much we will talk to you soon next time bye bye thanks everybody [Music] bye
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.