
Ready to revolutionize your approach to performance monitoring?
Join us for a deep dive into Performance Co-pilot (PCP) installation and management using Ascender Automation. In the world of HPC, optimizing performance or troubleshooting issues can be challenging, but not with automation by your side!
This webinar is perfect for those in the HPC space, as we'll explore how automation can streamline your processes.
Transcript
[Music] [Applause] [Music] yeah [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on empowering the next generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and aper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source good afternoon y you're still on mute there you are I was making noises get off the mute what's up Happy February 1 oh my goodness it is I know in my journal this morning
I was like wait what day is it like two one 24 this is awesome D awesome life is a wild wild adventure and it's good to be here um I'm pretty excited about topic dude me too it's always good to have Greg on it's Greg is the best so it's definitely bringing Greg on Greg there he is oh see he's on mute too see we got I'm Greg I'm a Greg maybe not the Greg but I am a Greg you're egreg this is accurate so we're gonna be talking about monitoring specifically HPC monitoring with one of our newest and most wonderful products called Ascender um
but first first we want to tell you guys about a little giveaway that we're doing because hello if you're watching us on YouTube thank you thank you thank you make sure that you like and that you subscribe and that you share this with your friends so that they can see it and they can like And subscribe not just because it's fun and I'm waving my hands around but obviously what we do here is educational and we want you to be part of the um in the know of what's happening in the HPC and Enterprise world and what ciq is bringing to both worlds um but
we also have a giveaway coming up so as we're getting closer and closer to a thousand subscribers on YouTube which is very exciting and thank you thank you thank you so much for being a part of that we're going to be giving away a backpack with Rocky Linux logo on it now I was not told we were going to do this but can we get a picture of that up like is that possible maybe at some point just like flash it in the middle of uh Greg's talk about a sender that's a good question fun and cool I mean I can see one on my
View full transcriptHide full transcript
screen right now there is find just open one here hold on it looks pretty cool I love love backpacks like I have had share your screen who it's got to be using I'm looking to just get the image out hold on maybe if I could make it keep talking I'll get it I'll get it out in a second hold on okay let me see if I can make it bigger yeah because I don't want to like share my wrong screen would be really embarrassing but I'll try I'll try this thrilling Jess wait don't move that because I'm gonna share that one right there can I
do share screen uh shares in the last two this is not very friendly this is not a zoom I'm not doing it I'm not doing I got it hold on all right you do it can I sing can we stop you I think it's the better you can take control man take control okay so anyway so I'll give you some details while we're figuring out how to show it to you so what there it is Rocky Linux backpack that is amazing okay so when we reach a thousand subscribers on YouTube so it's not the 1,000th subscriber so don't be going around trying to like subscribe
and unsubscribe trying to get like that perfect match but when we reach that we are going to um go across our uh social media platform so YouTube Twitter so make sure you follow us on Twitter as well um what's the other one uh LinkedIn LinkedIn Twitter Youtube and we're gonna pick one lucky winner to win this amazing backpack so again thank you for being part of our like first little giveaway which is super exciting and also being part of our community it's awesome okay have about that I can stop sharing my screen thank you so maybe you want to do an intro Greg Soul just
for anyone who is not familiar with the magic of the this man uh sure sure so I am uh Greg Soul uh Solutions engineer over here at ciq I've been in it for about 20 years as you could tell I have no more hair left right um so I compensate with the the mustache obviously sorry if my voice is a little rough today I have the uh the virus which you're not supposed to mention on YouTube because the algorithm will do weird things so pushing this is my second my second um webinar to do while uh under dur rest and we'll see we'll see how
it goes but um yeah 20 years in the industry a lot of it heavy networking whether it's like uh doing ISP stuff or doing fixed based Wireless shooting internet to people all over the place I've got to work with some amazing folks from all over from I help light up some in Congo I'm big in Nigeria believe it or not like Indonesia to right across the street tell doesn't have to always be far away but um uh like most people in it wear a lot of hats so done a lot of different things and uh the last I would say about four years I've been
doing automation So anable based stuff for sure and I uh I truly enjoy that thank you Greg I truly enjoy you being here as well and as you're talking about like anible uh and kind of working with that talk to us about a sender and just give us like a quick overview of Ci's uh rendition of that yeah for sure so Cinder and this is an exciting time uh to be uh using a cinder because we have some really awesome stuff coming your way very soon but we are built from the Upstream awx we take it we make it easy for you to consume so
we've written I say we it's the the Royal we right it's like all the smart people on the team have actually done the real work I'm just the one who um writes tutorials for it so I get the credit which I'm happy to take anyway I digress we've written a lot of installers that make it really easy to grab and use both in just kind of uh in a testing environment I want to test really quick if I want to throw it in production if I want to push it to a cloud provider we've got a lot of installers that make that super simple for
all of that and and honestly if you're familiar with automation you're familiar with anible right we create these little scripture going to run called playbooks and it really tells the automation what to do well a sender puts Enterprise controls around all that right so that it's not just everybody running willy-nilly firing from the hip Wild West I mean there's probably still some of that but uh it's gonna crack down a little bit on it and essentially it provides an environment which you can run all of that safely right most Engineers are firing off their automation from a server everybody connects to the same server they
launch it the problem is anybody connect to that server can run any of that automation right often referred to as unchecked automation I learned that one from Ford that little turn of phrase there and essentially what this means is with a sender I can say who gets to run my automation when where and what they operate against and I can even safely share it with other people that I don't necessarily trust in the environment not to say I don't trust them but you know they're I would say Network and security are kind of diametrically opposed forces you know we're just sort of the light side
and the dark side fighting with each other so I'm happy when uh security just shares their automations and I don't have to have their Draconian Hammer cracked down on me did I alien enough uh uh security people yet I'm trying trying to get people to say something in the comments I feel like that they've kind of come together though like the network people don't want anybody touching their Network either so it's uh you have security making it painful and the network guys not want to open ports or yeah cre vlans it's always fun of course of course because everybody's the smartest person and and they
all uh they they all do it the correct way as opposed to how everybody else wants them to do it obviously and they do subnet math in their head which is amazing to me well I actually practice so uh like the uh the octets I haven't memorized I can say them forward and backwards believe it or not it's a special skill I have it just makes me nervous anyway 24 does that help yeah yeah yeah all I heard there was I'm the smartest person no absolutely not you said it I didn't say it you said it said it it's recorded all yall heard it no
no there's uh something I learned a long time ago somebody very smart and I respect a lot uh told me you never want to be the smartest person in the room because then you have nothing to learn so I enjoy being the dumb guy I I truly appreciate it so always happy to learn from anybody around me if I can but are you guys ready to learn a little bit about monitoring in the HPC space yes me because I am not an HBC fella uh but uh I am dipping my toes in where I can and uh part of that was done via let me
share my screen because I told you I used to be a network guy and what's a network guy without his diagrams and slides so I am learning a little bit more about how people go about monitoring stuff inside their HPC environment networking people they inms Network Bing system near and dear to our hearts we couldn't live without it like uh I don't even know if they could go to the bathroom if they didn't have their nms up and running it's like they would just be completely lost but in HPC it's kind of a it's an interesting world I've learned and you know it seems like
code nowadays or at least the code I write tends to be a little bit sloppy right that it you don't have to squeeze every ounce out of it anymore but that are some people that still need to be able to do that and that is the HPC folks right like even the the smallest incremental um processing saves like actually equate to really big either time saves or dollar saves which usually equate to the same thing right Z yes usually so a lot of what they do is looking at how they tweak and tune something or say they're moving to a new environment they're running a
new workload and it doesn't seem to be performing quite right or all of a sudden it starts performing differently well how do they actually figure that out well there's a set of tools that I've learned about called performance co-pilot or PCP that um actually is one of the ways a lot of those folks do that um so I was uh in our slack was talking about how he was going to be experimenting with PCP uh over the weekend and they were saying well I guess Greg's gonna get randomly drug tested but no it's performance co-pilot it's it's a kind of a suite of tools that
allow you to grab a bunch of performance metrics when I'm talking about metrics I'm talking about like uh CPU metrics right that could be util utilization weight time you're talk about memory stuff so memory utilization page faults cach hits your disk stuff right read and write throughput latency iops networking environment that's pretty straightforward one environment could be temperature and power which is interesting people don't often think about correlations of uh what's going on with kind of power or temperature but I have seen in some environments degradation of service due to high temps and you just don't realize that's you know a contributing factor but then
also like virtualization so I hypervisor performance or Rel resource allocation right so performance co-pilot will collect all this stuff they have kind of their own set of tools that allow you to look at all of that which is great but what I think is even more important is they allow you to Archive that information right usually for a set amount of time because it can get pretty bloated right if you store all of that so it's one thing to have the ability to connect into a machine when there's an issue and to do some troubleshooting but what happens if you don't realize there was an
issue until two hours later well this allows you to kind of go back in time and figure out what was happening in the moment to really kind of troubleshoot so there are I found a lot of different architectures associated with PCP and how you can use it and this was kind of the simplest one that I found that I thought I would kind of automate so in PCP you have the concept of a collection host right here it is really just any machine Any server that is running PCP to collect metrics on itself right that's considered a collection host nothing too crazy there and most
often you'll have your admin or you know anybody that wants to check out that information you're going to SSH into that host and you're going to run the PCP tools to see what's going on right pretty straightforward it works a treat well you end up with your Fleet of servers here you can I have six if I'm trying to troubleshoot an issue probably not that big a deal when I get to I don't know a 100 or a th000 it's probably going to be pretty painful to connect into each one of these machines to pull metrics right it's it doesn't feel scalable to me now
it's not to say you shouldn't collect them on every individual host I'm a big proponent of actually having them there if you want to be able to go back and kind of check that stuff out but also they introduced the concept of a monitoring host let's see if I can adjust that to be a little a little bit more in the middle so a monitoring host is one that will either push or pull or rather the collection host will push to or the monitoring host will reach out and pull that collected information and'll store it in one place right and so they suggest best practice
your monitoring host only pulls a thousand hosts worth of stuff in there although I'm sure if you beefed up that monitoring host you'd probably be okay but essentially now your admin has one place to connect in to to grab all this information and PCP is more than just uh I'm running like PCP htop or whatever to kind of see what's going on you can also like run reddis and grafana and you can do kind of heat Maps or graphs for all this information you can also take this PCP collected info and send it off to a monitoring server something like uh I saw Prometheus have
some plugins directly for pulling this information in so you can start graphing it for long-term stuff right so that would be great for looking at um scaling in the future right you're trying to figure out out I've got a new cluster coming in this is what our workload looked like before how to you know how much compute resources do we need it's good for stuff like that as well as doing triggered events right I see something anomalous let me trigger on this and alert in some way right maybe a slack message maybe an email maybe a ticket in the ticket system something like that now
going back to the push or pull model in my reading um if you have your collection host push that data that technically puts additional load on them which could skew your met even though it's probably just a very small amount like it depends on what you're trying to tune right so really the pull method is kind of the the preferred method your monitoring host is going to uh take a little bit more of a performance hit but again it doesn't really matter right if it takes it a little bit longer to to pull the information from all those hosts it's probably going to be okay
you can always just beef it up and you'll be fine there so I went with the pull method right so keep in mind collection host anybody who's pulling metrics the monitoring host the one that's going to grab all that info and kind of store it in one place I got a question for you Greg and it's probably a question for the audience too when you were reading through that did you find that people were doing that on the head node of a cluster or were they spinning up a VM somewhere else outside of the cluster where were they putting that monitoring yeah I didn't see
that specifically but I um in most of the stuff I did see people were putting it on just kind of a dedicated machine off to the I think it depends let me roll back it depends I think if you're getting to that point where you're really taxing that monitoring host where it's it's working hard it's chugging and doing all that you probably don't want to have any other critical applications running on there because it might affect its performance right so I would say keep that in mind if as you get closer to that 1,000 Mark or you're really starting to push the resources on that
monitoring host you know move it off to something on its side or you know what just bypass that all together and just start with it over on something you know some VM off to the side that's not going to affect any of your core functionality I would say it's probably a good way of figuring out where to stick that monitoring host make sense absolutely now I'm just curious I want to go to dig into how people are deploying this looking for specific use cases and how people are utilizing this isn't always the easiest um I'm not sure I guess HPC guys don't think it's gals
and folks and fellas and fetes and everybody in between um I like to tell people on from Texas so guys to me means a group of people there is no gender associated with that but um uh I think uh generally these folks just think it's not interesting so they don't really publish like how they're utilizing all these things in their use cases so I would love and I've been asking for anybody who has feedback on how they actually use all these tools kind of specifically in their environment like one person telling me is great but if I get like five or 10 uh I get
a better kind of more overall idea of how people really use all this stuff so definitely give me some feedback I'm friendly enough not necessarily friendly but I'm friendly enough friendly enough for those of you that aren't familiar with Ascender this is it so Ascender while it is a nice usable clicky click gooey which you see here um there is more under the hood so one of my favorite Parts which I never thought I would say uh is the API that's baked into this thing right so Enterprise controls but now that we have an API you've got all these existing Tools in your environment that
you know trust to love guess what they get to take advantage of automation now right so if you're a service now shop you can actually have it call in with say service catalog items they'll call in and perform automations or you can access the cmdb the configuration management database so like the list of all your host in your environment you can pull all that into here and use it you got a monitoring system it can call the API to fire off automations if it's sees something happen odd that it doesn't expect so it really gives you the ability to start using this for more than
just I have to interface with a tool I've got existing stuff now it gets the interface with the tool as well automated remediation it's like magic it's the dream right or as I like to say don't wake Greg up at three in the morning exactly what I that stuff but I'm going to I'm going to consolidate this a little bit and so in here we have job templates which really is the place that ties all of your required components for automation into one we have inventories all the hosts that could potentially operate against projects that's we're pulling in my get repository that has my playbooks
remember my kind of recipe for how you're going to perform this automation credentials how am I logging into the random hosts and then the job template again pulls all that together in one place and so I'm going to look for my PCP stuff here I should have three entries I have one that installs my collection lectures right so all of my collectors it's going to run through it's going to fire off all the required pieces for said collectors and then I have one here for my monitors so it's going to run through and it's going to build out the monitor host and all that stuff
and I've created something called a workflow just to call them one after another so I could run them individually and that would work fine or I could use something like this called a workflow where I can take individual job templates in here AKA playbooks and I can string them together an interesting way so I could here have uh while it perform this and green means on success go this direction I could have actually added a plus and say well you know what if that first instance fails for some reason run this Playbook instead right so you can see I can start branching and I can
have multiple branches and I can have them converge back together so it really allows you to start treating your playbooks your automations like Lego bricks so I tend to shrink down what I'm trying to do into like kind of little discreete things so that I can start reusing them in workflows really easily this is my favorite way to share automations with people so let me close that exit without saving don't break it Greg because you're about to launch it and I will go ahead and click launch on that now like like I say like any good cooking show I've already performed this because this takes
a little second to run to connect all my hosts and so again this is my monitor host this is my collection host right here so whenever you have your collection host as soon as you get PCP installed uh if you want to test and make sure it's up and running I can type PCP it should show me some information about uh what's running on the system kind of Hardware wise that tells me PCB actually installed so I'm good in that respect now something else that it does is it's going to configure this host to actually listen on Port what is it like 443 3 21
I think exactly and part of my automation is I actually tell it which monitoring hosts are allowed to connect in so I basically set up an access list right to keep it a little bit more secure so just not everybody can Hammer those um those listening ports right so the only the authenticated systems are allowed to uh connect in so as far as this guy goes PCP is running it's listening does it like the Playbook is actually fairly simple I can move over to my monitoring host and I can take a look at what directory am I in VAR log PCP PM logger I actually
set this up in my playbook as well my monitoring one and I can show you that in a minute if you're interested um but essentially I'm saying where do I store all of the info I'm going to collect from all these different hosts where am I going to stick it I'm going to stick it in this folder so I'm do a ll I'm going to list of folders I can see Greg Rocky 9 that's my uh collector over here so I'm just going to list the contents of Greg rocky9 there you are I can see it's collecting collecting various statistical information in here I can
see that it's actually getting pretty hefty and I installed it this morning at about I think it was about 10 o'clock and it's already getting pretty big on the file so I have these set up to store and they're kind of recommended is about 24 hours of historical information in here right because um it will get a little bit Hefty and I mean you can always adjust that if you want to completely configurable but as you could see it really was as simple as me clicking launch and it reached out and configured all those hosts and my monitors now if I add one more host
I can still runch rather launch the exact same automation I don't have to specify hey run it just against this one host because anable excuse me anable is item potent which means it will run you see all these green okays that means I didn't actually need to make a change cuz I'd already run this on this host so it was already preconfigured so I can run this say for compliance to make sure everything's configured the way it is if I'm update some of the settings I can run it and obviously it'll make changes against everything if I add one more host to my inventory I'll
run it and only that one host will have the change lines in there so it makes it really easy to rerun your automations without breaking things so when I used to manually script stuff back in the day it was a bad time if that script broke halfway through because if I tried to rerun it probably wasn't going to turn out well for myself or anyone else involved or probably those around me I was probably uh yeah yeah so it's a pretty cool function of uh ansible that I can rerun this stuff over and over so that was that workflow I can show you that exact
same collector install from my previous run you can see all these changed entries right that's where it's actually going through and configuring I can rerun it over and over and over and it's just going to say okay afterwards so this stuff is not the most exciting thing I could potentially show you um but it actually is pretty useful I'll pop into say the install Playbook really the only part you have to configure as a user is to tell it here are the remote subnets you want to allow the monitoring host to connect into after that it's pretty much just going to go in install PCP
configure all the the files it's going to update your firewall in your machine it's going to update SE Linux start the service and that's it so it really takes all the work out of it for you if you want to do any customization you can come in here directly you can override those variables at runtime really makes it simple in about five or six lines you can configure that on a suite of 5,000 servers right like genuinely like hands off and it'll just go which is like kind of crazy to me that's kind of mind-blowing it's very cool yeah so for sure any HBC folks
using PCP I would love to hear your War Stories how you actually utilize it your various scenarios um this is kind of a generic configuration right I'm just taking most of the defaults in here like I'm curious what the vast majority of people do to like tweak and tune it and change all that stuff if you would be interested in me adding say to the monitoring host like the the reddis and grafana piece that's like two lines I can add that I'm just I didn't know if anybody was interested in it so I'd be curious if they are we have one comment that was talking
about uh having issues around 70 to 80 noes on a single host so they have that's where they kind of have their monitoring machine set up it's around 70 to 80 notes nice well all the documentation I read said about a thousand so that's good information and I think examples real world precisely and I was thinking it probably also depends on how much information you're pulling right because you can um extend PCP for various other applications so if you're running some of those other applications that have a whole lot of information that it's going to be collecting I mean it's it's gonna it's GNA skew
that uh collection size very vastly I would think real world I love it me too it's always fun for sure that was great Greg I I'm I too now I'm ready for feedback I want to hear how people are using this thing I want to see it in the real world for sure or maybe it's inspired somebody to go turn it on that they haven't wanted to go to a thousand noes and do it before yeah it's pretty it's pretty neat because you do have what I think is neat about it is it emulates the regular command line things you're going to do to do
your troubleshooting right like kind of um an intop right you could do PCP HTP and it'll actually show you it'll like pull it up like it's just the application and you're you're walking around and and doing stuff in there but you can also at the same time export all that and then graph it right so you kind of get I don't know the Best of Both Worlds you can still troubleshoot in the way you're accustomed to but also have some long-term statistical graphs to go along with it which I think is kind of unique I'm not used to seeing that which is pretty neat you
know it's not like generally you can't go back in time and ping hosts from OU that's having problems right so it's like you know it's neat to be able to really see that stuff I also like tying stuff into monitoring whenever it will go do something something's happening it will trigger automation to go do that first line of defense the things that you have to go do and it can do it instantly instead of waiting for you to get to it for 15 20 minutes and then when you get there you don't have to log in you just look at the data so PCP does
have a binary in there that allows you to set up triggers uh so you can watch the the information the metrics that it's collecting and you can trigger into events now imagine you're doing that on a th000 host you'd have to have that trigger configured for all a thousand but maybe you do that on the monitor host but you're still having to monitor set up the app to do the trigger for all the things so to me that seems like it can get hairy pretty fast obviously if you're using automation I mean you could automate the process so it wouldn't be too bad but I
think a monitoring server pulling all that would be the best tool for that like to me it feels like that would be the best scenario as long as you could get the seven uh level of granularity on what you're trying to match in there I would think that would be the better place at least that's my opinion from experience in the past so is there like an environment that is either too small or too big for a sender type of automation uh I don't think so um automation well I say I was going to say automation doesn't care how big your environment is um as
long as your automation can scale it doesn't care how big your environment is uh which is a good point Rose because a sender can scale I mean we've worked with people doing um millions of hosts inside of Ascender environments right like being able to really uh get big and manage all that stuff but uh even down to I don't know an environment with 50 devices it's still going to save you a ton of time like uh I like to tell people a lot of my inroads were or a lot of the customers I would talk to were networking customers and one of the lwh hanging
fruit for networking folks is updating the firmware on a device that's 15 to 20 minutes of human time to do that first you have to upload it you have to you know reboot the device and then you start sweating when it doesn't come back when you think it should and then eventually it comes up you're like oh my God I was just about to get in my car and drive up there to make sure it was okay that has happened more than once to me uh and it's happened in things in other states anyway I digress um but then after that you have to do
all your checks right I have to test this stuff and as a seasoned network engineer of course I know all the things to test I have this big list that I could test but I'm just going to do these three and then everything's okay uh until 8 o'clock hits and user start actually using stuff and then I found out I missed something um but the good part about automation is it's happy to do all of that um and it's happy to test as much as you want without question it will test all of it it doesn't care and it will do it instantaneously which to
me is a major advantage over doing any of that stuff yourself so not only can you uh say you have 50 networking elements times 20 minutes a piece that's a lot of time that's a lot of human time you're wasting right there to do all that stuff all right say a zero day comes out you need to update all that stuff now you need to make a change window you've got to scramble do all that stuff or I can just put it in my Automation and let it safely do all all of that so it's huge even for a relatively small environment it also sets
standards around how you're going to do everything you don't have uh we love saying snowflake there really aren't that many snowflakes out there but every engineer will do things slightly differently if you don't have kind of a template or a pattern for everybody to follow and if your automation's the one doing say network configuration it's doing system provisioning or package Management on various things it really is going to be able to consistently do that in the same way for everybody so it uh it's definitely a GameChanger now does it take a little while for you to um get accustomed to using automation absolutely right everything
does you know at one point you didn't know how to ride a bicycle and guess what it's you know it's like an awesome mode of transportation once you get the opportunity to really get in there and figure it out um for people that uh don't always have the time to dedicate to learning it uh that's one thing that we're happy to help with is that we can come in and help get your automation off the ground right I'd like to say we get the airplane Off The Runway into the air so we can help you build out your environment so that it will scale to
whatever size you need to right we can help you engineer it as big as you need it to we can also help you build some of the automations right off the bat right so you can have some templates to follow so if you don't have a corporate standard on this stuff looks we can give you some stuff that will help you develop that corporate standard but also some automations that save you a ton of time and we can teach you some best practices so we can do all those things happy to happy to do that that's actually um while I enjoy doing these things that's
my one of my favorite parts of the job is really getting there getting my hands dirty and helping people kind of learn and and grow and see those Li ball moments where they really start getting in and just you know that um that genuinely uh Sparks Joy is that is that still a phrase people say of course it is I I I how can you not love that phrase sparking Joy like who would not want joy sparked I don't know it sounds like something a pyromaniac would say okay so as I'm thinking about automation I mean is it possible that there are people out there
that have no automation tools whatsoever like have you come across companies where you're like what are you doing this is insane okay so that that's a piece of the question and then the next piece is if there's someone who's like yeah that makes sense I should do it but I don't even really know where to start can you start with like one thing that you automate and then build and build and build on there and or do you really need to have the big Vision right like do you need to know where you're going to start or can you just start those are really good
questions and believe it or not not I've run into those questions from multiple five Fortune 500 companies um and Zan will attest we've talked to a lot of Fortune 500 companies that are enormous and are in charge of some like life safety sort of stuff and they're not using automation it's like how how do you guys keep your hands around all this stuff it's it's like to me it's like terrifying um but also I came from some environments but you partially my personality but I came from environments with high slas you know and so uh just operate on a different level of paranoia but yes
there are some people that get pretty big and they're not using any automation so um there you don't have to be you're never too small and you're never too big to start using automation I mean I think it's it's going to Aid and assist um always something I've noticed is companies tend to grow um a lot of times through mergers and Acquisitions and often times we're not getting additional bodies to help solve these problems so we have have to use automation right we have to start using tools to help us kind of grow and and most of the people I talk to have really good
teams like they have really good it teams full of smart people and they just need extra help you know and they're going to be able to do that through kind of an automation tool and I I think to your point Rose um the uh the methodology I would always talk to people about is crawl walk run right you got to start somewhere so whenever you're very first getting started there's always some hanging fruit and that is to collect information about your systems that are out there nobody has all their infrastructure documented I have never once ever talk to a customer and that's not to say
people don't do a good job but nobody ever has it completely documented because oftentimes it's a somewhat manual process and if it's manual it's at the whims or time or availability of a human and you know so it's like what are you going to do so one of the things most people will start with is information collection right it's safe I'm not going to break anything I'm not going to hurt anything I'm just going to go and pull a bunch of information from my host well once you collect all that you can learn to start doing very discrete changes which means I'm going to change
one small thing right I I'll do it safely and I'll make sure I feel good about it and we've seen literally seen it with virtually every customer I've talked to right you you just start small the safety you know until you get comfortable you kind of ease your way into it you'll do a discreet change and then you'll widen the scope of that change you'll get bigger and bigger and then eventually you'll start defining what portions of your infrastructure look like in code so in my git repository whether it is a uh a server or a networking element I'll Define how it needs to be
configured in code I'll have the automation take that and then push it to that device to that hypervisor to that cloud to whatever it is so now if I need to make a modification I never actually touch that device I just do it from the code which which is like again it's another one of those learn to ride a bike cuz that's also kind of like a mentally scary thing you know and have this automation maintain all of your infrastructure but whenever you do a Dr Drill holy cow will you be a believer when you can stand up an entire set of infrastructure with just
a couple clicks because I've seen people where they are required for compliance reasons to test their Dr plans right Disaster Recovery how recovery is something catastrophic happens in my environment and I've seen people where it takes them two weeks to do that and I don't know too many businesses that can really sustain a two- week hit on their systems being down obviously that wouldn't be all of their infrastructure Parts would start coming back up but still that's crazy um and through automation you know you can turn it into a handful of hours and part of those hours are humans feeling good about what's going on
and humans testing some extra steps right so I think the automation piece can certain certainly uh be Game Changer in any environment no matter what you're trying to do or where I love watching this go through cab whenever people start getting to their change advisory boards somebody started automating something and then at the end of it cab is asking first thing did you automate this is this automated it starts eliminating problems and tickets it it's fun to watch everybody get jealous of everybody else who's automated and they start together it's just fun or I've talked to um some teams and what they do is every
month they'll get together and they'll look at all the tickets that came through you know whatever their platform of choice is for like their team and they'll look at what did we get the most tickets on all right let's automate the top one or two things this month and then next month they do the next and the next the next right so it's like they're slowly getting the most problematic things and making them just go away which is incredible like I haven't met very many it folks um that got into it so that they could push the same I mean it's not George Jetson where
you just show up to work and you push the exact same button for eight hours a day they want to be challenged in new ways they want to architect they want to try and engineer new solutions they want to research the new thing they don't want to just sit there and keep the lights on all the day you like run on the hamster wheel so the automation absolutely allows them to do that I'm very passionate about this stuff I don't know yeah I get super jazzed about it you explain it so well but I imagine is one of the issues of security team kind of
looking at this and being concerned about there being one platform that kind of touches everything Rose you're kicking in my PTSD yes we've uh we've talked to many um many security teams and I think at some point there always we say there always has to be a password zero right like at some point some system is going to have to have like a password stored on it right and so it's generally your automation system it's going to be able to connect into all of your infrastructure to maintain it and but even then we've seen people limit the blast radius so in Dev they'll have like
an Ascender install in Dev that only touches Dev stuff right and then over in test they have one for test and private or uh they have a clean side in a dirty side right so the dirty side will have its own automation the clean side will have its own and the clean side can reach to the dirty and pull but not Vice vers it's like there's so there's many many many ways to to stack this stuff up but essentially if you want an organization to run today and in the future you're going to have to utilize automation so at some point your security team is
going to have to come to terms with the fact that machines like software will be touching and managing the systems inside the infrastructure and it's it's what do they call it in a negotiation a real one everybody's unhappy so they'll be they'll be compromises obviously one way no win-win but um I can't think of a single instance where we talked to a security team that was um feeling hinky about it and we couldn't help them feel better or come to a consensus that this was a good idea yeah but absolutely they will they will have problems with it um there's mitigating things you can do
is one is to protect it inside your network behind firewalls and you you know you just can't get to it from the internet and then say you do two Factor authentication on it which a sender supports right so you can do uh tofa um but ultimately you'll schedule automations to run in there and you have to trust that the system can run on scheduled intervals and reach out without human intervention to trigger that event right you have to you have have to when you're teaching your kid to ride the bike at some point you got to let them go and they pedal on their own
you know you got to you got to trust and believe right but you also put a tracker under their skin right here as well is that is that what you do well also everything that happens inside of a cinder can be um logged and shipped off for logging and Jimmy's putting in some awesome work in The Ledger and he's putting in some stuff that it's going to blow folks mind like I talk about excited I cannot wait to start steing his credit for that when I get to show it to everyone it is going to be Sensational like it is legitimately people will be excited
about this they'll say finally so I I can't wait to really start showing showing that off but pulling information making it usable for folks like security as well but system wise I'm sorry Zan I interrupted you but no I was that's what's going to be exciting for the security team is they may not like the fact that people are touching things when you show them what information they can get without having to go ask and what policies they can enforce without having to guess did it get done did it not get done like they can just go do it so it's GNA be very exciting
yeah yeah and then get a clean report back like on a schedule that whatever they pick every minute every day every month whatever they want it will just show up to them and they don't have to ask anybody for it that that lowers the friction so much that cool excellent okay do we have any um questions from the world questions I think so world I haven't seen any on this side okay yeah I don't see any on this side as well well we do get questions that actually come into um the website so feel free guys to go to our website ciq doc anytime that
you type anything in there that message is going to go to me I will grab Greg or zaye or Jimmy or any one of the team that can really answer in detail the questions that you um have put out there uh and we are happy to help please put Greg to work okay put him to work off of these webinars if I can tag on to that I am always looking for people's suggestions on Hey Demo this Hey Demo that like I would love to see this or how do you do that with automation because I only know what I know and I'm not exactly
sure what you're doing in your environment so if you've got something really interesting or um you're just curious how something works even if it's basic like I want to make a demo on it so send us as much feedback as you can comment on YouTube or yeah I think it's info@ cq. whatever get a hold of us let us know yeah yeah okay um the Dickerson 87 feels like the rent theme song needs to be remade for security automation reports can you can you do it Zane can you can you it nope not even GNA try no all right but I like it I like
it a lot thought you were in the singing mood today Rose what happened I actually don't know it so I would have to Google it but we're live so like I don't know if I like pause I'll be back in a second get it in my brain well we did 10 minutes to pull up a a backpack earlier I mean what's what's another we backpack okay so reminder you guys on YouTube thank you so much for liking and subscribing when we get to a thousand subscribers on YouTube we are going to randomly pick a lucky winner to win a backpack that's got a rocky Linux
logo on it it is an awesome black ppack with the cool little you know Circle Mountain Rocky Linux logo we are very excited to give that away we are very um uh grateful for your uh you know community and for you guys watching and showing up and and Greg mentioned a lot of the vide so we we bring him on here live to do these kind of demos but he does his own uh where he is demoing different things to do with a sender so there's lots of videos inside of this channel of the ciq um YouTube channel so go check that out if a
sender is of interest to you and you can leave a comment on any one of those or reach out to us on our website if you have something specific I know there's um there's a couple of specifics that I'll pass on to you Greg as well people that have have written in okay so yeah definitely uh we also have a Twitter we have a podcast flops and threads we are on YouTube we are in LinkedIn we love you we want you we need you uh definitely reach out to us and thank you so much we'll see you next time same time thank you Greg thanks
bye everybody [Music] bye
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.