
Harnessing AI and Machine Learning in High Performance Computing
Take advantage of this opportunity to discover innovative ways of integrating AI and ML into High Performance Computing environments!
Our discussion will cover the challenges and opportunities of combining AI, ML, and HPC. We’ll address integration hurdles and scalability concerns while also identifying potential areas for advancement. Learn best practices for implementing AI and ML in HPC, including algorithm selection, data preparation, and collaborative strategies.
Speakers
-
Zane Hamilton, Vice President of Sales Engineering, CIQ: LinkedIn
-
Rose Stein, Sales Operations Administrator, CIQ: LinkedIn
-
Forrest Burt, HPC Systems Engineer, CIQ: LinkedIn
-
Allan Sill, Managing Director of HPCC, Texas Tech University: LinkedIn
-
Jonathon Anderson, Solutions Architect Manager, CIQ: LinkedIn
Transcript
[Music] thank you foreign foreign [Music] [Applause] [Music] [Applause] [Music] foreign foreign foreign [Music] thank you foreign [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on powering the next generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and aptainer support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source happy Thursday it is a happy Thursday it is a wild and new Thursday Mr Zane it is good to see you
Rose it is very good to see you as well thank you sir so I'm assuming you saw some press release today I know this and I imagine that we're definitely going to have some questions from the world specifically the Linux World on like what does this mean what is it what is actually happening um so yeah give us a couple of details just just overarching sure so high level well you can read the press release there'll be more coming out about this in the future but uh in cooperation with Oracle and Suzy Linux ciq is entered into a cooperation where we're going to actually have
the open Enterprise Linux Association so open Ela that's what's going to be called and that's going to be a place for Enterprise Linux to have a standard for everybody to pull source code from to actually go build your compatible Enterprise Linux distributions very exciting it's a great thing for the community and I love seeing this I love having these large companies involved with us and helping helping us do this very exciting so thank you to Oracle and susei thank you thank you ciq Oracle and Tuesday for coming together absolutely it's fantastic it's really exciting it really is so cool so if you guys have any
questions about that you can go to our website or you can just go to the openela.org and there is a lot of information there you can join the slack Channel you can join the movement you can ask questions all the things so very exciting very happy to have you all that is not what we're here to talk about today no no we're actually going in a little bit of a different direction sort of because that kind of like you know having open Enterprise Linux right available kind of supports a lot of what we're going to be talking about today specifically we're going in the direction
View full transcriptHide full transcript
of harnessing AI machine learning in high performance computing and so we're going to get into the details of that but I think we have some friends here are we are we adding some other people besides us yes Allen in the house Mr Jonathan Forest very excited thanks for being here hello I simply came to this one because I know how much Alan loves this topic and I just I actually thought about bringing popcorn I'm looking forward to this one welcome out hey um no it's nice to see all of you and I'm just here to keep you guys in line see you Allen I need
you can I can I are you for hire to come to my house too to keep me [Laughter] yeah that was a great topic gets a lot of uh discussion why don't you introduce the topic Zane I'm sorry Alan I miss what you said uh I don't think we know what we're talking about yet uh we are talking about Ai and machine learning and HPC okay great your favorites see yeah whatever you guys are talking about I love it let's do it I love it you want to introduce the others and we can launch it absolutely Forest I feel like it's been a minute how
you doing great husband can everyone hear me all right you can wonderful yes thank you Zane it's good to be back I've been uh out of the office for a little while but yeah excited to be back in this week and excited to once again be here at the Round Table thank you Forrest Jonathan yeah hi everyone my name is Jonathan I'm a Solutions architect here with ciq and uh know a bit about HPC uh a little bit less about AI so I'm also eager to hear more opinionated opinionated AI voices in this conversation and participate as best I can Allen okay I was hoping
we'd get more uh more we have had some fantastic people on this table over the last few years and I was hoping you would pop in um so this is of course uh buzzword compliant topic um and uh you know we're moments away from the the mashup of quantum and AI buzzwords so uh watch for that to appear on your screen shortly right uh you know Quantum AI startups are just around the corner right so um so what I think I'll go back to exactly what you said Zane what what do these terms mean what what do we mean by AI do we mean deep
learning do what I mean you know generalized uh sorting algorithms do we need generative learning do we need uh can we mean you know linear regression it's all linear regression by the way um do we mean predictive language models what what uh so I think the first thing we can do as a community it would be beneficial to everyone is the old high school debating tactic tactic you could always bring the other team to its knees by asking them to Define their terms right oh so what what are the what do these terms mean so I'll just I'll just throw it right back at you
guys what what do you mean by AI what what are you seeing it play out in the marketplace that is to be asked it that way it's a great question Alan thank you for that so uh it's been fun guys I really appreciate you coming uh I don't really have anything to debate anymore no I'm just kidding I I'm gonna let Forrest answer this because I feel like whenever we start talking AI I typically default to Forest he spends a lot of time on this topic uh he enjoys this topic so my opinion Alan is it's all of those things I think it depends on
what part of that thing and what part of HPC are you looking at doing at that moment in time so I feel like this could just be a never-ending conversation but I'll let Forrest how about it overall I would say AI um is a pretty broad definition in general like you said Alan are we including these you know kind of sophisticated um you know framework based AI models that are built on something like Pi torch or something like that are we looking at more you know the use of mathematical models you know like I said it's all regression in the end um I tend to
define AI as being the genuine creation of one of these you know model based architectures in some type of framework so whether you're building you know one of these perceptrons one of the more modern types of uh you know large language model GPT type things um I kind of overall just refer to AI as that catch-all of um these specifically built you know usually layered to some extent uh neural network type architectures that uh are a little bit more than just taking a bunch of data and applying you know a few different statistical models to it and kind of seeing you know what you can
predict with that um overall I would say AI comes down to having that um you know neural net like architecture where you actually have uh some type of you know you can go really deep with this but you know you've got activation functions you've got neurons you've got um a thousand different ways that people do that type of thing um so overall I kind of you know look at AI as that um layer based creation of some type of potentially autonomous agent trained on a specialized Corpus of data um for some type of purpose such that it can make informed decisions uh on its own
um so that's kind of what I would say is machine learning is a little bit um you know included in there is uh more so I would say on the statistical side of things and it is the creation of artificial intelligence but I know some people kind of lump those a little bit more together um I've kind of found in general that's sometimes a gray area but for me AI is neural Nets the GPT type things these sophisticated agents that are meant to take in some amount of data learn something give you a more sophisticated result than just you know basic statistical modeling Jonathan yeah
so I'm with Forest about ai ai is pretty broad um and I I think of these as viewed from different perspectives so to me AI is most useful as a a catch-all term for a user experience than any particular technology or or field it's it's you know it computers that are trying to interact with us with with the users of the computer like another person would let's say so you're you're talking with a an AI chat bot and your experience is more similar to uh speaking with another person than you might otherwise expect from from a machine things like that I'm a systems guy so
where I Look to these terms is to help me differentiate different parts of systems architecture and so my definitions are less accurate than they are useful to me so when I think of HPC like if you just pull that term apart and you're like any high performance Computing system that's not useful to me but when I think of HPC as that culture of computing that grew up in the kind of Nash National Lab space in the University space that kind of distilled itself down on something that can do a Beowulf model with a high performance interconnect an MPI things like that that's what high performance
Computing is to me is distributed memory fast interconnect running a process across multiple system images and so for me when I think of machine learning I'm less concerned with the actual machine learning the application and more they're the guys who actually use the gpus I installed uh and so that that's most of where it begins and ends for me they're also the guys that tend to fill up my storage but then don't use my interconnect and so that that's that's where it breaks down for me is is less like that's not a useful definition of machine learning to anyone except for a sysadmin who draws
the line at well the application does something but what I know is it makes the gpus get hot and it fills up my storage great okay Alan well so um you know I I think uh there's so many directions um thank you those are the answers I I wanted and I think they're accurate um the problem I think we have to face is that uh we have to be aware of the fact that we're the folks with a hammer that in the proverbial when you have a hammer everything looks like a nail right some people come along and say AI we say HBC right so
uh oh you want HPC so uh so uh you know we have to be sort of careful what what they want um what they seem to want is machines that think and so so this leads me to um what Zane was expecting me to say you know my uh off repeated but I think still accurate statement that there is no AI there is only uh obscure scripting with unexplored failure outcomes the problem with that that statement is it doesn't make a really good marketing t-shirt right so um you can put it on would probably buy it but uh but we've made we've made freeloader t-shirts
surely we can put yeah oh yeah I'd be rocking that on the daily okay so uh so you know if you talk to this from the sales you talk about different perspectives so from a salesperson's perspective of Hardware software you know operating systems whatever it's um you know what would you like it to be right so um you know there's the uh the famous Dilbert uh comic strip I I try not to quote Dilbert much these days for various reasons but uh but I can't ever forget the one or the pointy here boss meets with the with billboard and his his colleague and says while
he says uh you know we've just gotten our consultant report and he's identified our biggest problem and while he says you know I recommend we build a tracking database and Dilbert says we can put it on the network and the boss says would you like to hear what the problem is first it's awesome you know we're sort of in that situation what would you like it to be well you know uh and the reason is it is a set of Technologies not a single technology sorry about that um so um I'm just gonna let that ring and let the rest of you talk I just
point point that that aspect out and then uh you know um the beginning oh Fernandez here good so you know then we can talk a little more technically about what are the ingredients of AI and why we end up applying HPC to them that would probably be a useful Direction thank you welcome for net it's good to see you hey hi everybody I'm just listening I don't know what's happened so far I was double booked no it's fine so we're kind of discussing what is everybody's view of AI ml when it comes to HPC so what does that mean to you uh I mean I
think it's all HBC now I don't think you can extract or differentiate uh HPC from what's happening in even in the broader uh workloads um so they're all scaling now they're all using gpus they're all doing really large data sets so I don't think there's a difference between Ai and HPC because I think it's all becoming HPC it's very interesting so what one of the topics I know Forest you and I talked about this quite a lot at SC 22 and we spent some time with a group who was doing a lot of data collection that we talked about wouldn't it be cool if and
I know Alan you don't like this topic either that's what I'm gonna talk about it having AI actually watch HPC infrastructure and stop problems before they happen or help prevent things or make changes I do I do like that topic okay so in the past I thought we'd talk about not liking that guy I think that was me that oh is it Jonathan yeah okay my apologies Alan no I think this is a great topic and my uh my take on it in a sentence is I don't think we have enough information in uh uh and we've been shy about getting it and so a
lot of my work in our industry University Cooperative Research Center has been uh drastically expanding the amount of monitoring um information we get in a useful way we had a super Computing talk I'll put a link up uh but yeah uh two or three orders of magnitude more information from the baseboard management controller so you can do things like watch memory power and CPU power separately on uh tiny fractions of a second time skills without ever instrumenting your code you know getting so what we work what I like to say is we work on precursors to AI you know you can't actually do you know
anything really intelligent if you're only monitoring something once per second right you know um and uh then the question is what do you instrument and how do you get the data um in you know in modern uh clusters yeah so I'll find a link and put it up thank you Forrest I don't know if you remember when we went and spent some a little bit of time having this discussion and the amount of data they were collecting I feel like it was like 66 000 pieces of data they were getting every second something like that it was some per minute gigabyte right it was a
lot shocked to see or to hear about yes um yeah I agree with that one completely uh just like I say precursors to AI the biggest you know we have the model and then you have the data set and one isn't particularly useful without the other um and so without you know some massive Corpus of data to actually look at that like Allen says gives us you know high granularity of what actually happens on a system that some type of agent can look at try to determine Trends from you know learn how to optimize a system based on um this is kind of what stalled
AI in general as I understand for a while um you know imagenet and things like that were very big because they were some of the first times where people had collated all in one place large you know for example tagged data sets of images and things like that that people could just go and use for AI um so when those those kind of things started to come around we kind of saw this explosion of you know image recognition and things like that and so yeah I'm I completely agree it's there's a lot of uh there's a lot of metrics and data that are left to
be able to effectively collect on a system um that we would have extremely large amounts of I mean the uh the basic um Corpus of data that the gpts are trained on they call it the pile which is this big huge bunch of like random web uh like stack Overflow Wikipedia stuff like that I think it's like 770 gigabytes or something like that so you can imagine you know in flat text you know if we can find interesting ways to compress down you know the specific types of data that we're pulling off of an HPC system um you can imagine you know the sheer amount
of data that we would end up with there at certain levels of granularity so um yeah it I have a huge problem at the moment is there is no massive kind of organized um Corpus of metrics and data and things like that that something could actually look at and be trained over yeah in addition to the link that I supplied there's the toolkit called uh like I know what I'm doing l-i-k-w-i-d liquid and uh that group has put together an astonishing variety of tools and uh and uh you know features we're working with one of the vendors on uh instrumenting um MPI code in a
more natural way again without interfering with the code you don't have to compile anything and you just use Linux perf and MPI tracking and uh so we can do things like you know many of the things you'd have to normally uh dive into total even do so let me wind back though my criticism about AI is that the reason I say it doesn't exist is that what we have uh in in general are things that uh can be trained uh everyone would feel free to disagree with me but they've been trained on pre-existing data or models and used to predict the next most logical outcome
in the case of uh you know languages uh you you predict the next most likely phrase or sentence uh uh whether or not it actually makes any sense so there's no thinking behind it that's my criticism same thing with art you can train it and it can make art but it's basically just a plagiarism as a service right so um you know so I don't think we are anywhere near actually getting machines that think but going back to what you said Zayn I think that problem is highly approachable machine learning based steering of computational codes has been a a positive influence on those codes in
a long time think about minimization algorithms you know you can write something to try to approach Minima or you can throw machine learning at it and this letter approach when you have those constrained problems it's eminently sensible people write code with with AI these days and you know we even have a whole major repository framework that does that steals your code and gives it excellent but uh I want to come back to that Alex I think it's an interesting topic and kind of diving a little bit to what you're doing and having to not instrument code but I want to get fernanda's opinion on this
too are you comfortable letting AI make decisions on your HPC cluster um not if it's a production system um if it's a research system where um you know a few jobs get lost and they have a 24-hour you know cap on running sure if it's um something that I'm running some you know week-long computational chemistry code or running something that makes planes fall out of the sky then you know oh I agree I don't want that either so I think we do have a couple questions I'm gonna go back to Allen's if I can touch on something else out there really quickly I just wanted
to mention that uh especially like you know coming up on a year post chat GPT release it has become incredibly apparent that these systems are not the um you know dynamic answering engines that we perhaps anticipated them uh to be at the start of it but there is a lot more of just kind of you know the prediction of the next most likely uh word in the chain essentially people have seen kind of the quality of some of these major AI systems like it can't pick out prime numbers and stuff like that anymore versus versions of it that were out months and months ago um
and they are completely unable to solve the problem of hallucination in the um even with all the resources and time and stuff that these places are putting into it they are unable to stop them from just making stuff up at random so it turns out it's unsolvable yeah I think it's completely unsolvable because in the end just like you said Alan it's not this is this is not machines thinking it's just some agent predicting what the next most likely thing in the chain is the big misconception is that these things don't know anything they just know what a lot of people have said about something
in a certain case um and so yeah I'm I just want to say I completely agree there there's there's not really genuine thinking it's you know sophisticated Markov chainers and stuff like that essentially in the end yeah you could you could point the you could point that criticism in both directions on the human boundary right you know you have to keep in mind that the stuff the ring training on comes from us and we're not exactly a source of Truth ourselves right so [Laughter] I don't know how you think but I only have act to what I have but right seeing her weird source of
chaos so we do it already I mean we're writing a script I said it's obscure scripting without explored failure outcomes right so we've been doing you know hopefully a little less obscure scripting uh with hopefully less obscure failure outcomes for a long time this is not new right we're applying some predictive and and uh neural net based you know machine learning techniques to it so you know I give an example of hierarchical storage management we've been you know for years these the problem has been in in place and solutions have been fielded to uh you know make sure you fetch the right data for the
next part of the calculation so you know it's not that different but the question is uh you know uh so so let's go back to the original thesis of the webinar right uh is AI useful in managing HPC I think absolutely yes um but you know are is that is it AI in the sense of machines thinking for you arguably no it's an interesting take on that island thank you Fernandez when you wanted to add so you come off mute before no I was saying that we're a source of chaos we're not really a source of Truth um and there's not a lot of people
spending a lot of time rebalancing data um we're still in the especially now in large language models we're still in the realm of dump everything and the kitchen sink into something and hope it comes out and um until we start to refine um our data and rebalance our data we're not going to get better answers we're just going to continue to see the kinds of answers that you would expect if you asked you know a thousand people on the street absolutely we want to throw the question back up we have a few questions to get to you so I I think this question is in
response to my my statement earlier about HPC being about distributed memory and just in case that is it to be clear um what that means in my experience uh isn't any particular technology about Distributing the memory it's not a large unified but distributed memory space it's it's a partitioned memory space that you then let uh um let processes speak into each other's memory uh with RDMA remote direct memory access usually over some kind of fabric using MPI or in other cases some kind of P gas or gas net system so that is what I think the answer to this is but I'm not familiar with
using Plan 9 to create a rem FS that you would then share with a control node yeah but points for recognizing the abbreviation true I don't I didn't know that it was uh I don't know that I'd seen it written the other way around but uh it was at least enough to make me confirm and okay yes I haven't run plan time myself though you certainly know where the name comes from oh yes oh good now we all have to know Alan because I don't really it's plan nine oh it's a famous old uh black and white uh um science fiction movie The Plan nine
from outer space I'm gonna go back and re-watch it I don't I hadn't even thought of that movie in so long it's not even if I can watch it what if you can stream that what are the odd friend who lost years to play at nine so we have another question okay I like it Sylvie thanks for watching and uh writing in a question so maybe Fernando will start with you how is chat GPT used and involved in HPC in your experience so I don't have a HBC specific example but it's a possibility for HPC we just wrote a Blog over Voltron data on the
startup that I'm at now where one of my engineers in my group augmented data uh the data set that we use is very common open data set which is um the IMDb open data set so listen movies Stars right stuff like that and what they did is they added a column and then used Chan GPT to give a synopsis of each movie and then in that column you know basically each movie now got augmented or abused of the entire data set got augmented with that and I was thinking about how we could do that with any other language model because you know if we take
the perspective of AI is compression and we can compress uh other kinds of data we could get richer data sets that have um you know uh extra features added on to these data sets that we can use then to compute on is it going to be you know great is it going to be 100 correct I'm sure that if we went deep down and looked at the synopsis of each movie it wasn't going to be perfect but it's it's a preview of what we can do with augmenting data set especially data sets that have been historically you know lacking I can think of examples and
projects that I've been in in healthcare where you have missing data uh you know maybe those those data can have we can have a better model by completing some data with like demographic information or address or city some sort of expectation of where that person might live based on other things that are within those data so maybe we might have state missing but we have the city stuff like that and I can see us getting a better more complete data than leaving it with nulls and and then removing them completely because I've got nulls right it's very cool thank you I haven't thought about it
that way foreign go ahead oh so it's easier to answer the other the question in the other order right how is HBC used in in these kinds of algorithms uh going back to the um the sales technician's response you know how would you like it to be so what is what is hpc's role in chat GPT right now Forest it would it would be everything yeah I mean everything the I mean it's um you know chat gbt exists probably in some way shape or form as between one to maybe half a dozen single files sitting on some system out there that something is running inference
on and that file was almost assuredly created that model uh based on you know heavy computation on gpus you can look at like um their partnership with Azure I think uh to bring more cloud computing into that whole project another example that's not GPT related but just is kind of emblematic is that uh Tesla runs a 5000 node HPC system called Dojo with a whole bunch of custom silicon that they have in it called the dojo chip that just is sits there and dedicatedly trains AI so at the moment you know not only are we seeing the massive repurposing of things like gpus when we
have you know 8 16 cores on the system it's much more efficient to not be using you know multi-threaded pipeline highly sophisticated CPU cores to do tiny little you know basically multiply some operations over and over again it's much more efficient to do the thousands millions of those that we need to do on the thousands of cores in a GPU um and so the GPU you know like you said Alan has become this kind of generic Computing device for a lot of different workloads um not only just AI for all kinds of different workloads but something interesting about AI is we're seeing a lot of
development uh like novel Engineering in HPC around creating devices beyond the GPU that are ultimately a gpu-like architecture you know a whole bunch of little tiny cores sitting on a chip but that are essentially dedicatedly made to just train AI um and so overall yeah uh um we see like I said the GPU has become you know I mean the de facto standard in uh the AI world for what you do your AI training on and those are mostly involved within the purview of HPC systems and like I said not only that but the entire industry of HPC is in like a technological Renaissance at
the moment trying to build um you know the cerebrus The Havana all these different um dedicated little AI chips um so it's yeah just like it's huge it is the industries are completely intertwined like you said Fernando you can't really differentiate the two of them at this point they're everything AI that you see is powered by an HPC system ultimately in one way shape or form right we should probably make the distinction uh sometimes you see the criticism and the media about uh oh every time you type a check GPT query you're using up so much carbon or something and that that part doesn't use
processing it's the part you're talking about Forest which is the training and the inference and the developing of the models for a lot of computer science goes on people you know run on our gpus to develop these algorithms and then when they get a good one they think you train it so when you actually just type a query you're not burning a lot of carbon you're you're just basically doing a web search among uh well a database search in the results so uh you shouldn't feel too guilty about asking it to draw a new picture for you or something thank you Alan Jonathan before I
ask my next question yeah I I don't have much Insider experience in how chat GPT is used in HPC today I mean we've we've done mundane things like using it as a glorified search engine for config file syntax and things like that or or helping us to develop containers in containerization languages that or you know definition languages that we have maybe less experience in but that that's pretty well understood I did have one idea though as the conversation was happening so you know I I acknowledged I've been um the critical is probably too strong I've been skeptical of attempts to use General machine learning for
fault tolerance or sorry like fault prediction and failure prediction and avoidance in HPC systems before not because I think it's a bad idea but I just I've seen multiple projects targeted at that and none of them come to anything and I do think like Alan said it's it's largely a data issue but in my experience it's also just been a putting the right people on an issue I've experienced it as a side project of a systems team or a purely academic experience that's not collaborative with a systems team so I think there's there would be value in continuing to pursue that either at a commercial
level as a product or um as as a joint research effort that has a bit more interaction than I've seen before um but I I was thinking you know we mostly experience uh GPT and it's ilk um as as a conversation National system or at most like it generating something that you would generate and now you don't have to uh but I would be interested how how effective it would be and I don't know I'm not an AI researcher but I would be interested how effective it would be to train it on machine log data or maybe even machine Telemetry data and then have it
predict future log Telemetry or or log data and then use that prediction to highlight divergences from the expectation and I I think most of these systems most AI systems are best thought of as a tool to enhance you know an existing operation something operators will do and there's so much data you know if you you get all this Telemetry data out of a BMC like Alan was saying and now what do you do with it you need something that's going to help call out the important information out of that and uh that that's what I would be really interested to see not necessarily for it
to take automated action though maybe eventually that would mean there are some areas where that's already happening um but at the very least to start I mean maybe is an evolved grep the it's showing you these are the things that diverged from what we expected or that you know more than some error bars expectation on either side of it these are log messages that are Divergent from the norm and someone needs to take a look at them that could be really interesting the general statement I want to make about that is if you have a closed form system I don't know exactly how to phrase
it mathematically or or with that kind of language but if you have something that's a closed form like uh Forest mentions uh writing set in Ox scripts you know so uh you know anything like that where there's a complicated step to be done but it's completely um within a set of existing examples or in your example you know it departs from the existing examples those I think are solvable problems and that's where I think uh the short-term win is um you know Fernando has been typing in our chat and she made the interesting point that uh that AI is basically compression you want to expand
on that oh so to speak I I think that's a light bulb but about three four years ago when somebody mentioned that AI is compression and then there was actually a chat on our company's uh you know slack where somebody had presented a paper um just this uh this year not you know just a few months back where they had used gzip as essentially their model um and they were able to you know query it and do things with it as as if it were an AI uh model that they built so interesting that um that we keep talking about AI as some sort of
Revolution but I think in the end we're going to find it's a lot simpler than that it's just trying to get things closer together and um you know just go on in that next little node in a graph and that's just a really giant graph that we don't know what it looks like inside but it's able to just kind of make Leaves and Browns in those little notes and in a giant graph and then tell us an answer and whether or not it takes a wrong fast sometimes it just um it depends on the kind of prompt that you uh that you put or even
how you word your prompt will take you a different path altogether so Fernando are you saying that this is going to be like virtualization was and then containerization was where we were doing it on mainframes years before we had any kind of version but it's the new thing it's the new thing exactly uh but I think I don't know if it made AI to me feel more mundane uh it's just a fancier way to do compression that's not a deterministic algorithm thank you do we have another question that came in from art oh the economics change well they already have it's the tickets to try
to get a GPU these days have you seen the prices of GPS yeah and it's let's be clear it's all it's all driven by the fact that the um the major Cloud providers uh made massive investments in huge amounts of Hardware because of anticipation of this AI need uh so it's already changed the economics by making the parts scarcer uh in terms of on-prem versus Cloud uh listen I always say that when someone invents another form of transportation you don't stop using the previous formulas of transpiration you change the balance of how you use them maybe but you know flying cars may come along and
and you will probably still have bicycles and trains and stuff so how does that apply to Ai and ml it's another application of HPC it's what Fernando's been saying we've been saying as a group uh you know you need HPC to solve these problems apart from a part shortage I don't actually see any you know first order change to the that balance of how you use this stuff lots of researchers you know here and at other universities use the local systems the clouds are still as expensive as they were uh that's why they bought all that Hardware to sell you time on um so I
don't see it changing the balance myself I want to reframe this question though because I want to challenge Arthur's comment that the Innovation is being driven by an AI and ml HPC is what drove Ai and ml it was putting down Titan and having a project with Nvidia to put those same math libraries same linear algebra libraries that were running old-timey AIML Frameworks that were sort of okay and then we put them on gpus and then somebody said you know I could use the same blies lay back whatever math Library it is that they originally put in there and they can substitute it for gpus
and that was what exploded AI so the Innovation came from HPC and we've known this for years we're sort of this you know um we're in this field where people think we're we're old-timey field oh you're still doing Fortran but we continue to drive Innovation and we continue to be left behind as if the Innovation is not being driven by us but you know passing is why AI isn't doing uh anything to change what we're doing we're doing a lot to change Ai and I wish more people would appreciate the kinds of um deep you know algorithm work that we're putting into scaling all of
that stuff I mean even like horovod right horovod was the first scalable framework for AI for training right that uh was it um uh the the driving company built it was built on NPI we built these tools we're the ones that are driving this kind of innovation a lot of stuff that's running today even in a large hyperscaler still using TCP they're not using you know the sorts of uh communication Frameworks that we've built in HPC we are the ones driving those Innovations they're you know Nvidia is hiring people from our field lots of the hyperscalers are hiring HPC people in fact we're struggling in
lots of places and labs and whatnot for HBC people we're the ones driving this Innovation AI is a consequence of HBC it's not driving HPC anyway that's mine so I don't even know how you follow that up I mean that's like a great way to end I yeah exactly Mike I'm sorry oh that was awesome thank you Fernando it's a great perspective I really don't even know how to follow that up because that was perfect I'm sorry no well I'll follow it up with a link because I absolutely agree with everything Fernanda said um so if you're curious how to learn these I love the
old-timey AIML myth I just did a picture of 1890s a AIML here um but there's a reference that I put in the chat that maybe someone can put up in the notes uh from Andrew Rings uh Stanford computer science 229 course last spring it's uh it's a PDF of these lecture notes uh and it just is a absolutely great uh appreciation you know covering all these basic uh techniques including the ones we started this conversation with uh you know what do you mean by neural Nets what do you mean by Machine learning and what do you know what are the various forms of predictive algorithms
um so uh I wish we had I wish I could come up with an inspiring oh but AI will someday be a thinking conscious uh counterpart all I know is if that day ever comes then uh you know all those prove you're not a robot captures we're gonna have some explaining to do yes no kidding Freddie you mentioned a Blog earlier uh from a colleague recently did you want to post a link to that uh did I mention it no I mentioned that paper nope I thought early on there was another blog that someone that you work with had posted um oh right oh yeah
I can do that yeah be great thank you well we are getting close on time but I think Fernando wraps it up perfectly so I will give everybody the option to say something if you want to but I don't know that you're going to top that Jonathan's not on mute sure uh so the only thing that I'm this kind of still resonating with me you know uh Alan has has challenged here uh repeatedly the the kind of notion or what we expect of of AI systems that we build and and how we conceptualize them and what they're capable of and this is a little bit
off topic here but I I like to use that as a reflection of what we're actually doing when we think and we conceptualize and and the more capable we make our machines the more we understand about what it is that we're doing that's different and I I look forward to continuing to be challenged uh in in how we think and conceptualize the World As We compare it to the machines we build I think it's a really interesting opportunity to learn more about ourselves thank you Jonathan anyone else before Rose wraps us up here no thanks for letting me start with uh with my criticisms because
I think it allowed us to go in in the direction of what is it actually good for and we've only just scratched the surface there's a tremendous number of application areas and you know a sturgeon's revelation right 95 of everything is crap so yeah 95 of AI is crap but you know that's par for the course right so uh we're looking for the good stuff and there are lots of very clever people doing really great work I don't want to diss them at all absolutely good point you know um for us you have we've had conversations about this before and I I brought it up
to him I was like so like are do we need to be afraid right because there is like AI is going to take over this is the one thing you should be afraid of and there's like a lot of that kind of chat and and Fernando you just said it so well and it was one of the things that that Forest really studies that well you have to remember that it's like it's people that are creating these these algorithms right and so it's it you got to take it back and and realize that you know we are we are creating this the way that it
is so I don't know maybe if you wanted to like rephrase that in your own words for us I think it would probably you said it better people are generating the data these uh AI models and and types of um methods that we have and you know if we look at them show notes uh uh Andrew Yang's uh um you know lecture notes these are all statistical methods that have been around for a long time there's nothing really new about AI other than speed speed is what made it possible the increase in memory is what made it possible the capacity of of systems and nodes
is what made it possible even the interconnect can be slow right now and is still making it possible that you know the slight speed of data collection ability digitalization of data that's what made AI possible there's nothing really new here it is a convergence of many things that it makes it more possible today so I think we need to be a little bit less scary scared of AI a little bit more skeptical that it can give us any new insights or new innovation it might speed up the innovations that would have happened anyway I agree completely it's mostly an imitation engine in the end and
just like you know Fernando a lot of these techniques have been known since like the 60s 70s 80s they were researching these things and it's only just been with massive storage and computation that we've been able to actually research and act on them so yeah um like I said in a post one year after GPT or so Chad gbt is out um yeah we're all still here and stuff so yeah it's it uh it's definitely not you know the boogeyman that's turned out to be it's it like I said it's mostly just collections of techniques that we finally have um the resources to be able
to apply to in the end but we could get Innovation if we can get better data and that's the next level right now it's dump everything and the kitchen sink next level is going to be be more Discerning about your data be more Discerning about the balance of your data and again to reference again Andrew Yang and I think I referenced him the last time I was here is the his whole premise now his whole startup Now is better data that you don't need Behemoth models you can do better with smaller data sets and you can actually do a lot better in terms of supporting
say industrial applications where you have very low signal to noise ratio then let's move on to the next stage like the novelty is great we can chat with these robots and sound pretty human but let's move to the useful part of this stage now true take that that's a quote we're going to make that into a little short it's like all kinds of t-shirts yeah a lot of issues I think a lot of nice Snippets for your little snippet series what do you call it uh the shorts yeah Snippets series I like that I think we're just we just are on fire today oh my
God you guys thank you so much super fun conversation I know that we could probably talk for another few hours about this but appreciate your time thanks for being here and make sure that you like and that you subscribe put in any questions that you have follow us all over the place reach out to ciq we're happy to chat with you more about everything HPC and get all of your systems supported and all the cool things that we do we love you very much we'll see you here same time same place next week bye thank you everyone so foreign [Music]
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.