Deploying Meta's Llama with Apptainer on GCP Optimized Rocky Linux with GPU webinar poster

Deploying Meta's Llama with Apptainer on GCP Optimized Rocky Linux with GPU

Watch Now

Join us live to watch CIQ's Forrest Burt demonstrate how to deploy Meta's Llama using Apptainer on GPUs, with Rocky Linux on GCP.

Speakers

  • Zane Hamilton, Vice President of Sales Engineering, CIQ: LinkedIn

  • Rose Stein, Sales Operations Administrator, CIQ: LinkedIn

  • Forrest Burt, HPC Systems Engineer, CIQ: LinkedIn

  • Dave Godlove, Solutions Architect, CIQ: LinkedIn

Transcript

[Music] yeah [Music] [Applause] [Music] he oh [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on powering the Next Generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and aper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source dude what's up how are youc was so weird while your like little thing was doing the intro I got nervous I think it's because we haven't been on for like a

couple of weeks live and I started get like yeah like the like the Jitters like the panicky I started trying to make sure everything was turned on it sounds different my headset today so I'm we're off yeah we're off yeah right I hear you I think we all have these issues right it's like where's the camera where's the screen where's the monitor where's my volume like it's just all over the place to see him what's that yeah you too I said I hope Forest isn't off today oh yeah that'd be weird well you know I have a sneaking suspicion that he has been planning and

scheming and looking forward to today for quite a while this guy is like yeah yeah yeah he I mean to like he gets so interested and curious about new things and new technologies and ways to put things together and so he has been deep diving into what is going on with this I want to call it Yama I know it's not Yama I know I know but in mind as I'm looking at it I want to call it Yama So today we're going to be talking about deploying meta's llama with aper on GC CP optimized Rocky Linux with gpus that isn't mouthful and it's going

to be rad it is a lot and I fully anticipated you having a llama behind you but I guess I was wrong oh my God I expected it I have disappointed all my fans out there I am so sorry it will never happen again it's all right maybe Forest will have a llama okay that would be amazing that would be amazing all right so where where where is this man bring him on yes Andy we got Dr God love too what amazing yes it's like a surprise Cameo this is kind of exciting good to see you guys you thank you very very cool great to

View full transcriptHide full transcript

see you all is this working it is working awesome great to see everyone Dave hey I decided to just drop in and see what Forest is working on it looks super duper cool I know I that's awesome I'll I'll actually I'll be I'll be right back oh oh okay so while he's gone Forest I think we we've talked about what these models are before and kind of what you're going to go into but I think it would be good to level set kind of the base of what is it we're talking about not just llama but in general what is it and kind of what

are people doing with it before you dive in yeah so good afternoon everyone good to see you all today uh yeah so this year has seen uh kind of over the past year in general we've seen a huge kind of rise in the large language model while there's hasn't really been um the same level of kind of development around you know text based Ai and these language models like this before um with kind of the huge explosion in generative AI in the past year and a half or so in general um we kind of saw you know this come out around text based content which

chot GPT first in um December so of last year um and since then the sphere has just kind of been getting more and more developed um one of the biggest things that's gone on uh kind of in the space is the race to create U models that are not proprietary so everybody wants now large language models and everything it's you know there are uh you know whole startups that are being funded just on you know rappers around there we go there we go made it he got it done from LOL CS to the llama's coming through for us he's got the whole Farm there sorry

to interrupt I just heard the llamas were requested so that's awesome magic dude magic see you still have young kids that's what I was thinking I was like man if my kid was still young I know I'd have one oh super cute thank you very cool um all right Forest startups being funded on these models yeah so there's you know there's whole uh like startups at the moment that are uh basically just rappers around the open AI apis um which is kind of funny uh but um since kind of like I said this generative AI boom started and especially since chat GPT kind of showed

us uh what you know text based generative AI uh can bring to us as far as capacity goes there's kind of been a race to free the entire llm space from kind of the proprietary World um like I said everything that you know open AI chat GPT all that's proprietary you can actually get uh access to those underlying model files and go like deploy them out on your own architecture so there's kind of been um over the past even like before chat gbt came out there's kind of been a little bit of a history of open- source language models that people have used um kind

of generally in the GPT space uh you know there's been GPT 1 2 3 four Etc there's you know kind of been development in that space for a long time um so there have been uh kind of open-source large language models before um like gptj and GPT Neo x2b uh were the two biggest ones and kind of the standard in around May of uh March May or so of 2022 um one thing that's interesting is that at the start of this year I saw tons of businesses kind of creating their own little come and train your AI build a custom version of one of these

models with your you know custom finetuning data and they basically offered gptj and GPT neox 20b um that was basically what there was in the space as far as open source AI that was kind of the standard of what people used there's a couple other models uh things like Bloom for example but in general uh what I kind of saw was that the space had mostly standard around uh standardized around those two GPT ones because they were the easiest to kind of uh they were they were um the ones that were already most kind of Suited to chat based applications and they were kind of

the easiest and some of the lightest weight to go and take and find tune on your own um even then they took quite a large amount of architecture behind them to be able to F tune uh and so kind of into this year uh there were a few different developments that came out there were a lot of places that kind of started to kind of seeing this chat GPT thing becoming bigger and bigger around um December January uh February Etc a lot of places kind of started to take these uh gptj and Neo x2b and build kind of a custom version of it um so

that like I said they could either um you know put it out there as like oh this is our open source version that we've you know kind of made better um or so that they could you know kind of put it behind you know pay wall and like sell access to it uh ultimately um like I said a lot of these it's it's basically just kind of providing an easy interface to get into those models um but there wasn't there was kind of a lot of places that uh you know you would start to like look I was trying to find this whole year I've

kind of been looking basically at you know creating a you know of course I'm getting in on the GPT craze too so I'm you know kind of working with my own projects in the space uh but kind of looking for different solutions and open source models I found frequently that you like I said around February or so uh there would be places that would announce oh we have this open source model it's amazing come take a look at it and then you get there and you know at the end of the whole you know model card it's this was on our custom you know proprietary

whatever chip you have to go and have one of these specific types of systems or you know get access to our Cloud platform to do this I'm like okay well this is not really the open source that I'm looking for um the whole Space kind of really switched up uh when in about March April or so uh meta open sourced their llama AI model um basically their original intention was to only provide access to it to like researchers um academic institutions places like that like they had kind of a vetting process for who they were giving the actual model code to um but then the

model code got leaked like a week after this out on the internet and at that point the cat was kind of out of the bag um this uh there was a and there are other kind of I should just mention really quickly to digress there are other kind of attempts in the space that I looked at um smaller models that turned out to be you know basically not you know usable for commercial use um stuff that uh there's like one called dolly that I looked into for a while that um ultimately uh just kind of got out moded when llama came around uh so anyway

they announced they're going to have a vetting process and that it's going to all kind of you only can get access to this you know by getting a hold of us etc etc uh and like a week later it leaks and suddenly everyone has access to llama so then the big race after that I was kind of like okay well you know I'm going to hold off on kind of experimentation more with these uh you know two GPT type models because uh you know we'll just kind of see where the space goes for a little bit and kind of what new standard maybe emerges out

of this uh and sure enough the race was on after that was leaked to produce a commercially acceptable version of that code that people could actually go and base their models on I'm not an attorney so I don't really know how you take leaked code and turned that into a I don't know you know I'm I'm also you know only really an amateur AI engineer uh so I'm not sure exactly how people were able to somehow create model that gave basically the same performance same functionality or close enough to this leaked model um to be able to release this but that was called open llama

and that came out about a month and a half or so after llama itself was leaked um so I kind of looked into that but I was kind of a little bit like well open Llama that's a little you know that's how do you desani leaked code like that you know realistically um and so other projects kind of pulled my attention away uh I kind of ended up you know just looking at you know just other things that you do kind of e and flow of work uh I kind of also like I said put this aside just to kind of see what new standard

kind of emerged from you know the cage fight that was going on between all these um open Llama like I said dolly all these different models were kind of competing at once um kind of into around June of this year uh and the space kind of stayed that way for a few months um it was kind of you know there wasn't any like huge massive new like huge news that came out until just about like two months ago or so um when meta came out with llama 2 llama 2 is their um essentially genuinely and I should mention that there have been people that have

kind of called them out for calling this in general an open- Source model uh because they didn't share um well they've shared the code they haven't shared like the underlying data sets and stuff like that that um it was trained on uh of course they you know we never did quite find out I don't think how many parameters chat GPT has in it I think last I heard some people were saying that the its trillions wasn't quite true and it's actually hypothetically like 800 billion parameter models that are all kind of fine-tuned um for different subject matters all kind of you know Wizard of Oz

behind the curtain together um but anyway uh So Meta 2 or llama sorry llama 2 is released um and like I said you can some people have kind of take an issue with how uh meta calls it you know open source but in general it's a really like robust chat llm model that they uh have essentially provided for free for any type of use the only caveat in their terms of use because the first thing you know I got back into this sphere I also see LL or llama 2 and I'm like oh it's open source wait what and um kind of looking through their

licensing agreements and stuff it's you know pretty permissive you can basically use it for whatever but they do have a caveat in there that says like uh if you have over 700 million users on your platform or like if you're over $700 million in Revenue uh something like that I think it's Revenue you have to go and negotiate a separate license with them to use it and they essentially did this so that it's kind of funny they've enabled to you know basically everyone on the small scale to use it but they have cut out like Google their larger competitor AWS these larger competitors from taking

it and using it because you gota you gota license that with them so um there's a lot you can do uh with this new AI model it is super robust it has as far as I can tell all of the Care and attention of a professional engineering team behind it basically as opposed to kind of I shouldn't say as opposed to but kind of some of these other projects um were not like the models themselves but like some of the GPT things the uh the ways that a lot of people were deploying them wasn't actually like the model itself but it was basically kind of

like some basic scaffolding code that people had put together to enable like parallelism on gpus with it um and so should mention this whole time I was kind of trying to containerize these different things with obtainer along the way um ultimately I kind of ran with like the GPT ones I ultimately kind of ran into um problems even before I got to containerization with just getting like the basic kind of you know GitHub readme type stuff up and running on it um like I said those models had kind of aged and the code bases that they were based on were kind of like this was

a one-off I wrote a year ago to enable you know this model to run that I decided to put out on GitHub that then kind of you know enough people over time contributed back into that it um it stayed a float for a bit but it was definitely by the time you know llama was leaked it was kind of starting to become obvious that uh there was maybe kind of a need for something newer I was like getting cutic compatibility problems different things like this um so it's got it I just I really like how they it's one of those things that I find uh

my experience with llama 2 so far has been one of those things that uh kind of just works if you know what I mean it you know I was able to pull the GitHub I was able to get the link from meta I was able to use it really easily it just uh you know pulled in the models just fine it's a so the actual the I should back up a little bit the process of actually getting the models themselves from meta um you go to their site they've got kind of a form that you fill out that tells you a bunch of or that

gives you a you it's kind like name your organization that type of thing what models you want uh they've made available I think a seven billion a 13 billion and a 70 billion it's either a 13 or a 20 billion in the middle but it's a 70 13 when I looked at what you sent out I think it is 13 you're right 13 okay yeah it's a seven a 13 and a 70 billion parameter version um and they actually have two code bases you can get from them uh there's the Llama model itself which has basically an un trained version of the model that they

have just trained on their training data but they haven't fine-tune for chat applications and then they have another version of it that they've done some fine tuning on um to enable to kind of make it better for uh um for chat applications essentially um the trained one you don't have to do what's called f shot prompting um f shot prompting is where you give you actually give the llm a few examples of like question answer question answer question blank fill in the blank and you have to like basically prompt it beforehand like let it know what type of response you want um their base model

you have to do that for queries but their chat one will just take you know what is this and it knows what to do with it so that's their base model they also have a code version of it that I think they also have like a python optimized version of that that they've like specifically fine-tuned for python um and so yeah they they've got a whole bunch of stuff that you can get from them um once you submit your info pretty quickly like a minute or so later get this link in an email that it says go pull this GitHub paste this link into this

thing um you know do the git pull do all that type of stuff um here in just a little bit I'll kind of jump onto my instance maybe and show you guys uh kind of what that looks like um I probably won't pull the models because they're significant um that's a lot of data but I'll kind of show you what the end result of it is uh you get in the end your like GitHub that you pulled with all these different you know basically model weights files stuff like that to work with um and after that the sky kind of limit like I said I

uh I got this working uh my environment for this was I went out to the Google cloud and spun up one of the um custom gcp optimized Rocky Linux images that are up there uh that we've gone and specifically optimized with gcp to run best there so I spent up a version of that uh with one of their um what is it called it's an A2 Ultra gpu1 G instance I have like this one instance in the Singapore region that I'm able to get access to in general I can never get these things to spool up in any other place and when I requested like

a a variety of uh I'm running this on an A180 gab uh GPU when I originally requested availability I got basically one in Singapore and I'm like yes my GPU uh so I was able to get one of those instances um while you sometimes you know availability can sometimes be a um you know always in the cloud you can sometimes not get a GPU all the time uh for the most part it's pretty reliable and uh I've been on there quite a bit since then just kind of tinkering around doing different stuff um like I said uh gcp optimized Rocky Linux has worked great I

was able to uh take basically just a standard Nvidia uh Cuda and Driver install bring that up on the instance um and have my GPU and everything working really easily um and after that like I said I attached some block storage to this so I could basically have two terabytes of scratch space to go put all the model files containers stuff like that into um spun this up and uh like I said did that whole process of getting the link from meta pulling all the models um and then getting all that kind of situated uh once I've kind of gone through and figured out the

basic steps that you have to do in order to get the uh model up and running in general just on you know bare metal if you will I went ahead and uh translated those over into an Apper overall it was really simple in general to get up and running it's not a very complex um definition file it's basically just kind of installing python 3.9 so you can use that uh installing some requirements files doing a pip update and then just uh it it's prettyy simple we'll take a look at it here in a second um but before I kind of maybe hop onto my instance

are there any questions or anything like that any commentary I think that was the one question I was going to ask was how hard was it to get into an Apper container but you it already yeah I I'll show you here in just a moment it was uh it was surprisingly easy I was very pleased let me clear this exp something yeah yeah right was very thorough than like every question that came up I'm like what about this one and then you just like went right through and like answered everything so we're gonna get like a a what is this so you said it's not

really a demo but we're just going to kind of get to see inside the container or what's happening here I'll show you guys running it on my instance really quickly I'm just G what my workspace looks like after all of this um like I said I kind of just have been running this on scratch face I've also got this maybe if I have a chance I'll show this at the end I've also got this running on our fuzzball platform at the moment the same example that I'm going to show you here um which has been kind of fun to Tinker around with so maybe I'll

show that here at the end uh but yeah so let me just jump onto my command line here really quickly and actually let me make the text bigger before I do that let see here while you're doing that I was told my microphone was a little a little bit quiet so I turned it up what oh sorry I can hear you what so you made a reference to um llama being different than chat GPT in like a few different ways but I don't know that might be interesting I guess you're going to show us so we're going to be able to see it so well

I'll show you like just kind of like what the model looks like running um GPT stand for generative pre-train Transformer I think which refers to essentially a specific architecture of model of large language model there are different ways that these are built um I I think it's thank you I appreciate that um I think on some level it's probably fair to say all gpts are large language models but not all large language models are gpts uh um GPT just happens to be like the architecture that open AI has found kind of works really well and so there's been a lot of work around kind of

generative pre-trained Transformers in the field um gosh and I hope that that's actually what gbt stands for uh it's a very intuitive name yeah total sense yeah like I said once again it's um you can kind of describe it at a very high level that all gpts are llms but not all llms are gpts and so there's um just underlying architectural differences uh between kind of you know between llama and probably what open AI is doing that make them uh kind of distinct uh families of model and I don't I haven't like read an incredible amount into that it's possible that llama takes a little

bit more from the GPT side of things but I imagine that open AI you know that's kind of a lot of their more in-house stuff so I imagine meta probably has their approach and I've read um the paper is very interesting and I wish I could remember the details from it but they're not coming to me at the moment I have a copy of it here but um they have all paper where they go into how their architecture works and it's uh yeah there's essentially just fundamental technical differences between that and like how a GPT would work um and of course this is open source

so we can take this and I mean I can interact with it here on my uh see if this works yeah I mean you're not sharing anything yet if I don't know if you're still trying to figure that out but we're not seeing anything Alex is my I just started sharing you sharing okay it I just started I just yeah I just started trying it cool um so yeah uh one thing that's interest like I said chat gbt uh one another big difference is that you can only really interact with it through apis or through kind of proprietary portals um you know you're asking it

something through their web interface uh they are I'm gonna make this just a tad bit bigger um you know you're interacting it through API calls you're interacting with it uh um through their web interface that type of thing it's all kind of proprietary there is no like take the chat GPT model file put it onto a Unix file system like we're looking at here and you know interact with it directly this is you know meta just gives you a GitHub they give you all the files I say here this is you know the dependencies that you need in order to run this on stuff and

uh we will kind of hop into what we're doing here so my workspace so far I'm uh here in my instance I've got one of these Nvidia a100 sxm 80GB gpus um you can see that it's mostly unused right now it's just kind of sitting there idle I'm pretty sure that I could do this on a smaller uh GPU but I just haven't got around to doing scaling tests up under the 13B and 70b versions yet so this is just the 7B model but I think even that takes like 14 uh gigabytes of GPU memory um they make reference to a like a model parallelism

value that's either 1 four or eight depending upon what model you have and I haven't quite figured out if model parallelism is essentially the number like uh if the model parallelism number is just kind of the fancy way of saying how many gpus you need um but in any case I'll get around to figuring that out uh the test space here is just kind of where I was testing out um you know demo beforehand to make sure everything worked this is the Llama Barn which is where I keep the various my yeah it's the uh that really what it's called is the Llama Barn I

love that yeah yeah you'll notice that the the instance host name is BTS llama Farm and then you'll notice uh that's at least the mount I think that even on I think even on gcp I might have named that like llama barn or something like that so great yeah so just kind of you know little fun techy host name humor there haha do better than that that was uh yeah anyway sorry bad joke there I I'll laugh at my own humor um so we've got code llama llama um obviously a couple different directories and stuff here if we look into this um this has got

all of the uh base llama chat model stuff in it if we look at this this is all of the code llama stuff so you also have the instruct version of um their code one which is essentially fine-tuned for like to give instructions or to like take instructions it's kind of um I hav quite figured out exactly what the distinction is that makes something better defined tuna for instructions but there is a there are a lot of places that we find tuna model for something like that um so in here um we've just got a bunch of different stuff uh the example scripts that they

give you um just kind of the requirements.txt ETC uh I'll spare you looking over the read me it's just rather dense would be uh probably a little difficult to citate in this small frame so I'll just show you the end result of it um you could just ask the model to give you a synopsis I could the uh what's cool about this is that it's all um you python code queries against it so you can just essentially um write python whole scripts Etc that call against the model so it's quite fun um where was I at oh yeah I want to C the lumo barn

uh okay B that file uh sorry so yeah so this is the uh the actual container file in the end that I put together to get this all running uh like I said I kind of just um took what they said to do in their read me and a couple of just best practices around you know updating stuff like that and uh put it all together to make this work I was honestly really impressed by how easy this was to get containerized in the end I basically just took my um uh like uh bash history put it into an obtainer cleaned it up a little

bit to format you know added the environment stuff like that and uh it it's just super simple it works great um these days when I built what's up Dave I was just going to say look at that FR line that's that's a pretty cool uh from line up there um in the uh header because it's nice to see uh you know official Nvidia Cuda installations on Rocky 8 yeah I switched over I used to just uh normally base my containers on um just like the Rocky Linux image and then do AA install and like well Nvidia drivera install Etc as much as you need to

do in the container over the top of that I eventually kind of figured out that a lot of things need this qnn um and with the development version of this container that they have out on Docker they basically include all the runtimes uh not the run times but like all the development headers all the stuff that things expect to be able to compile against that I usually um was like trying to install Cuda in order to give like make available um so these days yeah I prefer to just use their official images they're pretty solid um and like I said everything's based on Rocky you

can go see the um uh like Nvidia Cuda install instructions and all that and they've got a whole section for Rocky 8 and a whole section for Rocky n by name so and CNN used to have some weird licensing requirements too so this is kind of like uh a really nice way to do this I can I can talk about that at length at some point but probably not right now since you've got so much interesting stuff to cover that's not just licensing it floats back to me that my old module file system deployment on the last cluster I was managing had a separate module

file for Cuda and then a whole bunch of separate stuff for Q DNN that had to be loaded so yeah having it all just in one base container is fantastic uh within an obtainer usually uh I actually edited this in the one in my article to make it a little more obvious the post section runs first and then environment adds basically this to the environment that's present in the container once the build is done so we're going to skip over environment for a second and go to post um like I said this is pretty simple we just do a dnf update to um I'm not

particularly concerned about like the CUA compatibility things being updated Etc stuff like that um I'm not sure if they I've never quite ascertained honestly if they do this in like um like a local type install or like a network type install like that will update this when you dnf update um in this case I just chose to do that because this is a pretty new model it runs on pretty new gpus and it uh pretty much expects the latest Cuda especially because it wants to run on python 39 um so I didn't have update the container just to update that plus all the dependencies that

type of thing um you know make sure I can actually find the get python 39 uh 39 devl packages and stuff um go ahead and install those um make oh sorry I thought that was make deer llama farm and make deer llama Farm double repeated for a second um no so we make a just kind of a working directory to store everything into and CD into it um and then we just go ahead and clone the llama and code llama um code bases these are not the models themselves uh meta only distributes the models um from that link that you have to get by filling

out their form and stuff and they say don't go post these anywhere else for people to ingest so this yeah so this container is mostly a wrapper um that allows you to take the models that you have on a large file and you wouldn't in general uh while sometimes it is useful and kind of the right thing to do um generally you find that very very large containers are kind of unwieldly they're difficult to push they're difficult to pull they're difficult for um you know systems to ingest as an image that type of thing um so usually to keep your containers as minimal as possible

and storing model files inside of them and like pulling those gigabytes and gigabytes and gigabytes of data whenever you need to um use the model is less efficient than just storing everything on dis and like mounting it to a container like this mounting it to an instance and then just binding it into a container like I'll show you um this brings up an interesting I think this is an interesting thing to think about though I mean um you know it generally I'm sorry to jump in I but um generally like I think that it's uh it's a good kind of practice or a good thing

to to think about your containers as being like you have you've got your code you're actually you know whatever you're going to functionally run inside your container and you've got your data outside the container that you're going to operate on but it always gets a little hazy when you get into these use cases because then you get into like a model and you're like wait a minute is that data or is that code it kind of is functional it kind of is like the thing that you're going to use to do stuff but it kind of isn't too because it's kind of just a bunch

of Weights that you could swap out for a different set you know so it's it's like um it's always kind of a gray area I think what do you think about that I mean I know just like from a from a technical standpoint it's it's in this case it's better not to have it in the container because you want to move the container around um in general uh I'm sorry response and it just kind of popped out of my head um how flexible are these models to be able to swap in and out with different uh you know like once you've got once you've got

llama installed um how tied to the version of llama or the version of whatever framework you're using is a particular model or are they usually very very flexible for the most part it's pretty static um you can pretty much uh in get one version get its dependencies together in a container um and then be able to operate on that uh and yeah to touch on like data versus things I think that yeah just like you said that's a huge problem just in general whenever we're doing use cases what needs to go with the container what needs to be ingested in from somewhere um it's always

kind of that fine line of what like I said is actually functional versus what is just data um in this case models uh would just how much like they used to be pretty small I mean a few megabytes you could pack it in the container if you wanted no big deal but with you know models getting five 10 gigabytes or larger um even if you're you know at a certain point you could um you know I sometimes recommended like packing them into a layer on top of the container um like an extra like an overlay image basically but it's just uh a lot of the

times easier to just put them on dis and have something that can persistently mount that to a cloud instance or just have something similar on Prem um ultimately uh AI is truly another HPC use case so there's a lot of kind like oh is is HPC I've heard people asking oh well is HPC ready for AI HBC and AI have already been the same thing for a long time so a you know uh the there are larger use cases that go on um even just in HPC that have you know more storage more data going into them Etc um and so AI while it still

has some room to scale its biggest problem is just like getting larger and larger gpus um definitely uh kind of interesting deployment strategies on how you actually get um the models out there but yeah the uh as far as like swapping them in and out um I would imagine you can probably like this uh container I imagine as it sits is probably good for um you could probably hot Swap this with the 13B the 70b and it would work the same way because they all have basically the same dependencies um and so yeah they uh they tend to be fairly flexible sometimes used to run

into a lot of issues with like older gpus because sometimes like if it's got a newer Cuda version it won't like older stuff um but all in all I've always found uh well the the old GPT ones were kind of starting to get a little bit aged but um stuff is Cutting Edge it tends to you know have a period that it works pretty well for um and meta this llama stuff is pretty actively updated as well so I anticipate stuff will stay compatible for a bit um yeah so just to uh finish up what I was saying about the container here um like I

said we go ahead and pull in the actual code bases themselves that we need uh so we bring in you know like I said the functional part of this container that isn't just the data we're going to end up ingressing in the end um we see the end of that file it's literally just an upgrade pip so it can find um the latest versions of torch and uh the different QA compatibility wheels and stuff like that um upgrades pip uh upgrades kind of just um like I said the uh the package repository when you do that for python um you just do a pip install

dase dot which pulls that requirements file and that type of stuff in um feeds that into python it reaches out finds torch all the dependencies uh puts it into a um python environment and at the end we do a dnf clean just to make sure that uh we're not storing a whole bunch of large uh Cuda RPMs and stuff like that um and just to in general make the container smaller because that's a good thing to do whenever you get done with a container uh so yeah we go ahead and build this on the architecture um and I'm going to go over to my actually

test space to show you this well actually whoops yeah so we've got uh this is the result of building this it's a 9.2 gigabyte container in the end even with just that um uh even without all the model files in it uh just because GPU code and libraries and stuff like that those containers tend to be weighty so you can see how if we were putting you know another five to 10 gigabytes of models pretty soon we'd be moving 20 Gigabytes over the network whenever we want to Ingress this um so instead we're sitting here on this uh like I said mounted volume to actually

go ahead and run this container here um we're going to go ahead and scroll back up in the history to get the command to make sure that I get it right there we go so this is is kind of a long command um but once we've gone ahead and done the well let me show you when we build this container we do an obtainer build d-nv llama 2.if ll-2 dode and it's just that simple to put this one together um we have to do the uh NV stuff so that it um kind of correctly finds what it needs to find out there it's always a

little ambiguous um sometimes uh I find that these container builds will not find what they need to find without that generally when you use the base container from Nvidia it's pretty resilient um but just to make sure we're getting um all the most accurate stuff we go ahead and build that yeah I I saw that and I was I was wondering about that you probably anticipated I was going to wonder like about the uh Das Das NV during build so it it it it like it's my understanding that typically when you're building this kind of stuff there's stub files in the Nvidia containers that it

can uh link against without actually having the actual libraries I wonder if you if you if you uh add that d-nv option do you know so what what that's doing obviously it's is it's grabbing the the driver libraries from the host system and putting them into the Container at at build time now and I guess presumably linking against them do you know if that's going to uh negatively impact the portability of this container at all um it's possible that it does I I'm half thinking that this NV and the dnf clean all were holdovers from an older version of this when I was installing the

Cuda manually before I just decided to switch over and use the base container in this or for this one um okay but it's also possible that I just added it out of habit and forgot to take it out so that's a good um that that's good to point out Dave it's possible you don't actually need the NV here especially because we have that base container yeah it' be it'd be interesting to see so yeah may have just been that I added that of Hab I think it's uh for the most part um does it uh does it inject the GPU device files into the Container

if you don't use d-nv because that's the only one thing that I can think of is that it might be like uh you might get to like the python section it'll just say no compatible GPU or something like that found because it doesn't have the device file but anyway it probably warrants maybe more testing yeah I don't know that it injects the I don't know that that adding the --nv option will inject the device files um it could because that's a Rel that's a relatively new Option with the build command um so it could be that it sets up device it where it buy mounts

device files into the Container but I'm not sure I I I kind of don't think it does but I'm not sure yeah yeah well Envy has always been a little bit of or the Envy option has always been the vean of my existence for a while now maybe one of these days I'll finally I'll finally get it um anyway this builds um we get this large container out of this um and then we're free to begin doing this right here um so just to kind of explain this entire line uh first off and some versions of obtainer um it will bind the current working sorry

I thought that switching Windows made it actually switch Windows on this thing um you'll find that in some versions of obtainer it binds the current working directory that uh of the container itself sometimes it binds the working directory of the script being run um so it you might have to mess with it a little bit depending upon what version you have in order to get the bind to work right um in this case we're just is binding in this Mount llama Barn in the mount llama Barn in the container I'm to make sure that's there I use- MV like I said once again it's potential

this isn't necessary um but in general that option always kind of confuses me a little bit I'm going to switch this to uh llama 2.if for this working directory that we're in uh this was from what I was using the test space earlier and then we have this torch run a long command here that basically just kind of is their standard way of running these scripts um so we're just running this example chat completion script which has just kind of some uh random queries that they put in there I think it asked me what's a recipe for mayonnaise what's uh like how do you get

from Beijing to New York but explain it in emojis which I was amusingly which when I first saw I was very amused to find showed me that the terminal has Emoji support which I didn't realize uh wait wait wait wait for us I feel like you are teasing us right now no I'll show you I'll run this in a second this is okay yeah okay thank you so yes okay so you're you're telling us what you did and then you're GNA show us yeah I'll go ahead and run it and we'll see if this uh we'll see if I I switched everything over that I

need to switch but yeah the um torch run and proc per node just won because we've only got one GPU we're running this example chat completion script oh yeah there we go um this right here is the check Point directory oh wait I put my tokenizer oh live demo okay I think that's right okay so yeah um we have to provide the checkpoint directory which is basically just the directory of the um uh where all the model files are at so like I showed you earlier with the Llama bar and the Llama file and then all these directories inside of it these are the directories

that you get from meta when you run their download script with the link that they give you um so you get basically all of these uh directories that it downloads everything into and then you can just reference them with checkpoint directory like that the tokenizer path this took me a little bit to figure out I got kind of tripped up on this because I was like where does the tokenizer come from but then I realized I think the tokenizer comes from when you uh like pull the download from them so I was trying to run this container and it kept saying it couldn't find the

tokenizer and finally I was like oh the tokenizer isn't actually there um so you might have to get that into a certain place ah there we go so let me scroll back up and finish what I was saying and then I'll show you these results um uh Max sequence length and match batch size are just kind of random uh AI uh options that kind of tell it what type of response to give how long Etc um but you can see that this worked so we've gone ahead and run this uh it goes ahead and finds the GPU it starts to initialize the model um and

once again we're using obtainer exact here so I'm just calling basically this command on this container right here with like the bind and envy options applied to it um but you can see we get these initializing model parallel pipelines um it loads in a pretty quick amount of seconds because we're here on this uh a100 and then like I said we get kind of these random queries that uh meta just I haven't changed their default input script but just um you know recipe for mayonnaise I'm going to Paris what should I see always answer with a h coup always with emojis how do you go

from Beijing to New York so you can see the terminal does actually have Emoji support um they've got the you know the the initial prompting you are a helpful respectful and honest assistant you can sometimes get open AI like chat GPT to divulge its initial prompting text that it starts before every conversation that's like you are a large language model created by open AI it is your job to helpfully answer questions and not to be you know this list of you know harmful and ethical racist sexist Etc um they've kind of made it so it's more difficult but there was a while where everyone was

like oh look you can get it to tell it like it's one like the one shot prompt that they give it beforehand um you write a brief birthday message to John stuff like that who's John John who oh I have no idea I thought it was somebody's birthday no it's not no definitely not it's uh like I said it must be a John at meta somewhere that's just their example um it's always kind of funny the text that AI produces it's always like the most like versy version of that text that something can be like you ask it you know write a cover letter and

it'll give you the most cover lettery cover letter you've ever written it'll cover all the bases to the point where you're looking at it like you know it's almost like too good like you got to leave out you know you got to remember to forget something in it you know the uh but thank you for your consideration you know that I did this that there's got to be some type of flaw in it um I noticed that the uh recipe for mayonnaise was a legit recipe was it yeah I know too let's take a look yeah what's funny about these is that uh sometimes they've

really struggled with logic um so like a lot of people have kind of said that chat GPT has gone downhill over this year which is not entirely out of the realm of possibility as they add more like content controls kind of stuff like that the its ability to kind of maneuver in a response gets more and more it thinks it at least is more and more limited um and so people have like asked it at the start of the year can you pick out these prime numbers and it does it perfectly now people claim it can't um spatial reasoning is really difficult for large language

models like try asking one uh you have a book a sphere and an open cylinder how can you stack these items most efficiently and it'll say like stack everything on top of the sphere and the cylinder just like in a way that would is totally illogical um so it's fascinating how it's not just about you know human-like text but uh it's incredible how much you can encode just in text itself like this whole you know the concept of logic spatial reasoning all that it's it's not just text it's well I wonder if that's really a failing of the model or if it's more of a

you know if it's more of a an artifact of the data like if you give you know if you take a random sample of people off the street and give them a list of numbers and say pick out the primes um you're going to get a lot of people who can't do it and so maybe it's you know correct for the model to be unable to do it exactly so it's yeah so there we have a we definitely can see how large scale um sociological and psychological factors from The Human Experience can be baked into these models without us even realizing it so you know

this is one of the greatest arguments against giving these things too much control and stuff is because there really are you know even Beyond just like the obvious biases that things can have there really are like large scale patterns of human behavior that are unintentionally encoded into these so it's it's you know they they are a fascinating Black Box to jump into while I have a couple of minutes um let me do this really quickly and I I will point out that in my previous example I did not say say if you just gave me a list of prime numbers or numbers and I might

be unable to do it because of course I could you know everybody in this call could it's just the people on the street prove this Theory yeah don't give me a list please don't give me a list next time we're hanging out Dave we're finding a street doing this can you like preface by explaining what a prime number is again it's been a while no just straight into it JK I think I probably figure that one out I think some of like the largest distributed computing projects in history of people like pooling their computing power have been like the prime hunts I think sometimes in

recent years those have come under Fire because there's some of like the most inefficient Computing possible because it's usually just using like whatever Back Room desktop PC people have so it's the power for performance ratio is a little low um but here in this uh right here the power for performance ratio is definitely not low just jump into here this uh article and or this webinar is also kind of in an article um that we've got posted up on the ciq blog that I put together so feel free to take a look at that to get the model file the examples um kind of the

uh the basics of what I've shown here let me hop onto this quickly I've uh seen some updates recently on what we're doing with our platform here with the ability to persist storage and do like persistent volumes that I'm very excited to apply to this um but for now just to show you really quickly so I'll I'll riff a little bit while you're bringing that up so just uh I think that you're bringing up now fuzzball MH and so for those of you who have not been religiously watching each and every one of our webinars fuzzball is the orchestration platform for HPC that uh you

know orchestrates based on jobs rather than services and is optimized for things like AI workloads yeah absolutely um go Rose there we go Rose thank you perfect yes yeah so fuzzball is a really cool platform uh I'm always thrilled to get to show it off um we've this is basically the exact same thing that you just saw me do here but as a fuzzball workflow um so I can show you we've got you know our definition over we well actually I'll start with the jobs uh this workflow is pretty simple we just create our llama Barn storage volume we pull our container image which is

the same image that I just showed you um us running out of on this gcp I basically just um took that image uploaded it directly to our uh our artifact registry and pulling it into this uh you'll notice that in from S3 I'm pulling the uh the model file and the tokenizer like I said this is kind of what I was saying about not wanting to Ingress these because like always bringing in this 10 gigabytes over the network uh is going to you know be slow um so like I said I'm very excited to test out some new stuff we've had recently with uh persistent

volumes more to kind of Tinker with that um but this brings in the tokenizer and the model file um we then basically just untar both of those get those in the right place and then we run like I said that same bit of code that we just ran in the last um uh demo um and you can see in our workflow editor we've got that torch run chat completion basically like same type of thing going on there we're requesting uh four CPU cores 58 gbes of memory and a GPU this is going to land us on a v00 so I think that 14 um gigabyt

of memory Works within the confines of a 16 GB v00 and we can of course and that'll uh well we won't have time to see this finish um you can see that uh this workflow runs on fuzzball and like I said that um I'll just show you the previous result that I was on there we can see that in here we have that same what's the recipe of mayonnaise I'm going to Paris that type of thing so that's uh that's been my experience with containerizing large language models using uh Apper this year it's it's been very exciting like I said there's been a lot of

kind of just sitting and waiting for things to um kind of move and kind of create standards uh but um kind of with this huge recent success and getting llama working and getting it containerized and stuff like that I'm super excited to kind of like I said dive into my my own projects around AI more especially now that we have uh what seems like such a solid standard to base off of so it's very cool Force are you gonna post somewhere in your blog the actual definition file that you had there it's in there right now excellent y you go copy that yeah like I

said paste it right onto um you know wherever your gpus are sold and uh get that uh get that up and running of course like I said I and I think all of us we in general highly recommend that gcp optimized Rocky Linux it like I said Works fantastic and uh I had nothing but smooth performance the entire time while I was deploying an entire AI workload on it excellent thank you for showing that for us Dave it's always good to see you thanks for being here too yeah thanks for letting me hang out this has been a really really cool demo for us absolutely

thank you for joining Dave and uh thank you for being here Zay and Rose to host as well yeah that was awesome all right well I'll ask my millions of questions later um and for now though you guys we are done so make sure that you like and subscribe share this with friends enjoy llama call us if you want to talk about fuzzball and uh we will see you same time same place next week enjoy your day thank you everyone bye [Music] everyone

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert