Fuzzball videos

Announcing Fuzzball Federate: Deploy anywhere

This webinar marks a busy day for Fuzzball with three announcements. Fuzzball is now listed on the AWS Marketplace, initially as a tech preview for people who contact CIQ, so teams can install it in their own AWS account rather than waiting for an on-prem deployment. Fuzzball Federate, also in tech preview, joins multiple Fuzzball clusters into what looks like a single cluster, whether they sit on premises, in different AWS accounts or regions, or a mix of both, opening the door to cloud bursting. The third addition is a workflow catalog of ready-made templates for applications such as Jupyter notebooks, ParaView, LAMMPS, BLAST, and Stable Diffusion.

The demo tours the Fuzzball interface running on AWS: shared accounts and workflow history, the visual workflow editor and the YAML it builds, push-button provisioning of GPU and non-GPU resources plus persistent and ephemeral storage classes, and a Fire Dynamics Simulator workflow with MPI that a colleague assembled quickly and the presenter reused from a shared account, visualized through a Smokeview virtual desktop.

The Federate demo submits ten hello-world workflows from the CLI and shows them land on both AWS and on-prem clusters. Audience questions cover other clouds, moving checkpointed jobs, splitting workflows across clusters, and how Fuzzball tears down cloud resources when jobs finish.

Key takeaways

  • Fuzzball is now on the AWS Marketplace as a tech preview; the listing is visible to people who contact CIQ to be onboarded.
  • Fuzzball Federate presents multiple clusters, on-prem or across clouds and regions, as one cluster, with users submitting exactly as before.
  • A new workflow catalog offers fill-in-the-blank templates for tools like Jupyter, ParaView, LAMMPS, BLAST, and Stable Diffusion.
  • Push-button workflows on AWS provision GPU and non-GPU compute plus persistent and ephemeral storage classes right after cluster start.
  • The Fire Dynamics Simulator demo shows shared-account collaboration: one engineer built the workflow, another reused it the same afternoon.
  • Fuzzball Orchestrate runs on a small always-on server and spins cloud resources up per workflow, then tears them down when jobs finish.

Questions this video answers

How do I get access to Fuzzball on AWS?

Fuzzball is listed on the AWS Marketplace, but at launch it is in tech preview, so the listing is only visible to people who contact CIQ and are added to the preview. Once onboarded you follow the Marketplace instructions to install Fuzzball in your own AWS account. CIQ plans to open the listing more broadly later.

Does using Fuzzball Federate require changes to my workflow files?

No. Whether a workflow runs through Federate depends only on where you submit the fuzz file, not on anything inside it. You log into the Federate cluster as a CLI context or in the web interface and submit as usual; Federate then places the workflow on a member cluster according to its rules.

Can one Fuzzball workflow split its jobs across different clusters?

Not currently. A workflow runs start to finish on a single cluster, so Federate decides where the whole workflow lands rather than placing individual jobs. Shared ephemeral volumes between jobs are one reason this is hard, though the team has discussed it as a possible future direction.

About this video

Fuzzball will deploy and optimize placement of your workflows across disparate clusters that reside in multiple regions or even across on-premise and cloud providers. This optimizes your environment based on data, compute and storage requirements and gives you the flexibility to develop in the cloud and deploy on premise or develop locally and deploy to the cloud for scale.

Watch the webinar to learn more about Fuzzball Federate and how you can get started with it today!

This video is part of the Fuzzball playlist. Browse every CIQ video by product and topic.

Transcript

Good morning, good afternoon, and good evening wherever you are. Thank you for joining. At CIQ, we're focused on powering the next generation of software infrastructure leveraging the capabilities of cloud, hypers scale, and HPC. From research to the enterprise, our customers rely on us for the ultimate rocky Linux, Warewulf and Apptainer support escalation. We provide deep development capabilities and solutions all delivered in the collaborative spirit of open source. Okay, wonderful. So, we are back, Dave and I, and we are so excited to talk with you guys today about fuzzball. Um, again, thank you to everyone who is on the fuzzball team and has been doing all of it.

Um, okay. So, I know that we have some things to show us, Dave. So, is there some context you want to kind of give to people first and then we're going to dive into the juice? Let's give the context. Yes. So, um today is a big day. Today we have a bunch of announcements to make. And so, um that's one of the reasons for this webinar. It's one of the reasons, you know, we're here and everything. So um you know one of the things that we can talk about is we have released fuzzball in the AWS marketplace. Um and so what does that mean? Should we do a little dance?

We maybe should. I don't know what the AWS that Thank you, Rose. Perfect. Um yeah a so and so so this is awesome because um we've been talking about fuzzball for a long time and we've been talking you know we've been developing it for for a while as well and every time we talk about fuzzball to somebody you know at a conference or something they get excited about it and what we hear often is when can I get my hands on it when can I try it out and that the answer to that question has been a little bit like well because we've been developing

it mostly for on prim and we've been focusing even though this is a cloud you know offering as well and we have our own um AWS cloud uh instance running it's been hard to like let people know how they can get their hands on it. So now there's an answer to this question and the answer is um you'll go to the AWS marketplace and you'll search Fuzzball and you'll be given instructions for how to install it in your own AWS account and get your hands on it and try it which is amazing. That is amazing. Okay, I know that this is like our proprietary stuff. So I you know people are going to ask you know like can I play with it for free?

View full transcriptHide full transcript

But it is it's not it's in the marketplace. So it is something that you would pay for. Yeah. And right now just with a little caveat too, right now we're just starting this and so we're in tech preview. And so right now the AWS marketplace listing will only be visible to people who contact us and are added to the tech preview. But that is for, you know, a limited amount of time. And, you know, we're going to go ahead and onboard some folks and do the tech preview. And then we're going to open this up and it's going to go to where you can just go to the marketplace, look for fuzzball, install it in AWS.

So this is huge. This is a very, very huge announcement. Okay. I'm so sorry. I always have a million questions. Who do they contact? I mean, I'm sure that like inside AWS it says contact CIQ. Boom. They can also contact any one of us. Right. Right. And if you go to the CIQ website, there's a little contact us button. and you just click that and send an email and we'll be happy to, you know, talk to you and get you get you ready to roll. Oh, so that's one of three big announcements we're making. Oh gosh. Okay, zip. Carry on. So announcement number two is that we've been talking for a while about something called Federate.

And Federate allows you to take more than one cluster and federate them, hence the name, into a single cluster. And so this has become, you know, we've been working on this for a while as well, but it's become much more relevant now with the AWS release because now you may have an on-prem cluster and a cloud deployment or you may have more than one cloud deployment. And if you want to, you can take all those different um clusters and you can present them now to users as if they were a single cluster. And so this is also huge. We've also gotten all kinds of people who've been interested in Federate because they're talking about, you know, trying to administer an on-prem cluster and trying to burst to the cloud.

Well, guess what? We're releasing that today. That's also, you know, it's in tech preview. And so, you can also contact us and um start to use Federate as well. Wow. Because all of this time, these last, you know, couple of years that we've been so excited about fuzzball and it's been growing, that's the piece that people have been talking about the most. Yeah. Is Federate. And so what you're saying is that they can have So how does it matter where in the world their cluster is? Is it kind of still a little bit local or could it be anywhere that they are kind of using the resources from all of them in one place?

Yeah, I mean it can be they can be distributed anywhere and yeah that's that's that's the idea is that you've got um you know maybe you've got more than one cluster um you know in two different sites that's both running fuzzball maybe you've got one cluster on prem and another one running in AWS um maybe you've got more than one account or more than one region and so yeah that's the idea is that you can take all those different clusters and and it's pretty cool because when you use um federate it looks like you're just using a single fuzzball cluster and so you can do things to send your jobs to to one cluster versus another.

Um but you know the the interface essentially just looks like a single cluster. So it's very very simple and you know once you know fuzzball it's very intuitive. I'm thinking of 15 people right now that I'm going to call and say that this is ready. So this is very exciting So, very excited about that. And, um, okay, I'm sorry, that was number two. What was number three?

Third big announcement is we've also been working on a little project called the workflow catalog. And so, what this is is essentially an entire catalog of um general kind of workflow templates that have already been created using different applications. So now let's say you want to start up a Jupyter notebook example, but you don't know how. And I'll show this in a few minutes, but just as a broad overview, you go to the workflow catalog. There's a template in place. Um it has some fill-in- thelank uh you know, kind of, you know, pre-populated um values that you can change if you need to. And then you just spin up that workflow and go ahead and you've got yourself Jupyter notebook running and you can connect and do what you need to do or you know one of a lot of other applications as well.

So yeah, so now it becomes really really easy to get started with Fuzzball because we've got this catalog of precreated workflows that's there and is going to be growing as we continue to uh create more. This is really exciting. It really brings up like a million questions. Um, but I don't know if we want to like dive into those particular details, you know, like, well, then exactly how does it work and what does it run on and what do you need in your system? No, Rose, let's just get excited about what it is. So, I do have one question that is a little bit relevant because we're actually going to Google Next in a couple of weeks, like you know, CIQ, a bunch of us.

Um, and you did mention AWS and that is very exciting. We're going to just shout that from the rooftops right now and roll with it. But of course there's going to be the question, what about the other clouds? So I know it's a little bit in process. You is there anything that you can speak to about that? Yeah, I mean basically all I can say is yes, we're definitely interested in running on Google. We're definitely uh interested in running on Azure. Um the engineers have done a really good job of trying to um create processes that'll be as generic as possible and be as portable as possible so that we can um you know get fuzzball running on these different cloud providers as quickly as possible.

But yeah, the the answer is that we're not there yet. It's we're in process. Um and that's that's something that we're definitely excited about and working on. Yes. Yes. And that is often a a good enough answer I find in the tech world, right? like as long as it's on the road map like you know it's going to happen, right? Because there's just it never ends. There's there's always more. There's so much that you can continue to add. So getting excited about something and then knowing that it's just going to continue to to iterate and iterate and iterate and h very exciting. Okay, so those are the three big main announcements.

Thank you Dave again and team who's working on that. Are we ready to see it? Is that Yeah, absolutely. So, um, I've prepared some some kinds of things to sort of highlight some of these announcements and talk. This is a this is a kind of a, you know, it's a it's a big exciting webinar because there's a lot of stuff going on. And so, I had a hard time trying to like, you know, usually I kind of try to hone in on one particular thing and talk about one particular thing. Um, and I do that because it's it's easier to digest that way. This is not that kind of webinar.

This is a little bit more of a highlevel everything in the kitchen sink kind of webinar. Um, I I want to encourage folks, we did a webinar back in January also that was a bit more of the what is fuzzball and let's do a little introductory kind of like um step-by-step introduction and I'll talk about some of that stuff as well in this webinar, but I'm not going to go into as details. So, like in the comments, uh hopefully we can drop a a link to the previous webinar and we can make sure that folks have access to that so you can watch that, you know, to to kind of get up to speed if I if I don't go into great enough detail with this.

Um, but yeah, I do want to just kind of like give a really broad overview of what the look and feel of Fuzzball and then talk about some of the specific things um that that you know we're we're announcing today as well. Let me go ahead and share my screen. And just full disclosure, um, you know, we we've done a lot of webinars, but in the past we tend to use Streamyard. Um, we're kind of we're kind of, you know, we're not you new to Zoom, but we are a little bit this team new to uh Zoom for webinars. So, if something doesn't work right, as as far as me sharing my screen or if everybody gets booted out of the Zoom or something like that, you couldn't tell, could you?

Sorry about that. Like, but we're we're we're uh we're getting better. We are. We are. Um, there was a question here. I think that we answered it. Um, so Erin, let me know if we didn't answer it previously, but I'll just read it for you real quick, Dave, because I know you kind of got like a lot of screens open. He says, "Can you give some context on what is available today with Fuzzball and Fuzzball Federate interested in onrem and other cloud providers like Vulture?" I think you sort of answered that mostly, but if there's any more you want to add to that. Yeah, so um, right now, uh, like I said, we have AWS and tech preview.

We also have federate and tech preview. So you can you can technically use those things yourself but um you'd have to talk to us first to get access to those things for the time being. Uh and that's that's essentially just because we are anticipating a lot of interest in this and we want to make sure um that we don't just open it up to everybody and suddenly have questions and comments and things from a zillion different people and we're overwhelmed, right? So we want to make sure to kind of um onboard people responsibly uh essentially. So if you are interested in using those definitely reach out to CIQ.

Um Vulture is kind of an interesting cloud provider because we use Vulture internally quite a bit uh to spin up um so if you're not familiar with Vulture, Vulture is is a cloud provider that gives you access to uh bare metal. Um, so you can have clusters that that are not VM or you can have um cloud resources that are not VMs, but they're more bare metal and that allows you to do a lot of interesting things. But we use Vulture a lot internally just to um uh set up uh kind of like uh onrem simulated clusters and to use that. So yeah, um I would be interested in talking to you about that as well.

I'm sure we could help you set up um Fuzzball and Vulture if that's something you're interested in doing. Um, yeah. Does that answer the question? I think so. Yeah, it was in the little Q&A, Erin. So, if if it does not, um, it does. Okay. Awesome. Thank you. Cool. Thanks, Dave. All right. So, I'm going to do just a really quick demo like look and feel of Fuzzball. I've brought up like the main screen that you'll see after you log in. I'm going to go ahead and select an account. Um, you can either select your own personal account or a shared account. I'm going to select a shared account in this case.

Um so and I should I should also highlight as I'm going through this I usually try to highlight that all of this is API driven and so what that means is that um the web UI is hitting the API we also have a command line interface and that all hits the uh API so you can do all this everything that I'm showing you through the command line also if you want and we also have um software development kits SDKs so that you can you know do all this in Python or something if you Okay. So, I'll just kind of go through and I'm going to highlight some of my own workflows.

So, one of the cool things that you can do in this um workflow history view is you can see everybody who's a part of this shared um account. You can see their workflows. And so, I can jump in and I can see, oh, well, Brian's doing something right now. And I can look and see what he's doing. I can't actually get logs from the workflow because those are for Brian, but you know, I could see what the definition is and I could open it in the workflow editor and I could, you know, use this as a basis for my own workflow. So, it's very, very collaborative and I'm going to be talking about how collaborative that is, uh, you know, in a few minutes as I go through some other examples, but I'm going to go ahead and just kind of show you.

Um, here's a very simple workflow that I created a little while ago. Um, it's just a couple of jobs. number job number one and job number two. And so I can open this up in the workflow editor and I can get this visualization of what what these look like. Right? So this is the first job and there's a dependency here for the second job. And if I click on one of these, um, I can open it up and I can see things like what's the command that runs, what's the environment that this runs in, how many resources does each of this job does do each of these jobs take to run, and um, I can also open a YAML file.

So in the background, this this, you know, graphical user interface is actually building a YAML file. And I can go in and I can see the YAML file that's being built and I could edit that here, you know, if I wanted to um do something like that, for instance. So, you know, this is just kind of a quick look and feel um very very basic of what it looks like to create a workflow. If I wanted to run that workflow, I can just do like this. So, I'll just say new test job. Go ahead and start that. There we go. And I should I should tell everybody this is actually running on AWS.

So this is this is um the CIQ uh you know it's one of our main clusters that we admin and we use and it's been running on AWS for a while. So, this is the same kind of cluster that you can now spin up in your own personal account and you can use or you can onboard your own users and have them use and things like that. Um, cool. So, that's kind of like the uh the workflow history and what it looks like to create a workflow. I showed you the workflow editor and so this is how it it looks to monitor a workflow. Now, one of the new things that we're talking about today is this workflow catalog.

And so this is one of the new things that we've been working on. And like I said, um you know, you can go into this and you can say, "Oh, well, I want to run a job that uses parview or I want to run a job that uses lamps or blast or stable diffusion or any one of these." And so if you're interested in in one of these, you can um view the details. And now you see that there's kind of a fill-in-theblank sort of thing here where you can say okay well um you know I there's there's some text that I want to fill in.

I want to change um the the u the URL to the container for instance um or I want to change how many cores we're going to allocate to this thing to run. You can look over here and you can see the actual uh fuzz file, but you can also just copy the application yourself and edit it. Oh, sorry about that. I might I might not uh I might not be properly authenticated here. U could be an issue with that. Anyway, um if I go back to the workflow catalog, what if I wanted to add my own thing in here? Can I add to the catalog or is that like a administration thing?

How does that get in there? Yeah, I might not be properly authenticated right now, but yeah. Um, normally you can create your own application. I've got a lot of different clusters open right now. So, there might be some like some issues with uh, you know, what I'm trying to do because like I said, I'm trying to do an everything and the kitchen sync demo here. So, some things when you've got so many different things open at the same time, some things are are bound to uh, get a little messed up. But yeah, you would just go ahead and create your own application. Um, that's definitely available to you.

You can do that as well. And normally what you would do is you would copy. So I I did actually copy this. It just gave me a little little silly thing when I tried to do that. So when you copy it, then you can go ahead and change these inputs or you can go ahead and open this up in the workflow editor and you can just um, you know, run it from there as well like like I like I showed you before that graphical user interface you can use. So that's the application catalog just in a nutshell. Um you can see that we've got maybe you know a dozen or so applications so far that's going to be growing.

So we're going to be continuing to add more to this catalog. We're going to be going going to be continuing to make this a a better user experience and we're going to be continuing to build this library out so that you can basically get started, you know, as fast as you want. Cool. Um, all right. So, let's see. Um, so I wanted to talk a little bit about, you know, what it looks like. Uh, okay. So, once again, I'm I'm I'm kind of going high level and I'm skimming a lot of this stuff pretty quickly. Um so one of the things that we use um a lot in work in in Fuzzball workflows is we use um volumes.

And so I've got a workflow here that um I'm going to be talking about in a bit more detail here in a second. But it's got a a volume that we're creating. And this is a a persistent volume meaning that I can save data here and then I can you know see that data again in the future if I if I run this again. Um and another thing that we use obviously anytime you run any kind of a a workflow you spin up resources you spin up um actual compute resources. So um one of the things that we've been able to do so you have to configure those things right you have to configure what kinds of volumes you and your users want to have access to and also what kind of resources you and your users have access to.

And um I wanted to talk a little bit about how easy it is to do that in AWS. Now, we basically have created um a couple of workflows for you that you basically just press a button um after you start your uh after you start your cluster and you say um I just want the default uh AMIs that that I have access to and it it automatically provisions for you um two types of compute resources, one without a GPU and one with that you can use with your cluster. And then it'll als it'll also automatically provision for you um two different storage classes, one that's persistent and one that's ephemeral.

So that workflow makes it really really easy um to spin your cluster up and also to get it configured right off the bat automatically just by pressing a couple of buttons. So you know that's how that's that's you know one of the ways in which you can manage um you know very easily uh these resources in the cloud. Um so I wanted to talk a little bit about that and then I wanted to talk a little bit more too about a realworld use case that we're working on right now. So, we've been talking to some folks who, you know, thought it would be really cool for us to try to see if, um, a software suite called FDS would run within Fuzzball.

So, FDS is fire dynamics simulator. And so, um, I talked to one of my colleagues here at, uh, CIQ, um, Adore Soena, and very, very quickly he was able to, he had never heard of FDS. I've never heard of FDS before. Very, very quickly, he was able to go to the GitHub repo, um, get a container and create, um, this workflow that runs this simulation. And essentially what this simulation does is it simulates how fire burns and how smoke emanates from a fire. So you can do things like figure out how a fire is going to spread through a room, how it's going to spread through a house, how it's going to spread through a forest, um what the impact of the smoke's going to be, all that kind of stuff.

And so in a very very short period of time, he was able to go from uh and shout out for doing this great work. um he was able to go from not even knowing anything about um this software to actually getting a workflow up and running that'll do a simulation. And so I was going through and I was thinking about what would be a good thing to demo because you know a lot of the things we we've demoed lots of different things with Fuzzball over many of our webinars and so we're always looking for new stuff and I was like well I know that

is working on this maybe I should I should go and I should see and this highlights once again um how collaborative fuzzball makes it throughout your team because I basically just went and looked on the shared account that he was using and I saw the workflow that he had run. I grabbed it and I ran it myself and I was able to save the output almost immediately to a persistent um volume that I had. And so this is the actual workflow that I ran based on what he put together. And then I was able to run he's got another workflow that I'll show in just a minute that um actually shows a visualization of this uh this couch burning simulation.

And I was able to get all that up and running based on the work that he had done um without even contacting him and I did it in an afternoon. And then I basically sent him a Zoom or sent him a Slack message and was like, "Hey, would would you be okay if I highlighted some of the great work that you've been doing recently?" And he was like, "Yeah, absolutely." So like that is like the definition of collaboration, right? Somebody else in your team is is putting in the work to to get this done. They're able to do it very quickly because fuzzball makes, you know, makes it very very easy to get things up and running.

And then you can just kind of go through and look at what other members of your team are doing and you can grab those things and use them yourself. I don't even know what to say. Thank you. Yay. Harrah. Like golf clap. Um it's and with these sorts of things, time and energy and resources and collaboration are lifesavers, right? So yes, that's really awesome. Thank you. Yeah, and I'll just kind of talk about this workflow just a little bit really quickly, too. So the the workflow um you know, it starts off with this this uh program and a container FD FDs and it also has some configuration files that we pull in.

Um and then there's, you know, some commands here that we run. I could go ahead and I could actually show you. Okay, so a little bit of a random question. Is there a way to share this with someone who does not have fuzzball or someone who's not in your group? Absolutely. So, I could always um let me see if I can open this in a new tab so I don't lose what I've currently got here. But yeah, you could you could always open this in the workflow editor and um you could always just uh download the YAML file. So, I showed you the YAML file here previously.

This is the YAML file. And essentially, you can just download this YAML. Um there might be a few things. So uh if they have used the pushb button workflow that I was talking about previously um then you know they should have access to p both persistent and uh uh ephemeral volumes. So the volume names might be a little bit different. they might have to change um a few things around but it should be very apparent you know once they've used fuzzwall for a very brief period of time it should be very apparent what things need to be changed you know sometimes there's secrets that we

used in order to access things like S3 buckets or um you know in order if if we need to access uh images that are secret images um you know not like like in private uh OCI registries those secrets are not exposed in these YAML files, but we reference them and they're saved in a secret store on um on the uh right right here within the secret store in Fuzzball. Um so you might have to update those. Um but you know, that kind of stuff is is is pretty readily apparent. Um it's pretty easy to do. So yeah, you can totally share this across like with somebody who does doesn't even have fuzzball.

You could send them the the file and you know, they could look at it and then they could get fuzzball and run it. That's that's really cool. I I love this word collaboration that you have said a few times and so it's yeah exciting. Thank you. So yeah, just to look at the logs again really quick also. So we can see that this is actually an MPI process that's running. We've got four different MPI processes running together. Um, one thing I wanted to really highlight and I kind of forgot to do uh at the beginning of this talk is um, you know, we created Fuzzball and we talk a lot about high performance computing, HPC and you know, Rose, you and I were kind of discussing this a little bit before the before the meeting as well.

Um, I think that when a lot of people think about HPC, what they think about is like, well, somebody at a Department of Energy, you know, national lab doing, you know, climate modeling or weather modeling or maybe somebody at a pharmaceutical company or also like, you know, something like that, like doing uh drug discovery, things like that. And it's usually thought of as very academic and you know very kind of like sciency. But a lot of the stuff that we have shown um over time with fuzzball is things like AI uh training and inference happening. Um you know people can use fuzzball for like in the financial sector for Monte Carlo simulations to kind of figure out what kinds of investments um they should be making in the future.

Um, people can use fuzzball like the automotive industry. We've worked a lot with uh folks in the automotive industry showing how to do simulations of crashing cars together for safety and and for performance and things like that. Um basically there's lots and lots of people in different industries who wouldn't um who wouldn't describe what they do as high performance computing but they really are doing that kind of you know they're using you know lots of different computers in the cloud all together to do one thing or they're sharing some resources with their colleagues that you know some data center where they can run um you

know simulations or where they can uh run models and things like that and these folks um we we've kind of started to talk about performance inensive computing as a a catch-all term because there's not really a term yet to kind of uh you know bring all these different communities together and talk about them all as one. So that's that's kind of what we'd like to talk about here with with you know with this with fuzzball is per is targeting performance intensive computing and this is you know this is a good I think this is a good um example of that.

I don't think that a lot of the people who use this this uh uh FDS program a lot of them probably don't think of themselves as HPC folks right this is a different this is a little bit of a different industry um but yeah so this goes through the simulation gets here to the bottom and it it saves some output files. And so, let me see. I don't know if my connection is still Yeah, here we go. So, here um I should show you too. I have another I have another job running and this job is still running and this is the FDS visualization job. And so, I just went over to this little VNC instance.

I'll I'll go back to that in a minute, but that's being backed by this job right here that's running within this workflow. And so essentially what what um did and then what I you know grabbed from him and did myself as well is um we ran this simulation and we did so with this persistent volume mounted and then once the persistent volume saved the data, we were able to run another workflow that used that same persistent volume and now it's got access to the output from that simulation. And in this one, we're running um you know, a pre-baked uh container that has within it a virtual desktop and it also has something called smoke view.

So smoke view is what's used by FDS in order to visualize uh the simulation that was run previously. So let me go back over. Um, so basically this this is running within this job up in the cloud in AWS and it's giving me this little desktop and I've got this terminal that I've opened up and I've got access to the data that was analyzed in the previous job. And so I can go ahead and and run that smoke view program. Oops I do. And once and once I do that um there we go. Once I do that, I can see this little couch, and that's pretty cool.

But what's really cool, let me go ahead and add to it um some of the simulation. So, that's smoke coming out of the C. Oops. Is that smoke um emanating from the couch? Let me go ahead. What is this? Is this someone dropped a cigarette or something? Yeah. So, it's something like that. So, basically, this is one of the um And I can rotate this around. We can look at the back of it if we want. This is one of the like proof of concept uh one of the suggested proof of concept workflows uh for for making sure that FDS is actually running appropriately.

So it's a pretty simple little simulation of like a foam a foam couch and as this fire burns away you can see the couch starts to deteriorate and you can see where in the room the smoke would end up and how the fire is going to spread and you know all this kind of stuff. And this is this is useful if you want to, you know, figure out exactly what you just said. Somebody drops a cigarette on the couch. And so, um, how does the how does the fire spread on this this particular type of couch that's made out of this particular material?

Um, I assume that you could model the entire room if you wanted to, and you could start to see how the fire is going to spread through the room depending on what kind of furnishings are near the couch, what kinds of materials those are made out of. I mean, you can already see too, like how does the smoke distribute in the room? I mean, people tell you if there's a fire in the room, one of the things you're supposed to do is drop to the floor, um, you know, and get under the smoke so you can still breathe and then get out of the room. Well, you can see, I mean, here's a simulation of that, right?

You can see exactly why why they say that. So, yeah, pretty cool. And once again uh because because of fuzzball, my colleague was able within a very very brief period of time to go from not really knowing anything about the software, how to run the simulation, how to visualize it, anything like that to being able to produce this and then I was able to go and say, "Oh, cool. What am I going to present tomorrow? Ah, a colleague is working on this cool thing. Maybe I'll grab that and show everybody, you know, what what he's been working on." And I was able to do that in an afternoon, you know.

So, but you said this was in the S3 bucket. Is that where it gets stored once it's like done running or does it does it do that automatically or can you tell it where to go? That's a great question. So, um so what's actually happening here is uh so this is not in the S3 bucket. So we we we do stage data ingress and data egress a lot of times using an S3 bucket to um persistent or ephemeral volumes within a workflow. But here what I have is a persistent workflow or a persistent volume rather. And this p and our volumes can be backed by lots of different storage types.

In this case we're backing these storage types on EFS. And so um kind of a long answer to your question, but that there's a storage class in Fuzzball using EFS, which is a persistent storage type, and I'm able to, you know, save um from one workflow into that uh EFS volume, and then use it in another workflow if I want. The other thing we do a lot is um ephemeral um storage volumes. And you've probably seen, you know, me demo those in the past. And what those do is they just spin up. They're also backed well they can be backed by EFS. They can be backed by almost anything really.

They can be backed by you know NFS or you know local other local storage that you've got or you know other types of things. But in any case they you spin those up on demand and then we have um Fuzzball has a method of doing data ingress and egress. You can pull the data in that you need. You can save the data. you can, you know, share it between um different uh different jobs within your workflow and then you can push the data back out. One thing I wanted to show you too, I kind of forgot to show you. I showed you this really simple workflow.

I'm kind of jumping around. I'm sorry. Um I showed you this really simple workflow, but I I don't want to give you the idea that your workflows have to be simple. So, here we've got a more complicated workflow. I like to show this one um because it shows you how you can kind of like fan out your dependencies and then fan them back in if you want to. Um but this one is another one that uses if I go in here and I look at the volumes, we can see that I've got this volume created. And if I go in and I look at it, we actually have this data egress.

Um so what this is going to do is it's going to create an ephemeral volume. And that means that this job, these jobs, this job, they can all share the same data backed by the same volume. But that volume is going to disappear when the workflow is over. And so I can define this data egress here that says find this file on this volume called blast blastp.out and make sure to save it here to my S3 bucket before you end the workflow. And so, you know, this is how this is another way that we kind of look at um using volumes is by just doing data like making ephemeral ones and just doing data ingress and data egress um in order to you know get the data onto or off of an S3 bucket, a GitHub repo, an FTP site.

Basically, if it's got a URL, you can do it. Awesome. Cool. And then one last thing. So I've kind of talked a little bit about Fuzzball running in AWS, which basically all the stuff that I just showed you is running in AWS and I kind of showed you a little bit about the workflow catalog as well and how that works. And I also like to talk a little bit about Federate. Um, so I'm going to go ahead and stop sharing because I've got to like log out of my current cluster and log in and do a few things like that. But once I get that set up, I'm going to show you just a little bit of a demonstration um of of Federate.

Now would be a good time if we have any questions. Um I don't see any new ones that have come in, but um there still are many people on here, so thank you guys for staying and hanging out with us, even though I had a little blip earlier. Um, but that's wonderful. Yeah. So, if any questions have come up, I mean, I I imagine your brain is like mine. Is like, can I do this? Can I do that? U, so go ahead and pop them in there and, uh, we'll get to them either as we go or maybe at the end if we have a little time for Q&A.

I know you could probably like keep talking about this for several hours on Dave. Yeah, I'm just the only the only reason I'm taking a break to breathe right now is because I'm logging into a Federate cluster because I'd like to give you guys a Federate demo. I don't know what is going on with my camera. Sorry, that is like bizarre. I just don't move. Um, okay. So, there actually was a question that I had and it's gone. Okay. Well, it'll come back to you here in a minute. So, I'm gonna I'm gonna go Let me go just clear this. I'm going to go ahead and share my screen again.

Okay. A terminal So all right. So what I've done is I've logged into so I've logged into a new cluster here and um so this is it looks exactly the same, right? This doesn't look any different from the clusters I showed you before, but this is actually federate. So it's pretty cool because it doesn't look any different, but um it's actually two different clusters that have been kind of mashed together into one cluster that I can use at the same time. And you can see that because you can see there's these different workflows that I've run and some of them say Fuzzball onrem stable and some of them say Fuzzball AWS stable.

So these are different clusters. One of them is in the cloud and one of them is on prem. I think it's actually in vulture is where this actually is running. Um and so here in my terminal so I've I've logged in um to this particular cluster on my terminal as well. And this will give you a little bit of an idea of what the what the CLI looks like as well. So, I can do, you know, fuzzball stuff. Um, but I'm going to show you. I've got a couple of files here. Let's look at this submit.sh. And I have a fuzz file here. So, I've got a hello world fuzz file.

So, I can submit this fuzz file as a job. So, let me go ahead and look at this um submission. I don't know if you can see any of that. Let me Yeah, I mean, maybe a little bit a little bigger would be good. I'm gonna make the background. Oops. All right. Is that cool? That's cool. All right. So, what I'm doing is I'm running a command. Fuzzball workflow start Hello World.fz. Let me just run that command really quick. And let's see what happens when I run it. It's going to start a workflow. And if I go over here to my cluster, refresh. Cool. I just started this workflow.

There's the workflow I just started. It's done. It said hello world. Awesome. So, what does this script do? Well, it runs it starts well, let me just you guys can see. Um, it starts this Hello World workflow that I just started, but it does it in a loop and it does it 10 times. And every time it does, it echoes, hey, I'm running, you know, this iteration. That's really all it does. And then at the end, it's going to run this other command to list the last 10 workflows that I ran. So we can just see, you know, what it looks like and where they ran.

So let me go ahead and run that script. That's going to take just a few seconds because it's got to go through and start 10 different workflows in a row. And so what you'll notice is that now I got all these workflows running, right? And some of these are running on AWS and some of these are running on prem. So the way that we've got this cluster uh configured is that it just sort of randomly farms the jobs out to one or the other to try to do really really basic load balancing, right? It just sort of says, "Okay, well, I don't know. I'm going to send you over here and you over there." Now, of course, you can put in other rules um you know, to make things more intelligent.

You could say, okay, well, maybe if the cluster has this many jobs already in the queue, go ahead and and burst up to the cloud instead of having to wait in the queue for a long time. Or maybe if the, you know, the job is a really big one, you know, go ahead and push it up to the cloud instead of running it locally. Or maybe if I've have asked for access to some volumes which are only accessible um on one particular uh in one particular environment, go ahead and run those there.

So those are the kinds of things that right right so those are the kinds of things you can do and so this this is the final command that I told you about that runs in that script and it basically is just telling me where did all these run some of them ran on AWS some of them ran on prem and so this is just like you know um just to kind of give you guys a flavor of you know what it looks like to use Federate Federate to me it almost doesn't demo very Well, and what I mean by that is that it's like it's so simple.

It's just the same as using any other fuzzball cluster. Same same. Yeah. But that's like that's the thing. That's what's so awesome about it is once your users learn how to use a fuzzball cluster, they don't need to do anything different to use Federate, right? They just go ahead and submit it to a different URL and then boom, now they're they're running on multiple different clusters and they they might not even know it. So money sometimes comes into play when people are doing uh cloud stuff. Um so one of the things that is awesome about uh fuzzball in the cloud as I understand it and correct me if I'm wrong but as soon as the job is done it there's a word for it brings it down right like it doesn't keep that instance open in the cloud and you're still paying for it right.

Yeah that's a great thing to bring up and I'm glad that you brought that up. Yeah, you you don't want to um you know, like sometimes when people think about HPC in the cloud, they think that you know, one of the first things that springs to mind is let me just spin up a big slurm cluster and have that sitting there spinning away and running and then I'll submit my job to it. That's a really wasteful thing to do because you got all these resources now that you've spun up and they're just sitting there and you're paying for them whether you use them whether you use them or not.

And so that's exactly the kind of thing that fuzzball in the cloud uh fixes for you is that you know the Fuzzball orchestrate platform which is what I have showed you consistently you know throughout this demo. That's something that runs on one server um and it runs in the cloud and it's kind of running all the time but it doesn't use very many resources. It's very you know it's it's only like six CPUs and 16 gigs of memory is like the minimum or something like that. So that's that's very very minimal and that's all that's your whole cluster. That's the only thing that needs to run.

And so then that's the thing that accepts those fuzz files and says, "Oh, okay. You want to run this workflow that requires all these resources. I know I need to spin up these cloud resources, land your containers on them, execute your commands inside them, and then as soon as they're done executing these commands, I need to tear these resources back down so you're not, you know, wasting all this money with these resources running." And that's exactly what Fuzzball does. Awesome. We've got um a a question here. Uh Wle says, "Cool demo." Thank you, Dave. Using Federate, is it possible to lift and shift in quotes along running jobs between clusters?

For example, if a job is using checkpoints. Yeah, that's that's a that's a great question. Um so yes, I mean the the the the short answer is yes, you could do something like that. The long answer is however it would take a little bit of um engineering at at the present moment. It would take a little bit of engineering on on behalf of the user themselves. So yeah, as you said, you'd have to be using checkpoints. Um so you'd have to already have that set up and you'd have to be, you know, saving. Um we've talked a little bit. I mean, this is, you know, this is something that's still in like just kind of the talking about and wouldn't it be cool sort of stage, but wouldn't it be cool if we were able to, you know, more automatically checkpoint jobs.

We got some ideas floating around, but this is this is very far off. This is very much in the ether right now. But yeah, if you were checkpointing your jobs yourself and you already had that built into your workflow and you already had, you know, things saved, then yeah, there there um there could be technical hurdles that you could overcome based on resource availability. For instance, like if you had an on-prem cluster that had GPUs and your cloud um environment uh didn't have access to the same type of GPUs, that's just not going to work, right? Um, so yeah, there's some things to to bear in mind, but in theory, you know, those sorts of things would work as well.

Um, so do we have more on your side of the demo? Um, that's pretty much everything that I wanted to show. Okay, this is very exciting. I do have one more question if it's okay. Um, have a couple of questions coming in over here. Evan, thank you so much. He says, "Can a workflow work across the different clusters? like do job one in one place then move to another cluster for the second job. Um as of now no. So a workflow runs uh completely you know from the start to end within one cluster. So you can't split jobs up and have you know one job within one cluster and another job within another cluster.

Once again I I've heard some talk about you know we we've kind of discussed things like that a little bit. um you know there that kind of falls along the same wait a minute I'm trying to understand this in my brain okay so there's the workflow which is all the jobs yes right so there's multiple jobs in there perhaps dependent upon each other so he's asking if you can because you need whatever a TPU for this particular part of the job or this particular job this particular part of the workflow as a job so you go do that somewhere else and okay so sorry my brain was trying to understand what he was asking.

So you're saying at this point, no, the workflow altogether runs in one place or another. Yes. Yeah. Visual visual aid because there, you know, if you if you've got that question, Rose, I'm sure a lot of other people do too. Yeah. So like, um, jobs are like the atomic, you know, parts of a workflow. And so jobs are what you build your workflow out of. Um, you can have one or you can have a zillion jobs. And you know as of right now um this entire workflow if I submitted this to federate is going to run either on prem like to be you know for a concrete example the federate that I showed you it's going to run either on prem or it's going to run in the cloud.

This entire workflow it's not going to say oh well this job is best suited to run on prem but these two jobs are best suited to run in the cloud. No that it's not going to do that. And you know this job also kind of highlights one of the problems that would be associated with trying to do that because one of the things that we're doing here is we're setting up volumes right and we have this ephemeral volume that gets set up and it gets set up within the infrastructure here and then it gets mounted into each individual job and these jobs share the data between um between each other in in order to to uh you know process the data.

So it it would be kind of problematic to like create all the directories uh on prem and then actually try to download the data, which is what these two jobs are doing, and not have the directories created because the volume's a different volume, right? So So there's like there's things like that that you'd have to kind of think through and try to figure out. I mean, you you could probably do this with like persistent volumes if you had access to the same storage classes, you know, backed by the same um volumes in two different environments. But that you're getting into some kind of, you know, difficult things to do at that point in time.

I really like that question though. Yeah, that one as time goes by might become more important. So, yeah, great. Thank you. Um, I think that's great. Okay, wle. Um, and following a follow-up question. Thank you. Is using federate or not a function of something hardcoded in the workflow file? No. Um, using federate or not is a function of just where you submit the workflow file, the fuzz file. So, um, yeah, you could either, you know, open up a web user interface and submit your fuzz file that way, or I showed you that you can submit your fuzz file, um, through the command through the command line.

Uh, the command line, it's going to it's going to submit that fuzz file, you know, wherever you are currently logged in in your fuzz in your fuzzball CLI. So, I I didn't show you the login steps, but you know, you go through a a series of steps in order to select which cluster you want to use. We call them contexts within the CLI, but you select which one you want to use, and then you run a command to log in. And so, whichever one you're logged into, that's where your fuzz file is going to get submitted to. Um, yeah, he followed up. So, it's a function of context, which Yeah.

Yes. Um, okay, great. So, yeah, if there's any other questions, definitely, um, pop them in there or reach out to us, right? We are definitely wanting to talk to you. We want you to be able to to download it and play with it yourself and see how it works out for you, especially if you um, you know, have the the AWS. And then if there's other clouds, still feel free to reach out, right? Because we want to make sure that we are in communication with you. And even if you were, you know, testing it out maybe on prem or just on our list to make sure that we reach out to you when other clouds are on the way.

Plus, and it also helps too because we're like trying to figure out how to prioritize which which cloud do we target next and you know I mean we we just went to a conference for instance um a Gartner conference in which almost everybody that we talked to was using Azure and so like if you know that that's that's pulling us in that direction. So if like if there's a bunch of people at Google next that are like no no no we we need to you know this needs to run on GCP like those votes count right those help us to decide you know what should

we uh what should we target next yes exactly your voice matters and we want to hear from you right because we are going to Google next so if you are going to Google next we would love to meet with you happy to do demos talk to you about stuff we want to know about your environment what's important to you all of these questions are are amazing I know that we Um, we got to wrap up. So, I am going to put my questions on a note card for the next time because I know we're going to continue to talk about fuzzballs. Time goes by. This is this is um this is our baby.

Maybe that's a weird way of saying it, but it's like what CIQ was was built to be, right? You know, so yes, it's very exciting that here we are and we are are presenting in this way. Okay. case. Any final words, Dave? Um, just that it's it's once again, um, it is a privilege to be able to demonstrate the work that so many other people put in. Uh, so, you know, thanks to the the Fuzzball engineering team for putting together a beautiful product that I'm able to go on and and show to everybody. And, you know, uh, just yeah, it's it's really awesome to be able to do this.

And it's very exciting that we got all these announcements today. It is. It is exciting to have lots of announcements. Um so again, yes, thank you Dave. Thanks for coming on and showing us. Um if you would like a private demonstration, you want to know more about Fuzzball, you want to get it, you want to be part of the tech preview, definitely um reach out. We are so so excited. Okay. Well, thank you Dave. Thank you everybody else. Thank you for being here, all these people. Have a wonderful day and we will see you next time.

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

9

Enterprise products

Spanning the kernel to the orchestrator

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert