Live at SC23: Demos and Tours webinar poster

Live at SC23: Demos and Tours

Watch Now

Join us live at SC23! We will be doing live demos, tours of other booths, and an interview with the winners (who used Rocky Linux) of the student cluster competition. Tune in, ask questions, and virtually join us at SC23!

You'll see these live talks:

  • Benefits of PBS Pro

  • Warewulf, Rocky Linux, and Ansible

  • Fuzzball Demo

  • Mountain Demo

  • PanFS Benefits and Connecting to Rocky Linux Machines

  • Sentieon DNAscope

Speakers

  • Scott Campbell, PBS Professional Product Manager, Altair: LinkedIn

  • Tannor Brown, Information Systems Architect, Sandia National Laboratories: LinkedIn

  • Forrest Burt, HPC Systems Engineer, CIQ: LinkedIn

  • Jonathon Anderson, Solutions Architect Manager, CIQ: LinkedIn

  • Richio Aikawa, Senior Director, Strategic Marketing, Panasas: LinkedIn

  • Brendan Gallagher, Head of Business Development, Sentieon: LinkedIn

  • Zane Hamilton, Vice President of Sales Engineering, CIQ: LinkedIn

Transcript

so Rocky and werewolf and and say let's make sure that all the changes that needed up the raspberry pies as of a few versions of firw can be put into network mode so we can actually in principle make werewolf work as well I haven't done it yet but I think all the pieces are and I've talked to several people in and they're excited so I think that all right we're here at Super Computing 2023 with the smallest cluster in the show floor at least the smallest cluster that's running Enterprise Linux it's right down here and it's a Raspberry Pi cluster that's running Rocky Linux on

two nodes running the latest Fedora 39 on the head node and it can actually accept Jetson Nanos or other types of CPUs the whole goal here is to take the Enterprise High School or used for software development or with exactly the same software so here we'll stop now and you need to get on to the all right so here we've got a touring p 2 version 2.4 motherboard it's got four Raspberry Pi cm4 modules and I've actually tacked on an extra ethernet here on the Node 4 just so I can have a dual ported head node like I have on a real cluster eventually I

I will actually build a separate head node out of a little $25 router card so the cost to get into this stuff is extremely low but it's running Fedora 39 as of 3 weeks ago it boots natively on Raspberry pies so that means that I have all the runway needed to build the rest of the software stack I've already talked to several of the werewolf developers we think we can make werewolf work on that setup eventually of of course it will TR it it will go Downstream to to the rocky product line to any of the Enterprise Linux product lines so we'll be able to

run Enterprise Linux on this setup it's early days yet but I think we're really excited five Alan that's probably enough that was awesome that was awesome cool I think yeah this was a good thank you for letting us have all of that SP they taken that whole process [Music] [Music] typically en yeah yeah [Music] [Music] the way that would work for us there would be a way to identify [Music] buz that up that would and then they would decide it be it'll be something that's so first just see this job needs to have access to that while [Music] job I don't think it's supposed to

View full transcriptHide full transcript

be out basement [Music] I that's where you so yeah like me [Music] back [Music] the absolutely it runs as a separate job that is part of [Music] [Music] go [Music] fails forever Reon [Music] [Music] answer [Music] [Music] how important that to I don't know asking about seems a little disconect I think you just do like they [Music] do so they're not up Prett mostly [Music] comptible all the D so like perspective we want to have versus like create a new job that tries to inherit the state and so if we allow the job to fail we set it up in a Wayan what kind of

me [Music] want e e e e e e e e e e e e hello is it loud on the yeah I can hear but I you tell me would say a little good uh hi my name is Scott Campbell I'm the product manager for PBS professional been working with Al for the past 18 years on PBS and today I would like to talk about the advantages of open PBS and alter PBS professional once we will be integrated with ciqs uh fball product uh just as a quick introduction to alter for those who may not be familiar alter is a very large company uh that

we where simulation HPC and AI converge what we mean by that is Alter Flagship probably most well-known at least historically business unit is the hyperworks division which is simulation and design platforms for automotive uh consumer products Aerospace Etc doing finite element analysis crash simulation and so on um we have alter rapid minor which is a relatively new acquisition in the data analytics and AI platform space and then alter HPC Works which is our HPC and Cloud platform stuff that where the majority of what I'll be talking about today fits into that arm of the business uh Al before I go there though in the hyperworks

sort of uh product development side of things alter's recently released our Ros crash solver as open rodos to the community and we've been partners with uh ciq in in launching that ciq has been tremendous help there in containerizing uh open Ros including open Ros uh in Rocky linet Etc so there's other areas outside of HPC that alter and ciq have been collaborating in the past few years drilling down into the uh HPC Works side of things this is a not going to go through all of these but this is essentially the portfolio of products that we currently offer today I'm mostly going to be talking

about PBS professional as under the workload management side of things but over the past few years alter has acquired different companies in the workload management space with alter accelerator for the Eda space um the grid engine alter grid engine used to be univer grid engine Oracle grid engine Etc very strong in the Life Sciences space uh liquid scheduling is one one of our new solutions for connecting multiple clusters together in new and novel ways uh for unlimited and unparalleled scalability come check it out at our booth if you would like to know more about that um in addition we have a variety of products around

user accessibility for job submission admin control of HPC resources under PBS and a job workflow analytics we have Monitoring Solutions including alter Insight Pro which is a new new topic that I will talk about briefly at the end here as well as different um these are IO profiling tools and this is a very complex job dependency manager tool that comes out again from the Eda space so in addition to PBS Pro we have the community edition of PBS currently called open PBS um it includes a lot of the foundational technologies that have made PBS a pillar of the HPC industry for the past 20 odd

years including scalability up to you know tens of thousands of nodes and millions of cores robust policies for deciding on what the right work to run is and where to run it so that you know business goals are met fair share is met whatever sort of inter interde departmental agreements that have been exist at at an organization can be met it's very resilient fail over no single point of failure when jobs failes um you know they're handled in a sane way so that PBS doesn't lose track of them and they can be rerun and the same results can be captured uh on a rerun of

the job there's a very flexible flame framework plugin where uh framework plug-in framework excuse me where during critical uh events in the life cycle of a job um python hooks can be embedded into the system that are executed with you know good visibility into the job data structures control over some of the job attributes and so on same for nodes depending on the uh depending on the event in in question implemented via that framework we have health checks framework so you can make sure certain devices are or certain file systems are mounted before uh before greenlighting that node to start jobs you can make sure

that users exist libraries exist whatever it is you want to check uh and in the past we've been voted open PBS has been voted number one HPC software by HP by HPC wire readers uh had previous supercomputing so we've got open PBS and we've got PBS Pro and why might you want to choose one versus The Other Well it all comes down I'm not trying to start religious wars here between open source and commercial but you know there's in certain industries there's risk avoidance uh in a lot of commercial you know car companies for example or an aerospace company versus the sort of risk-takers and

early adopters that want to be on the Forefront of what you can do with an HPC management system um so the experimenters versus the you know 100% reliability and we think this the current PB PS Pro and open PBS dual licensing solution addresses that in a way that hasn't really been seen before in uh HPC so I'm going to talk a little bit more about alter PBS professional specifically and why you might and and what it offers Above and Beyond open PBS um so just an overview of the capabilities a lot of it is what I already talked about with open PBS because there is

a lot of fundamental crossover between the two um some of the differences we get to in capability are the ability to simulate which I can talk to in a few minutes about simulating changes into your Scher configuration or job workload or uh node configuration or even node quantity to simulate with your actual job structure job data what the schedule would do differently if it had a different situation additionally we have an allocation management system there's some Cloud bursting capabilities um so with PBS professional you get regular patches we have additional security features in PBS Pro Enterprise security features that are not included with open PBS

including TLs encryption of all network communication and full SE Linux MLS support which enables government agencies to run multiple different classifications of of of jobs on a shared hardware resource uh we have container support for Docker and Singularity with with PBS professional podman and Apper are work as well if you've done an alias uh we're going to uh revisit that soon but that's where things stand uh we have a high throughput scheduling module that can increase the throughput tens of thousands hundreds of thousands of very small very short jobs if you have a bursty workload like that there's a high throughput Computing module that works

directly with PBS professional taken from one of the other uh the Eda scheduling domain that I had spoken about earlier uh Bud budget management outside of the core workload manager you can allocate budgets of different resources so storage compute CPU time Etc acrew that over time set budgets and goals and limits for the uh for the year for the quarter whatever the charge period is charge different amounts based on how and where the job runs um and then do all postpay budgeting or postpay model so that it doesn't do any enforcement while jobs are being scheduled but you can do accurate reporting and charging back

after the fact or you can have all these policies be enforced so that jobs will not be allowed into the system or will not run when they get into the system in addition to the normal complement of PBS Pro uh q and server limits simulation I'll talk about that more in depth in a second Cloud bursting um PPS includes a cloud burst module that can burst based on workload suitability resource requests Etc into all the major clouds so AWS or oci Google and Azure um PBS improves provides robust GPU integration uh so TurnKey configuration for device isolation for NVIDIA gpus topology aware device detection and

scheduling so that you get gpus that are associated with the cores that are that are that are closest to them setting up the environment and collecting GPU usage information from things like dcgm for example uh again not to go into too much uh open source versus commercial product uh Rial rousing here but you know a lot of times you're Shifting the cost from upfront to to later when you have a support and alter has has been known for years for our excellent support teams based all over the world so you can get support in your local language and again increased quality assurance so more Pat

more regular patches improved uptime fewer unhappy users coming up in our next release for version 20241 this is coming out in January we're introducing a graphql interface to cover common uh job and um query job submission job alteration and node and job query uh task so essentially covering the Q sub qstat qel and Q alter and PBS nodes commands um this is great this provides another Avenue for things like fuzzball where we can integrate other Technologies seamlessly into PBS using a web uh a modern API we have the CLI and so on so the admin can or the application can generate tokens that have you

know expiration dates and so on and do all the normal token management things for web API Act ACC back to workload simulation this is a feature that's unique to PBS Pro and allows admins to again take their current jobs in the queue and change the the change the uh scheduling policies that are in place change the amount of nodes that are available to the scheduler and sort of just vet out any potential change that you want to make to your scheduling policy or your cluster and make sure that it's going to have a beneficial impact on your workload using a different alter product in our

HPC um portfolio called alter control you can actually also do those same sort of simulations going back months and months into uh your past history your counting log history so you can run with PBS Pro you get the from now forward sort of simulation and or look ahead simulation and with uh alair control you can also do back look back simulation and again no no impact on productivity it runs much faster than you know it's a simulated reality so you can run your run your simulations much faster than it would take to experiment possibly have a negative impact on your cluster to find out you

shouldn't have done that uh coming next we have a alter Insight Pro this was just announced yesterday it's our latest generation of uh um monitoring and Reporting solution for your cluster so it can take in take data from jobs from system metrics Cloud Telemetry Etc and different pro products from the AL portfolio including these IO profiling tools that I had been talking about um and and integrate that and synthesize all the data and answer common questions for admins that want to know about who used what and when what did they leave unused uh was there was was the system being used efficiently how long How

would how what did job weight times look like relative relative to their execution time and so on and I have a few example charts here showing cluster usage by Q so what was the unused you know number of CPU hours GPU so that some of these Advanced scheduling capabilities and Reporting tools and so on can be available with fuzzball workflow generation and workflow management on top of it there's other solvers are as well we we've been really working closely with ciq and it's been great over the last few years so I can take any questions or anything else but thank you thanks for having me

in your booth and thanks for your time yeah go ahead please yeah okay so the question is about how does the cloud bursting work um so actually in this case this will actually use because we are using the actual scheduler to say here's the inst the scheduler has a menu of all the different instance types that have been preconfigured in the Cod bursting admin guei so it knows that I can do this instance size and that instance size and so on and it simulates well if I burst I'm trying to burst on behalf of this job if I bursted four of these instances would that

job run the answer is yes it'll burst them and it can do that multiple times try to reduce cost and hit whatever targets you want but it is actually leveraging the PBS scheduler to make that decision via the simulation technology does that answer that question yes the schedule will provision resources on the cloud based on the the the configuration of the bursting engine right right so resour jobs that have specific resource requests jobs that reside in a certain queue for example may be eligible eligible for running in the cloud but some jobs may be on Prem only and the provision absolutely after so the question

is after the the job finishes will the cloud resources anything else that we should talk about okay no thanks a lot for having Al at your booth and we're looking forward to doing a lot more collaboration in the future with CQ thanks bye that hello hello all righty hey everyone my my name is Sanner Brown uh I'm a systems architect on the Cy national lab's uh HPC test beds team um today I'm going to be giving a little bit of background on what we do on the test beds and how we are using werewolf financeable for a lot of our system automation so Sandia Labs

uh we have three main uh branches of our test beds team we have our kind of true test beds um it's a lot of smaller systems kind of the newer Cutting Edge stuff that we working on um a lot of heterogeneous architectures it's pretty low availability uh or or kind of low Readiness um we're getting a lot of of systems that we're kind of plugging in um straight from vendors and letting loose our uh kind of expert users on um to test break let us fix it and kind of make sure that everything is working um and kind of determining where where the future is

going that way um a lot of the times it's on kind of open networks it's it's still internal Wiz but but pretty open availability um then we have our our next kind of branch which is our Vanguard uh systems um a little bit larger scale still uh experimental since it's under the test beds uh kind of scope um more focused efforts on on advancing this new technology that we're testing from the test bed side um a lot bigger kind of user base that we're working with here uh kind of expanding it from from our expert users and in a little bit more polished of of

systems here um and really the idea is is to to increase the the technology landscape um a lot of these resources on the Vanguard systems are are doe systems uh along with the the nnsa trabs um Mo most of their resources there um and then our third branch that we're working with is is the uh application Readiness test beds um these are kind of just small uh smaller twin twin really systems of a lot of the bigger leadership class systems like like Trinity and Sierra that that are out there um they're Advanced system Technologies uh they're really built to support the uh code teams with

their cicd pipelines um things like that um with with more production like availability still not not production level but what we strive for for uh production level support um down on the bottom here we see our our test beds as CIA Labs as a group we're more on the left side of this um higher risk but a lot more diversity in our architectures um and as you go more and more towards the production-like systems you get uh greater scalability but kind of less risk and and kind of lower lower diversity um so the hello there we go um so the kind of true test beds

branch that that we work on it's it's a lot more hardware and software that's that's really experimental pre-release stuff we're getting straight from the vendors uh before they've really um even tested it um where we're testing it for for for viability and and the um uh really early evaluation to see if it's a possibility for future systems um pre-release CPUs gpus other accelerators um now that we've we've got um a lot more fpgas our announcement with next silicon um and really a lot of new storage systems things like that uh then we have our application Readiness test beds like I said it's small to midscale

normally one to a few cabinets of of the really leadership class systems um that really support the the code Readiness teams um a lot of the debugging and performance shooting happens on these systems uh before it's really users are going to to the to really big systems to use um often they're they're testing their their pipelines on these machines um dedicated for debugging they've got a lot longer reservations than the big systems so people can test their their stuff here before moving it to the to really big systems um and then we've got a lot of the really Advanced uh like Systems Operations activities um

research on new ways to develop and manage systems that's when werewolf en anible comes into play um we're testing all these these workflows and stuff and and seeing if they're viable solutions for the really big systems that are out there um werewolf fuzzball um cicd pipelines a lot of different Cloud Technologies Cloud storages things like that um and then a lot of container based stuff uh container based images we're running uh uh containerized system Services podman um kubernetes containers pods and stuff for system services and then a lot of containerized workflows and and ephemeral storage so the tras T Test beds aspect of this um

a lot of the hardware stuff with pre-release CPUs it's stuff that has high bandwidth memory on chip memory um stuff they use as a combination of both um a lot of their early h100 uh gpus um GPU raid accelerators um stuff like grade and and things like that um a lot of machine learning learning and data flow accelerators uh stuff like the next silicon cards like that that we're we're working with now um and then a lot of fpga stuff uh Linux or sorry zyink um bitware Molex things like that um and then a lot of the software side um kubernetes HPC um both on

the the management side of things and also on the user side of things um we're using where the HPC uh the provisioning um uh on the test beds using werewolf for on most of our clusters now for uh um node image management and things like that um we're experimenting with Rocky Linux with uh 8 and N um which is replacing Centos for for end of life um so a lot of that kind of stuff a lot of the node images from container images um and then we're using anible to configure our containerized node images uh things like that so then we have more of the

the application Readiness test beds um one of our new cray clusters Tachi uh we've been working on quite a lot um they're uh the small to midscale versions things like uh Trinity Sierra Crossroads elcapitan um we run these systems exactly the same as the big production systems um so we're able to use these um to test things uh before they're put into production so um a lot of researchers use these for testing their codes uh debugging Performance Tuning things like that um on these smaller clusters uh so they're able to take that code and run it on the big systems before um before they're reserving

the entire node um on the big clusters and things like that um and the nice thing about this too is we're able to use these systems to to run patches and do things like that before it's put into effect on the big clusters um and it also gives us the the ability to do a lot of like the the power capping CPU trimming things like that on here yep so it's mirror systems it's all the same Hardware um but just in a smaller scale um so it's say you have a system that's 5,000 nodes on a big system we're getting maybe 50 to 100 nodes

on a smaller scale we run it in the exact same way the exact same management um systems like that so that we're able to test on a smaller scale before deploying it to a production environment so now that I have you have some of that background it's a it's a question of why why we need all this automation um so on the test beds as you can imagine we have just a an insane amount of systems and different architectures uh a lot of different hardware and software churn we're constantly testing new hardware and software um and if you think about it from a system administration

point it's it's a nightmare um so I mean just in advance like a really diverse uh pool of of different clusters and node types we can have a a single test cluster that's 10 different node types it's got different Hardware different fpgas gpus things like that in the nodes um utilizing different pipelines to and automatically like generate these images um we're doing a lot of testing with with anible gitlab container images regression testing things like that um and then also we're we're building and testing our no node images as container images um a lot of times we're pulling Docker images say of Rocky 9.2 we

pull a Docker image in we can run our ansible plays on on our container image um and then we can mirror that or copy copy our image Within werewolf and make changes to that so that we can uh distribute different node types across our cluster so a big thing is containers um we're using gitlab like I said uh mainly for our anible work um for our uh different version control and things like that we can have a base anible configuration where we're um pulling a a ready to go Rocky 9.2 bootable container image we can run our anable plays on it to configure it for

our Network and things like that uh and then once we put that image into place in werewolf we can make copies of it that is specific to our our uh system and cluster we can make uh specific changes in either run times or overlays inside of werewolf and then we can uh deploy those different images across a lot of different nodes um it makes the management a lot easier for us so that we can um boot a lot of different container images we can do testing on specific nodes um to make sure things work uh once we have a working image we can modify the

container image so that it'll boot across the rest of the nodes um with different run times and stuff um and we're testing this uh as a test beds we're prototyping the this whole process so that we can try and deploy it to our different Vanguard um and possibly CTS and ATS yeah I'd say it's a lot it's a lot quicker um because our test beds team is only five or six members right now um so I mean if we were constantly building stateful disc images um doing a lot of stuff like that and and actually modifying a entire OS image on Rocky for each individual

system it it's would take a lot lot longer we wouldn't be able to do as much much as what we do now um a little bit of what what we're doing for what we kind to calling HPC 2.0 um we're doing research on new ways of deploying and managing um system and user workflows um so using fuzzball and things like that uh on the management plane and also on the user plane uh running system uh system services and things out of PODS and kubernetes um using werewolf and fuzzball are different cicd pipelines for our file systems and different Cloud Technologies um yeah our system images

based on containers um and then ephemeral storage things like that too where we're creating pods for storage as it's running in kubernetes and then it's able to wipe out that storage after each run um an example here uh say we have a thousand node cluster we have our 10 uh 10 node management plane um we're getting all of our cicd Pipelines are Jenkins get lab things like that they're going into the fuzzball orchestrate um over the network spawning kubernetes pods all of this is running on Rocky Linux um then it gets sent over to the fuzzball substrate where it's running werewolf for all of the

node images um and then it's going to all of our different cloud storage and things like that um so I think that's about all I have for now um my team lead is going to be giving a fuzzball demonstration after this so if you guys want to stick around for that um and yeah and if you guys have any [Applause] questions so we we so we're we're mainly using anible to to configure all of the little things that we need for the sandz network and and all of the different here my name is Forest BT I'm a Solutions architect here at ciq I've been with

the company for about two and a half years now and in that time I've worked a lot on our fuzzell product so I'm here today to talk just a little bit about uh ciq and about our where to talk a little bit about our ciq company and what we do here and about our fuzzball solution that we have I've got a little demo for us that will run here in a little bit uh but in general I've got a few slides stuff like that to uh show you guys first but then we'll jump right into the guey and take a look at what we've got

um as I'm sure many of you know we are ciq we have a full stack of HPC products from the operating system to the applications level so Rocky Linux werewolf obtainer Ascender different things like that to meet the needs of your Enterprise or HPC organization uh you can talk to us all more about these products if you'd like to I won't dally on this too much um this demo is on fuzzball so fuzzball is our big HPC orchestration and pipeline kind of workflow enging app uh that we've been working on for some time now uh basically we right now we're all very familiar with the

bolf model um this is traditional HPC this is what we've been doing for a good uh 20 30 years now uh most traditional HPC systems that we have today are built based on this architecture you know standard flat head node with a whole bunch of uh flat array of compute nodes underneath it this is a fantastic model obviously this has powered our science for a good 20 30 years now and has been um one of the primary drivers of you know everything HPC um this is a fantastic model but kind of as we get into the uh 2020s and we kind of start to see

HBC Advance a little bit we've started to see changing needs within both HBC and Enterprise that have kind of led us to develop a new model that we call HBC 2.0 uh HBC 2.0 is kind of implemented in our fuzzball product uh and why exactly are we seeing kind of this move away from traditional HPC uh we have a few different things going on at the moment uh first off we're seeing use cases becoming more and more complicated science is advancing pipelines are getting more and more complex just in general things are always getting harder and harder to compute um use cases are getting more

complex and we're finding that we need the ability to orchestrate larger and larger clusters in order to make these things continue to work uh we kind of have this split model at the moment where uh HBC has all of the high performance gement uh kubernetes that type of thing that they've been establishing for years um that helps manage these massive large systems that we're starting to need in HPC uh so in general uh we kind of see that split we see things like AI becoming a bigger and bigger use case everywhere and the need for larger and larger clusters to support that um kind of

in general we've always had this model where we half expect users to be a little bit of computer scientist themselves even if they're a domain scientist um we'd like to move away from that and get more of a focus back on users being able to uh just focus on their science and just have us you know be helping them set that up but not telling them so much how to do it based on uh you know kind of what we have to do right now with interacting on the Linux command line things like that um so we'd like to get away from that uh the

cloud is a big question in HPC of course there's a lot of interest in that a lot of uh kind of a be able to integrate that with our on-prem systems uh so that's something big that we're looking at uh supply chain security you know we're all using containers we're trying to bring those into HPC that type of thing uh that especially recently has had a big push to kind of make sure that we can track back where software comes from with s bombs and have more signing and verification and things like that around containers um so we in general like to be able to

do a little bit more with those in HPC than we can do right now uh and just in general like I said organizations that I've never seen themselves as HBC organizations are starting to consume this and uh HBC organizations are starting to need orchestration and management like they haven't needed before so we have a few different driving forces that are kind of leading us to building out this new version of HBC like I said HBC 2.0 uh and kind of the fuzzball product to give you guys uh basically the architecture diagram of what we're working on uh I see my slide is a little off

there um one of the biggest things about fuzzball is that we're integrating kubernetes with h HBC I think we've all kind of seen some of the batch Computing efforts over time that have come out with fuzzball uh and HBC uh but in general these all kind of fall a foul of the fact that kubernetes was never developed as a batch Computing engine and there's overhead that happens if you're trying to run HPC apps in pods that decreases your performance um in general kubernetes is a little bit of a Byzantine uh Concept in HPC in general we're all very busy so being able to sit down

and actually get into kubernetes and figure out how you can integrate it in your HBC environment is difficult uh yet kubernetes is ubiquitously used in the Enterprise you know basically anything you interact with on a daily basis is probably kubernetes under the hood one way or another um so why aren't we using it in HPC like I said because there's traditionally never been an integration between uh kubernetes and HPC that has used the best practices of both and been able to integrate those effectively so with fuzzball uh we are building kind of an all-in-one automation platform around HPC that allows you to instead of um

running this more conventional HPC stock run this fuzzball orchestrate piece of software uh in kind of a more standard fuzzball or uh kubernetes microservice type architecture so essentially what we're looking at here we have a control plane uh for our fuzzball product that runs somewhere up in the cloud this is on some type of managed kubernetes like eks something like that on AWS uh in just a moment the demo that I'll show you is going to be up on EOS um so that's uh where that's running um do you have some type of magic kubernetes that this control plane is running on the control plane

uh ultimately consists of a series of microservices that provide the overall fuzzball kind of orchestration stock uh ultimately what this gives you is the ability for users to codify their workflows you know that standard input data comes in pre-processing run your analysis postprocessing results go out you can codify those pipelines regardless basically of what type of you know vertical of uh science you're working with we've done AI training cfd FAA a wide variety of different use cases on this um regardless of uh what your use case is uh I'm so sorry I lost my truck thought uh we have uh like I said oh yeah

sorry users can like I said put together their workflows codify their stuff they send it into uh the API endpoint for the fuzzball cluster everything in fuzzball like I said there's no sshing onto uh a cluster there's no having to get onto a login node or head node or anything like that um everything in fuzzball is apid driven so users are able to either use a CLI or a guie to interact with the cluster um and essentially take these workflows that they build in yaml or with our graphical Builder send them into the fuzzball cluster and in this series of microservices kind of breaks apart

that workflow into its component parts and orchestrates the rest of the running of it so there's a workflow engine that takes all the different requirements sends them out to the other microservices there's a data mover that handles any S3 transfers or other data transfers that need to happen before the job can run a volume manager that sets up any that type of thing that uh the workflow will need to run uh everything in fuzzball is based on containers so all jobs that users are running are done out of a container so uh there's an image service there that manages that one important thing to note

is that hello there we go um one thing to note is that none of the actual uh applications are running inside of kubernetes hello none of the actual uh applications themselves are running in kubernetes pods only the orchestrate Services themselves are running in that so we've kind of implemented this best practice of only using kubernetes for the microservice orchestration plane and then that does all the the other uh interactions that a workflow needs in order to run so for example when a user's job needs some resources is able uh to be able to run out on the cloud our provisioner service takes those uh requirements

sends that out to AWS it spins up an instance based on um those apis it lands the Ami on it that has our custom uh container runtime fuzzball substrate and is able to land a container and a command on that and start running it independent of needing to have a kubernetes pod doing that with kubernetes itself just kind of monitoring those instances and waiting for them to be done uh and so yeah you have the control plane the stack of microservices that are providing what exactly we call fuzo orchestrate in this uh on the other side of it on your computer resources yourselves you have

some type of uh operating system of course we're real big on Rocky Linux um and then our custom container runtime fuzzball substrate as I said when nodes are spun up in this out on a cloud provider this is installed onto the nodes and it's a custom container runtime that the rest of the cluster can then drop a container and a command onto and have run on those arbitrary resources so a function that's a service um and yeah overall the process like I said Ingress you've got volumes you can persist data you've got uh volumes that you can use to process data back and forth um

through uh between jobs and a workflow uh and yeah and end this is kind of the concept of fuzzball orchestrate so I have uh a demo that we are going to take a look at uh once I move over to that um is there anything else I wanted to cover no I think we're good so to actually give you guys an example of uh what we are working on here and what this looks like from a user perspective I will go ahead and log into our instance here oh there we go so as I said everything in fuzzball is all API driven there's no users

logging in having to interact with the command line doing that type of thing uh it's all API calls versus a cluster endpoint so everything you're seeing me for the most part except for like the graphical editor here for the most part all these cluster management operations that you're going to see me do here are all things that uh can also be done via the command line through the fuzzball CLI it just so happens that we're here in the graphical one because that's more fun to look at so this is the actual workflow screen um like I said these are runs of different HPC pipelines essentially

whether that's uh some type of cfd job some type of AI training some type of uh like I said represented just in the EML so if you know the syntax it's easy enough but we provide the graphical Builder uh to make it easier like I said for those domain scientists that don't want to have to learn as much of the Linux command line and to just unpack those tarballs and then we actually run our uh example training script for llama 2 so if we want to see what this actually actually I'll go do this um this is the actual workflow editor so you're looking at

essentially a graphical interface to put together a uh HBC workflow end to end so this is the same type of thing that people might be doing when they're writing scripts except for that uh now we can do this graphically and you know all wired back into the uh cluster so we can run this when we're done on the resources that we've uh on the resources that we have here that we've defined here I'll show you that in a moment uh you can see that when we want to define a job in fuzzball we have a few different things that we can add uh first off

over here is a command so this is essentially a shell built-in a compiled application anything like that that you are uh looking to an MPI job or a task array basically uh task array being you know that embarrassingly parallel type architecture uh multinode being MPI and also gas net I don't know if anyone uses pgas uses uh you know the chapel programming language anything like that um but this also has support for that as a multi-node uh implementation alongside MPI uh we can also edit the environment of a given job so all jobs in fuzzball execute out of containers and in this case we are

uh accessing a gcp artifact registry with uh just a username and a basically access token here this is wired back into Hashi Corp Vault so you can store uh Secret within this cluster and then have them be templated out at runtime via this type of format so when it gets to the server side it gets inserted there and uh users never have to keep their credentials in their workflows we have a we can mount storage volumes to these so we can have essentially scratch space for our different jobs and workflows and things like that to use or our different jobs as a part of the

workflow to use uh if we go to volumes you can see where we've defined that uh we're calling this the Llama Barn uh but you can see I've got this Ingress configuration here where I'm BAS basically just providing my AWS keys in a region and uh specifying that I want this tarball pulled down with that model file uh and doing the same here with the tokenizer we could add an egress if we wanted to push back out graphs results another like fine-tuned model uh we don't have one here but we could configure that same thing just by pushing it back up to an S3 resource

uh like I said everything is done out of a container we can mount volumes we can also decide the resources that any given job needs in order to run uh so this one needs eight CPU cores 28 GB bytes of memory uh this one right here and you'll notice that this is a directed a cylic graph of execution so as long as you have no Cycles in this you can essentially build you can build this dependency Matrix essentially as arbitrarily complex as you need uh if we go look over here we can see a different resource request along with uh one nvidia.com GPU that we're

requesting uh this is our device plugin so specifying device files drivers and libraries for device and giving fuzzball the ability to bring that up and address that if you have those resources on your compute nodes um we also have support for uh certain fpgas and things like that uh but in general this is uh working for gpus at the moment um I don't have a demo of it right now but we also have uh the ability to spin up isolated network name spaces and use a port forward to bring things like jupyter notebook up on this um so you can run this workflow it starts

a jupyter notebook out on ad BS resources and you can use this port forwarding tool via the CLI uh to bring a port from that so you can then you know Local Host 8888 and get access to what's we'll go ahead and as a part of that job um but resources we're actually going to be running this on is contained in the definitions uh so for the moment we're focused on the cloud vertical with fuzzball uh so it works on AWS uh we're working on bringing it to other clouds especially gcp at the moment um but for the moment uh it's very oper on on

uh on AWS up to including like the elastic fabric adapter and things like that uh you can see that most of our compute sources here are instance definitions from AWS so like a p3.2 x large with a certain Ami there uh so as I said our provisioner service will reach out match the resource requests that a user requests in their job with the instance that best matches it uh prisoner will reach out via AWS spin up an instance matching that this Ami with our fuzzball substrate runtime is landed on that the uh instance is then leased back to the fuzzball cluster the job scheduler is

able to take that instance land the or give it to the workflow land the container on it land the command on it and boom your workflow is running uh and so these are how you define a different compute definitions that's a p3.2 x large so a GPU instance uh we can do you know t2. 2x larges for testing uh essentially any arbitrary instance you want to add to this uh it's essentially just adding the Ami the instance type and any devices that it has on it that it needs to be aware of um as far as the definition goes uh we've got Secrets stored in

here you can see different AWS access Keys container registry access Keys things like that um as far as container Registries go uh this like I said interacts with anything public or private uh as long as you can give it some type of username and a uh account credential uh you're good to go there's kind of a uh account structure that works with organizations we call organiz and accounts so in this case CQ is our organization and we can have different accounts down below that that all have their own users their own Secrets their own definitions workflows Etc so you can imagine that as something like

you have one fuzzball instance for your University and then different Labs have their different accounts users Etc out below it um this is uh going to basically run in the background here but just to show you one last thing we can retrieve the logs from a workflow uh those are automatically captured by fuzzball except for uh on this one apparently that's interesting see it grabbed it on one of these other ones huh that's odd um when I was shown this earlier there was logs there I'm maybe I'm on the wrong workflow or something uh anyway so as things like that are printed uh they get

collected um there's a log tab that prints that live there's also a terminal uh you can get directly onto the instance that the job is running on I would show you that but it's going to take a few minutes to to get to this run llama example here at the end um these onars are pretty quick but uh here in a moment I might be able to log in there um so yeah that end to end is what fuzzball is it's kind of in a closed beta right now so we are taking Partners in for our SAS platform uh kind of selectively um but if

you're very interested in talking to us more about it I would encourage you to come and have a conversation uh that is the fuzzball platform are there any questions I can answer what's up so the question was uh right now it's mostly on the cloud what are we thinking for on Prem uh yes at the moment we are focused very heavily on the cloud vertical for this we have a decent amount of sophistication in the on Prem case uh but for the moment um with everything we've got going on we're mostly focused more so on the the cloud side is that's uh a little bit

um I don't want to say low hanging fruit but it's a little bit more approachable than some of the on-prem stuff but we do have a great level of sophistication there um we've got for example an automated installer for on Prem uh that can basically TurnKey bring up a fuzzball cluster there um but we kind of have that at select Partners at the moment that we're working with there but yeah a lot of sophistication on it coming soon yeah ultimately our goal with fuzzball is to enable hybrid Computing between all of your data centers so we'd like you to be able to have a single

pane of glass that includes all of your Cloud resources your you know East Coast Data Center your West Coast Data Center your European APAC Etc data center all put together in such a way that users can codify what they need to do submit that workflow to one single endpoint fuzzball is aware of all these other clusters that you have that are also on fuzzball and you're able to score based on cost resource availability more esoteric things like custom uh parameters carbon credits stuff like that uh where exactly that workflow should run on all those different clusters that you have wired into that that's kind of

a road map uh we've got a little ways to go before we get there but like I said we've got some sophistication on on Prem we've got a lot of sophistication in AWS our uh road map for you know next year or so so is all about improving our Cloud offering and starting to provide that Federation layer so that we can have that true hybrid Computing model between on Prem and Cloud resources you know you're defining a workflow at this point so for the user it becomes one more workflow they have to learn in addition to all the other workflow Brides for example next flow

you guys youning to any so the question is what's the plan for the integration with other workflow engines that are out there yeah obviously there are a lot of them wdl snake mate Cromwell um all that type of stuff um at least in the next flow case something that we're looking at is for example how they build executors for different um uh you know kubernetes AWS things like that next flow has a um yeah so that is fuzzball like I said MPI task arrays even uh pgas support and stuff like that and uh that's the platform thank you all and once again thank you all

for attending I'm just curious how big is your Dev team behind fuzzball about one of the things that we have in it which is the mainline stable kernel or builds of the mainline stable kernel that we have for Rocky Linux so it's pretty simple demo today but please feel free to stop me and ask any questions that you might have so as I mentioned we'll be using two ciq products or you know oh hello yes we'll be using two products that uh ciq supports at least uh a product of vs Mountain the other is Rocky Linux which of course is a a community product um

that we have value ads on top of so what is Mountain uh mountain is a our content repository and content delivery mechanism so um right now we're supporting RPM packages predominantly it has yum repositories that you can subscribe to in it uh so you can use yum or dnf uh to subscribe and get packages just like you would with any other package repository but with that we also support oci um Registries so we can push uh oci containers and kind of the whole stack of that up into a registry uh to supply that as a product to our customers and we also can support CIF

container images in that same registry uh so if you're using Apper or Singularity or even werewolf uh you can pull those images directly down from Mountain from a subscription that you have with your CI support contract um but the future of the mountain platform is to expand it as kind of a hub for our partners so not only ciq products would be in there but we would want to have uh isvs as well as a growing catalog of Open Source software pre-built ready to go in subscriptions for you to run so that you can in ball deployment or pull it directly into aper and not

have to build those contain thing um and in the kind of parallel future we're developing a an edge mirror version of this I think I I just heard that this is basically done the uh the the kind of first well now second version of it that we're going to start being pushing out as a product but it's not available today but the idea here is synchronized products down be able to add your own packages to uh The Edge mirror that you deploy in your environment to not have to go out over the Internet uh for cluster of nodes um this is what it looks like

today uh this is kind of not the admin view but an admin's access to it you know list of products each one of these is some kind of registry or RPM repository um you log into this web UI generate an access key and then put that in with a mountain CLI application uh that we'll demonstrate uh so I mentioned a couple of these things uh things that you can find in mountain today for our partners that we doing uh fuzzball beta test testing with you can get Buzz fuzzball out of here for for certain on-prem deployments uh we Supply aper and werewolf images here if

you're familiar with uh werewolf it uses uh or it can use um oci um container images as it's node pull those directly out of mountain and and uh import them into werewolf directly oh oh I actually skipped a little bit Apper and werewolf packages for the open source stuff but if you want builds for you can get packages from there you can get werewolf node images um we do Supply CQ custom patches and builds for various products and and patches on top of Rocky Linux um in the F oh not in the future I'm thinking of the kernel LTS release uh we we Mark certain

versions of Rocky Linux as long-term support and backport security uh patches into even numbered Point releases of Rocky Linux so those are available there but last on our list here is uh the mainline stable cel uh which is a value ad that we provide on top of Rocky so just to be clear about what I'm talking about uh if you go to kernel.org you can see the releases of the Linux kernel as they come off of the developers and they have Mainline which is literally the the head of development uh all patches that are accepted by Linus to Torvalds are in that main line but

then on a release Cadence there's a stable release of that main line and so that's what we're taking we're taking that and building it in a version that integrates successfully with Rocky is built to work with Rocky but has uh more up-to-date Hardware support and additional kernel features than what you would normally get in an Enterprise Linux kernel for comparison if you're not familiar uh the the Enterprise Linux kernel distribution strategy is to take a a specific kernel release keep it stable at that point from an ABI standpoint and then for the next 10 years backport patches onto it maintaining that compatibility the security posture

that you need so the Enterprise Linux 8 kernel is 4.18 the Enterprise Linux N9 kernel is 5.14 uh they were released in 2018 and 2021 originally respectively uh but the mainline stable kernel that was in the ciq repository when I prepared this we'll see if there's a newer one when we run the demo here uh was released in August and that's the kind of access to the most recent kernel that you can expect through this offer not to put too fine a point on it but if you're still on sentos 7 it's running uh Linux 3.10 which is more than 10 years old now at

this point uh so just just you know another pointer it's time it's time to get off of Centos 7 uh so this is where we will switch over to an attempted live demo H one hand yeah that's the worst part I'm realizing now that I've left my notes over here not that I can't just do this but uh it's good to have a backup plan oh and now I've disrupted that okay so what we've got here got two windows here into the same system this is um you know a a VGA console off of a cloud provider we're using vulture here uh you can visit

their Booth down that way in fact um and it's running Rocky 8 and you can see that here in this uh SSH session and we want to take it to Mainline stable kernel so uh the first thing that we need to do is download this mountain application it's a simple command line uh binary and just a moment I'm going to put the mic down so I can type with two hands excellent so that's the first hard thing to type so what we're getting here is a package that's on a completely open yum repository uh I guess it's not even a yum repository we're installing the

Yum repository here uh but this is a package that anyone can install um that's just you know on our website and this enables a ciq repository that we can then install our Mountain application from that repository is this big enough as can people see this here okay great uh so now we have this mountain binary and this is our gateway to the the ciq mountain SAS platform so the first thing we need to do is attach it to our subscription like I mentioned before what you would do is go into the web interface and you just say I need an access key and you get

it for the purposes of this demo I've already pre-populated that in a uh an environment variable so I don't have but this is one command more interesting than the grass growing I mean mountain is the the delivery mechanism that we use to give it so there there's Mountain the uh the SAS platform that's where we serve it out of um but uh in theory you wouldn't need to use the mountain command line application to get to it you use that to generate the Yum uh configuration the first time at that point you could take it out and put it in your configuration management or like

that so that that is intended with the platform it's not in the platform today yes yes I I'm afraid I don't have a time for it hello um what I can say is that the the mountain platform in the Enterprise Linux ecosystem that we're building on top of Rocky is kind of number one priority for us right now it's where most of our attention is going so I I expect soon uh but oh yes within the next year absolutely any other questions it's it's like the opposite of a dial up internet connection there we go so package is done installing not too exciting at this

point but at that point we can just reboot our server we can watch the reboot happen in our console over here on the side uh if we look carefully we can see that it's already uh showing us that it's going to boot into that 6.5 kernel and here it is hopefully our SSH is back up and running yep there and from here we can see that we're still running Rocky 8 but we're running this much much newer colonel that was built a month ago uh so pros and cons here this new kernel should have you know more security patches in uh integrated uh newer Hardware

support newer kernel features in theory you have new file systems the uh the Enterprise Linux kernel doesn't support btrfs for example this kernel will uh you won't have the btrfs user land so we'd need to also bring in packages for that uh but it it's a gateway to much more functionality but there's a reason why this isn't the default kernel and that's because when we talk about the Enterprise Linux platform uh we're talking about a kernel compatibility layer that will allow you to bring in out of tree kernel modules as well things like a luster client or a luster server uh other file system clients

melanox driver Nvidia driver things like that that are built against the kernel for that platform so one of the things that ciq is doing is this trade partnership between ciq Susa and Oracle to try and drive adoption of a new standard for what Enterprise Linux is and one of the things that we're hoping to prioritize there is uh newer kernels and access to newer Colonels as part of that standard um so that's my presentation that's my demo uh as before if there are questions I'm happy to take them but otherwise thank you for your time and and uh and for sticking [Applause] around hello to

the ciq audience your strategic man strategic marketing manager uh direct at pan panassus I'm going to be talking to you about a little curveball for what you've been probably listening to relative to uh Rocky Linux and other ciq but it's within the HPC ecosystem we do parallel file systems we do uh software the stack we do uh Hardware Appliances that we provide to the market and I'm going to talk a little bit a little couple stories one on Modern HPC and where we see that market going I'm going to make a analogy from my past so we've been around since uh 1999 since the Inception

really of a lot of the parallel file systems in fact we're the first parallel file system to commercially Deploy on Linux um we have a was our founder was coinventors of raid at UC Berkeley G Gibson and from that you'll see this story about resiliency and reliability because it's like built into the DNA of the company everything else we talk about performance manageability all that usability all that is riding on top of the fact that we have this reliable system that's very self-healing fault uh tolerant and fault prevented this is effectively uh a a block diagram of panfs so the idea is that you have

client applications we have our direct flow client that would run on a Linux and we'll talk about later about Rocky Linux and where we support Rocky Linux in addition we'll have we have and we just announced an an S3 API interface support through a Gateway so it'll come through through our Gateway site SNB and NFS support and then we have a couple anary software products that help you to tie this all together relative to analytics and data movement one of the more semal patents that we have in our company is on a per file eraser coding so this is unique in the marketplace and allows

us to have granularity to the file level if in just in terms of being able to do things like uh set a raid level a eraser code level level at a file level uh it gives you a lot of flexib provides you a lot of uh fault tolerance capability the way we manifest our our software is we use storage nodes and we use director nodes so this is all the metadata management has happened on the director nodes all the data itself and all the metadata itself is being loaded on our on our uh on our uh storage and we call these storage object stores so

this is my story here about where we see the modern HPC so I think everybody's familiar with traditional HPC being more modeling and simulation type workloads Computer Engineering computation of fluid dynamics and then we saw this Mig ation happening uh this this integration happening between things like uh uh uh big data and uh Hadoop and uh analytics being incorporated into HPC so HPC had this new definition that said it's not just traditional modeling simulation in the National Labs and those type of places but is actually being used for analytics as well and then I like to quote a couple guys in inidia who actually coined

a term on um on this HPC and AI so it's the first level integration of HPC and AI where effectively you are doing pre-processing you do some computation you could do some post-processing and in this computation method you will like break out and do some Ai and come back and we'll talk about how this is changing and we're working with different partners in terms of something that's called augment yeah I think I'm touching the mute button um and so this is all happening on Prem it's happening at the edge it's happening in the cloud and so Nicholas debate from HP actually coin this term that

I really like which is HPC and internet workflows and it's really the change that's you're going to be seeing happening throughout any kind of design industry today that you can actually have integrated AI meaning that you can simulate like if you wanted to do a cftd flow and you wanted to do you can simulate your data and it'll get you instead of doing this wide variance in terms of where your simulation and what your accuracy and how you're going to close off on it you can simulate your data you can train on that data and help you point in the right direction and save you

a couple orders of magnitude of time in terms of uh getting to your solution and an analogy I use is like you're driving on a roadway and you're on Surface Street and you don't know you're kind of figuring out where your destination is and you hop on the freeway and it get you all the way to the nearest exit and you get off and you still have to do your closure at the end but meanwhile you save so much time and being able to get there except in Southern California where we know that there are traffic so here's my little story so I it was

a time where uh we were building an automated highway system I was actually working on a project that was part of a big conglomerate of people with calr and GTV back then and all and it was going to be the automated highway system that was going to deploy like in 2025 or 20 2015 2015 and it needed all this infrastructure in place so anything for collision avoidance anything for um just uh uh latitude and longitude in control of your vehicles and at that time the technology trying to use vision technology uh uh combined with we were actually using ceramic tiles on the road to be

able to give you a a Latitude control and of course you had all the different ladar Lars and lasers and radar technology to give you the the uh the control for the longitude and control and then we're using things simple like on-ramp and metering and stuff to try to keep the flow going so this is an optimization problem and it's not unlike an optimization problem for me when we look at HPC storage so you think of that and you think about what are the things that you need to worry about you need to worry about reliability the robustness of the system you need to worry

about things like flow control you need to worry about self-healing you don't you want and so so I actually did like a failure modes analysis on an H uh just like a just like an intersection today is optimized for you to have rear end collisions not to have hit- on collisions in an automated highway system the idea was to have no massive just like Airlines and stuff no massive collisions and to be able to have this this idea of platooning vehicles together you do the spontaneous platooning but you limit the number of vehicles that can be together and then you have this sort of like

this distributed system that allow you to be able to program your navigation system and there are many pundits who ended up saying this type of stuff is called a railroad system cuz you actually G gain uh railroad cars together and all but this was this was like being rolled out in Germany with the freight liner trucks and all and being able to create these big platoon of trucks down the highway to get better efficiency so with that segue this goes down to in that premise of how do you build a system for us a storage system so we have some core values that we we

so EAS it so so back in the day when parallel fall system was just happening in gpf s was out there and was running on a and we were working with luster to try to get commercial parallel f file systems to the market the primary thing was to have an ease of use scenario so that uh uh so that that that hands-off Administration we have customers who say it just works we don't even think about it um 99.9 you know 39s more uptime everything was about uh being able to make it easy to use but that's based on having something that's resilient and being able

to have a system that had a shareed nothing with no single point of failure system selfhealing our failure our F Eros coding this those two come together and of course the the main thing that in the parallel system that you want to have is a high performance high capacity unlimited capacity how much can you scale how much can you uh get performance perfor uh the other thing that we actually added in for our hybrid solution was this great technology and I'll show you in just in a second it's called it's an intelligent placement technology we call Dynamic data acceleration to take advantage of the hardware

that's available and for us that means that if you have small files you put them on as Prestige you put large files you put them on hard drives you have metadata DME flash so you're trying to optimize whatever media you have available for you in the system as it gets deployed so panys is our bread and butter it's our software it's a uh it but it's more than a file system and includes a whole software suite and we have a active manager product that all you to be able to set different levels of uh alerts and uh uh of just email um feedback to the

to the administrators we have a security that we launch a couple years in terms of of a hardware based encryption so the idea is that to let your compute nodes your compute nodes to be able to optimize its performance and so we use a uh secure encrypted Drive technology and we do and we worked with uh uh key manager companies like like Talis and Intrust to be able to provide an integrated solution so that you can actually have the key managements working with encryption at rest also support SE Linux so we have uh uh file security label support then I talked about these two different

products we have for mobility and for visibility so the idea is that you can know everything in your net our n our product but everything else so if you want to move data around and if you want to set from a Mobility point of view you want to set policies for retention levels you want to move stuff to the cloud we support the S3 uh and Swift uh and and I'll talk about that's like we can store things from our panfs system into the cloud and um and I've got a great ey chart for you later that'll show all this in place so this is

my first complicated slide but it's just to show you that how we and trying to figure out to let you know our job is to try to make any system administrator's job so simple and how do you do that is you prevent hop spots and how do you do that kind of thing so we we start off from the left here we do um fire level creation ratio so every file comes in we turn it into object components and we add the p and Q so we do a two parody uh value system and we use like raid level terminology layout so you can use

Raid five or Raid six and that depends on how many parody you're going to do and typically for the smaller files we will do replication because eraser coding doesn't really help with smaller files and so you can do but we'll use a standard xor for five and then exor plus a resolin for rate six uh so two parody really rate six just means two par in that contest and then we spread it across our networks and we try to use what we call a you know a distinct uniform random distribution right so this this aggregated solution try to keep things level it out and then

we do this active capacity balancing when especially when you have reconstruction happening if something fails and we're like fixing things on the fly to keep the system balanced and that's to make sure that there are no hot spots and then I talked about this we have MV dims mvme ssds and idea is to be able to leverage the media that we have available to us and then at the very T tail end then we want if you as a system administrator want to move uh files uh uh to different archived levels uh to different cloud different S3 objects then we provide that capability so we

instantiate our pan software on appliances today and we have a family of appliances different hybrid products we have an all flash product product and we have some smaller capacity products one of announcements that we made at the show today is working with Seagate segate announced uh uh uh collaboration with us on using their dual actuator s SATA drives so effectively you have one disc drive now that has two actuators on it and you get double the iops double the throughput off of it effectively for the same price and it's fairly much a drop and replace we had to do some software obviously to integrate with

it and they've done a lot of ecosystem work to get all the the Linux sack and um uh uh controller companies on board but the idea is that um that you will effectively have with a SAS Drive um two luns that would appear and we now address them as two separate ports for us to be able to use them in our system like we do as SAS drives we have our director products and we also have these are all ethernet based they have uh two dual 25 gig ethernet ports on them and then if you're running infin band on the front end and uh we

ourselves um have a infin band router so it's infiniband on one side ethernet on the other so that you can route infin band base uh compute cluster information down to the storage side we also work with cornales networks on omnip paath as a separate fabric uh so they have uh uh omnip paath hcas and omnip paath uh switches and in and as well as omnip paath routers so this is my big diagram just to show you and where we're so if if you look at where we sit and you're running these applications uh so direct flow client for us is a it's a direct and

parallel access from the clients in posic so it's cash coherent from the posit uh from the clients access to all the storage that gets distributed this data that gets distributed gets distributed as we talked about with the profile eraser coding across these if any of these drives or anything would fail we would reconstruct that drive that was being failed while the system is still running if you're using other protocols like NFS S3 and SNB that'll come in through our Gateway Services uh so if if for example you wanted to ask um if you had written something to a Seth storage and it was using S3

and now you wanted to use that same application and talk to panst storage then you can come through our Gateway service and use that application and come in and talk to uh panes similarly you can take our panes and you can send each file up and becomes an object in the cloud so in Azure blob Google file store AWS S3 or any Swift or S S3 Cloud instance or any private uh Cloud Object Store instance and if you're writing like microservices space Cloud native applications you can talk to these objects here through our s API so to be able to move this faster we use

uh our our technology partner Temple and all to use their concept of data movers and these data movers all support Rocky Linux as well as do our compute clients for NFS uh SNB uh as well as S3 and for our for our for our direct flow client that you can run um uh on Rocky Linux and that's pretty much it so call to action please stop by our booth you're here already in a CQ Booth uh go to seat Booth to learn more about this dual actuator this is like really groundbreaking type technology in terms of just doubling the performance just using multiple actuators and

of course uh one of the things that I was talking about earlier with alter had launched their hyperworks product that allows us integrated um AI workflow through their um yeah through the hyperworks product and they're in Booth 825 are there any questions any tests anybody know what uh parall Fus okay thank you very much I appreciate [Applause] it overview of that and we are partners with ciq as part of Rocky Linux and delivering kind of high performance solutions in terms of storage and like AMD servers and Rocky Linux and that's combined and then we are the um application that sits on top of that so

I'll give you a little information about the company um this is a table content because we have a ton of tools right we have lots of different tools if you're doing any kind of genomic data analysis at all um there's different applications for that I'll skip over most of those but uh we can we can talk about them um so at cention our mission is to enable Precision data for precision medicine and what that means to us is giving our customers the opportunity to process that data in extremely accurate efficient um and cost-effective way and we do that uh based on oops our engineering team

uh and scalable and accurate algorithms our core strength is um accurately and efficiently solving math problems in computers and a lot of people would say this is a supercomputing conference and a lot of people have that core strength um I would say one of our employees was the number one physics student in all of China in 1988 and he was two years younger than everyone else who was taking that physics test and he's a big contributor to free BSD which I guess is now the Apple operating system so he used to make freeing put it through a genomic sequencer um our data or our software

will help you process it um so the company has been around it's been founded in 2014 in the Bay Area I'm lucky enough to live here in Denver I hope you've all been enjoying uh the beautiful weather here it's a little warm uh for this time year but it is typically Sunny like this um let me know if you have want a restaurant recommendation I can give you a lot of that as well um so like I said we run on any CPU or arm um we work very well on top of um whatever Linux you're using uh ciq is a great partner of ours

um and uh we provide I'm the head of Business Development and we provide a very friendly um uh I leave you alone and you leave me alone which is good but also this we've won many awards in both that because now our field is has an opportunity to grade the accuracy better than it did when we started eight years ago but when we started we wanted to say you know if you like your pipeline you can keep your pipeline you just have better comput algorithms that are realizing the same math so we basically have a per core efficiency Improvement of generally like 5 to 10x

per step um and then we also have a better software implementation so it's self-contained it uses lower RAM and it's um automatically parallelizable to whatever size machine you have um and then also what helps with that right is uh there's no thread dependency so a lot of um the other tools also have thread they have thread dependency they down sampling that's kind of part of the consistency challenges um and I mean the software works great but we're also here to support you to kind of let you know um why the tools are calling the data in a certain way or whatever so typically the support

is um on the results like helping a customer understand the results or um understanding the tools and pipelines this is a command line based software so there's a lot of detail in the manual and I tell people to you know use our support team uh like the manual which um they're okay with they don't love it but I I sure you guys understand um also another thing to kind of highlight in that scalability um and this is also why we have consistency is we have no down sampling because we've been able to handle um High depth data better so the other um tools in the

field basically can't hand what part of what's going on is a hidden Markov model where they're looking at all of these strings of 150 base pairs so that's like a read and you generate you know basically these machines will generate like billions of reads a day now but in your genome all these 150 strings could get stacked up into a cover average of 30X which many tools will be able to do but some um applications you want to sequence deeper so you might want to sequence a THX or 10,000x or even 100,000x and in that case all the reads like under a region are kind

of looked at in the context of each other so then the compute problem really explodes and our team is able to handle that very efficiently so we don't do any down sampling um and it still run you know five or 10 times faster um also so this is kind of the main goal of bionformatics analysis software right and maybe this is also if you're familiar with other HPC data processing these are similar things that you want to be good at um so one is if you're accurate and consistent uh which we are we've won um Awards on this and this can all be verified um

also just efficiently meet your turnaround time goals um our software is distributable onto as many servers as someone would want which I don't really recommend because it already does a great job on single servers typically um and also it scales down when people have um like a braa panel like a smaller genan panel um which is also very useful in the cloud where if you're processing lots of tiny jobs you don't want to put those on a giant server uh you can just put them on smaller servers uh so our software really scales up and down so you can efficiently meet you know whatever turnaround

time you need need either locally or in the cloud um with extremely low total cost of ownership um this is not bu informatics is not like a crazy money thing right it's more it's not the it's more the Lords work um and then luckily that these engineer like people say where are the engineers who are really good who are solving this problem um and they're at our main competitor alumina uh that's also good it's a good product the dragon system um and then I'd say we are the other uh best software in the world for this um and we are the best sequencer agnostic software

in the world by far um alumina is kind of trying to do a wall Garden approach right and say you know start with this kind of Premium model and then pay us and go into our platform right um so those are kind of the two good options I would say um and if you're not using one of those two options then you should definitely talk to me and you can talk to me also if you're using the alumina dragon system as that's our primary competitor that we do you know compete well against all the time um and like I said multiplatform and um super easy

to use as far as enterprise software goes I am not a software engineer so um I can run it like in The Bash but not super well but um I talk to our customers all the time and they're always extremely pleasantly surprised with how uh nice it is hey Glen um so this now we're going to just start skipping through some stuff because this is like a long thing here's just one example though of just pipeline complexity so another tool that we compete with is or we replace is just a the open source BW gatk so we have made um usage improvements when we can

um in the command line so you can see in this like gatk has 3,400 lines um that are used and we only have 567 so there is less headache less scripting involved and we also um are working on releasing we you know provide to you kind of example scripts to use um but we're also working on as our world has kind of gotten a lot bigger and we're supporting more than just alumina um and we have kind of complex P bio pipelines complex Oxford nanoport pipelines um we're wrapping that kind of stuff up into python so that the user can say great like run the

P bioscript or run the element bi whatever it is so we are improving our usability um Beyond this um here's aot quote though from one of our customers everyone in America um has heard of Mayo Clinic so we use that as a hey like Mayo Clinic likes it and wrote a paper about it um and in their paper this is more of a computer science speed up we're not 50 times faster depends where you measure from and what you're measuring um here's another example uh this is kind of also what I was talking about this is some work we did with Intel Benchmark it was

taking us um 12.9 minutes so over 1 minute faster than the h100 and um about 22 times cheaper um and so and and the Nvidia is is a good system relative to the BW g8k um open source tools provided you have um gpus but with us you don't need gpus um also uh we have I mentioned before an improved accuracy tool so we have um you know if you want to run a BW GK pipeline we have that accelerated kind of like a software accelerated Pipeline and MCH 2 and Joint calling and all of those different things um but we also have a more accurate

Tool uh which basically has an improved local assembler um and some machine learning and that's what we use um for people who want who are more into accuracy than matching um and again I we support a lot of different sequencer so this is one example of a new sequencer um and the data is quite a bit different and so um we work we partnered with them and they also partnered with the broad Institute to kind of develop a pipeline specifically for their sequence data types and you can see um we're in the red here while the broad Institute tool is in the blue and you

can just see we have a significant accuracy Improvement um relative to kind of the Bro institute's best effort for this um and then also uh I think we're around eight times more efficient per thread compared to that solution um it cost us about a dollar on demand to process a 40x Ultima Bam um here are some old benchmarks uh this kind of I already mentioned it a bunch because I'm excited about this 12.9 minutes uh one but this is all basically our new release combined with the new um new servers available on Amazon or available locally right the servers have gotten better over the last

year um our software's gotten better over the last year so you know based on these numbers it's about um the lowest we've seen is on arm we can process a genome in 83 cents um on a current arm server um in Amazon and that takes maybe 35 or 40 minutes something like that uh while again with the the latest AMD machines it could be 13 minutes and obviously if if if you wanted to process genomes in 5 minutes um you could split it and do that if you'd like um you know and then it would cost let's just say it's double right so it's $3

instead of $150 and it's five minutes now or what whatever uh I think that's fast enough personally um and then also just it works across all instances just like I said we support arm and x86 so whatever um instance you want to run an Amazon whatever cluster you have locally and our license is just a timebase license so you can have as many software Li licenses running either locally or in whatever clouds you want um I don't care so there's no it's one simple contract with cention that I do with you and then um it's we don't have to talk for a year or two

years or three years which three year depends or one or a free trial for 15 minutes or something um and then I just want to go do there's we have lots of applications that go into um this is one though that the and long reads um according to this data we have the most accurate structural variant colar in the world um it's just a little better than this other good tool called sniffles 2 that my friend develops at his lab at Baylor um and then there are other tools um and you can also see that uh a lot of the tools degrade as coverage goes

down and coverage kind of equals money so people people do want to optimize on how much coverage they're doing um but that's in Long reads we're also we w a bunch of awards and long reads um for accuracy um it's also much more all of our tools they're faster and they also are generally require lower RAM uh which is nice some people will say oh I've accelerated some algorithm but now instead of 64 gigs of RAM I need 2 terab of ram um we actually use less Ram um and you can see in this example on the right right this is an independent Benchmark and

they're um testing Google's deep variant you know the noisy data potentially um so you know I would look I would recommend them also here's another example of just pipeline runtime so lines and this is um the structural variant calling this orange one to there so you can see our structural varing color is extremely fast and then here's an example of the pipeline where the blue um is the alignment time so that's kind of the open source best tool for pack bio I would say which is mini map 2 plus deep variant from Google and here's ours which is our sention mini map 2 which we've

accelerated uh by about 5x relative to this and then we have um a much more efficient variant calling strategy as well and it they're both very accurate we're generally a little more accurate but I'm not here to quibble over you know whatever a thousand variants or something they're both good they're both good from the accuracy side I would say um let's see the last some of my others favorite slides this is oncology we we do that great uh this is one of my favorite slides I think we have the um there's these Umi tools this is kind of like tagging of the data and we

have um developed an extremely efficient um accurate Umi tool this is lot for liquid biopsies or but really whatever umis you want we also just published a thing with Sher which I'm told is the Harvard of Germany it's a I don't know if youve i' never I hadn't heard of it before though so but I had I only knew about maybe two universities in Germany before then and I'm half German but anyway we just published a thing uh with Charity where we accelerated um something they were working on with Oxford nanapur by like 14x um with Intel um and part of that was um we

actually adapted this this tool this ctdna pipeline Umi tool to do single cell RNA seek so basically if you're doing any kind of um genomic application um you know know about cention and come reach out to us um if you're interested in a free trial now or you know anytime in the future genomics is a you know you might there's people with pipelines that they haven't updated for six years and that's fine um but when you do want to consider updating them um come talk to me um and again I want to thank ciq for having me they've been a great partner um and I

don't know I would take any questions if anyone has any uh yeah I just based on Amazon pricing because it's kind of the most um like just you can just do it right versus saying you know if I buy this Ser I don't know if that's what you're alluding to or not but like if I buy This Server locally like what is my depreciation I have all these servers and if I'm running them at 100% of the time or 80% of the time or whatever so obviously you can drive that cost down by owning yourself um I personally think the bigger issue with the cloud

rate is the long-term storage right the the genomes are big data they are you know it depends how you want to store them but they're you know compressed generally like 40 gigs um so it's really the I think the bigger issue is a where do you want to store your data long term and wherever you want to process it we're going to we're basically taking your compute cost almost to zero because especially if you're doing like other small panels or exomes right those are 8 cents or something um but yeah again on Amazon we process a a 30X short read genome um with a alignment

and variant calling for you know 80 whatever 880 some cents to you know $2.50 kind of depending on the instance you're deploying it on and that's all on demand prices so then there's also the spot Market which my general rule of thumb there is just cut it in half by 50% is it should be and because we're running this so fast um it enables the spot Market better right so because you are it's already only 35 minutes so like it's a that's a lot easier to then get that 50% savings relative to being like Oh it cost me 7 and now I'm trying to use

the spot Market to do it in 350 and then you're chunking it up and trying to do you know trying to keep things rolling whereas this is you only need the spot Market to be on for you to complete the job in you know between 15 and 40 minutes per instance per sample of 30X genome obviously much tinier if you're running you know a genome is a three billion base pairs long um whereas you you know you might be running 4,000 base pairs and then it'll be done in you know zero no no time uh anything else um what storage solution are s our fast

so we just support all the open um like storage or whatever like what we we support all the current file form so it's not like we don't use a different file format or anything um but so there's your fast Q file which you can zip and you can also do reference Bas compression you can do different things with that uh usually uh we recommend and people like to keep the cram file so all of our there's the BAM file which is UNL is the non-compressed version of a and then there's cram file which I guess maybe stands for compressed um so the cram file people

like because then you can see um kind of the alignment of kind of what the variant collar was looking at it doesn't show you there's a little in the weeds there where you could um actually have our tool process out exactly what the variant color is seeing but the cram file is storing at least what the alignment is seeing and then a lot of B informaticians like to look at that by eye so storing the cram file is definitely what um we would typically recommend um so our to be pretty it consumes the data pretty quickly so as long as um if the E I'm

not that familiar with like networking so if they is what I know is our software we want to have like the local ssds or enough speed in your storage in your IO to be like local ssds um we also have done some things where um like I mentioned we're low Ram so some of these new instances have huge have enough RAM to where we just use the ram as the local storage um if that makes sense so it just dep you know what your goal is and you can optimize like the software will fit to structure you know you have yeah thank you uh we

have snake make we don't we have there are we don't support um a nextflow implementation ourselves there are are also some kind of Open Source consortiums that are working on we were supported in the old nexow and then nextflow got upgraded in the Sak Consortium and now they're they're working on adding us back into that um and then we are I don't know if you here before we're working on a python impl implementation that then you can just put into your next flow so it'll and it will all just run for you so we are working on improving the efficiency we do have whittles um

because of our work with um Amazon omix so we have we're integrated into Amazon omix as one of their main Partners um so you can also run our tools there from a um whatever like an API call that'll just run a pipeline um but in that we've also developed whittel uh that are publicly available so there are nextflow snake make Whittle um cwl there are many um workflow manager job like description languages and our our tools are designed to work with all of them and we can help you at least get you started with all of them but as a small company we don't support

most of them as a commercial product if that makes sense all right well thank you for your attention I hope you're having a lovely super Computing conference oh and yeah uh talk to me afterwards if you'd like hello hello hello test test test test testing testing testing testing yo bless you w 3 2 1 thank you yes e probably would he did that AC C do Rocky L live weinar you guys want to talk about what you did all right we good Alex so welcome to sc23 and this is a student competition and we're here with Boston University Brown University in UMass B3 just finished the competition how did it go um definitely a little rough at the beginning uh we had some obstacles um with shipping the computer actually maybe like day one of like 400 p.m.

or something um so while everyone was set setting up their computers we're kind of just twiddling with our fingers cuz our vendor messed up but once we got everything set up um things were working smoothly um some some like errors on NPI installation but um we got everything to be working I think the end that's fantastic you want to tell everybody your name yeah I'm David where you from I'm from New York and I go to Boston University excellent thank you very much so have you had a good time so far this week or you just been focused on this the whole time yeah honestly

it's been really exhausting um some all nighters um but yeah I feel good that's over good now what are you going to do um get some food get some rest um celebrate with the team that's awesome yeah very good thank you how about you guys how do you think this week went yeah uh once we got our like actual cluster set up with us uh things were pretty much smooth from there kind of uh we'd seen most of the applications before uh as they gave it to us and we kind of just did our thing but then the mystery app did throw us for a

loop really really different content than what we were expecting and whenever you talk about the applications that they give you so explain to everybody kind of what that means yeah so I guess with the the competition there's the benchmarks and applications and the applications are sort of high performance Computing uh things one that I worked mainly on was 3D mhd it was a Magneto hydrodynamic simulation so obviously like with all the physics and everything there's a lot of computations that need to be done so it's an HPC application and and all we really need to do is compile it and then sort of set up

the parameters so that we could run it as fast as possible given our hardware and yeah there was just a lot of testing around like playing around with numbers and trying to just optimize our time that's fantastic you want to tell everybody your name where you're from yeah I'm shammer also from New York and I also go to Boston University great thank you very much so what do you think about Rocky uh I think it was pretty much what I'm used to working with and no issues uh no issues on like the OS part good uh I guess we were kind of new with sort

of server uh server level thing so there's like a BMC controller that apparently we had as a security vulnerability and that's our bad but you know you learn absolutely you going to do it again next year uh hopefully so yeah I have one more year left of school and I'd like to revisit this it's fantastic what are you thinking about doing after school uh well as for the competition uh it's only available for undergrads so no more competition but high performance Computing I guess is something that'll always be on my mind uh I am somewhat in more software interest aligned so I'll see where I

go that's great have you had a chance to walk around the floor since you've been here or is that what you get to do maybe maybe now you can go look around uh well yeah with it um running the applications you kind of after you press the Run button you kind of sit and wait so I've had some time to look around collect some merch and it's been a good week so far great what did you see out there that was interesting to you uh some of the talks they just uh a lot of certain HPC applications where it's like making Energy Efficiency and just

uh bringing computing power to just users through Cloud which is really interesting like being able to train a model even if you don't have Hardware yourself so really cool things that's fantastic well thank you very much and congratulations on finishing that's great thank you appreciate it did you guys get shirts Rocky shirts stickers

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert