Fuzzball 4.3 | Serve a model in one command

Fuzzball v4.3: your GPUs serve a model in one command

Contributors

Jonathon Anderson, Principal Engineer

Running a large language model on your own GPUs should be easy. Without Fuzzball, assembling all the pieces and wiring it all together and making it accessible is complex. Now Fuzzball puts a large language model on your GPUs, served and ready for traffic. It also ships with the serving stack already assembled: the catalog entry, the replica pool, and with an OpenAI-compatible endpoint.

Fuzzball v4.3 makes it easy to self-host AI models on your infrastructure and to connect those models to AI agents running inside the cluster or locally.

Fuzzball v4.3 also continues our support for environments that don't fit the traditional HPC structure, with expanded support for node-local storage in both local and disaggregated environments.

Turn-key models and agents

There's a new vLLM catalog entry that serves a HuggingFace model from a pool of replicas on NVIDIA or AMD GPUs:

fuzzball workflow catalog start vLLM --values Model=hf://openai/gpt-oss-20b

That's all you need to start hosting a model, ready for inferencing. Point any OpenAI-compatible client or agent at it and start incorporating locally-hosted AI into your development or process.

The vLLM catalog entry has a number of parameters to tune the generated workflow for the model you want to host; but there are eleven preset catalog entries that have already had the GPU count, context size, and vLLM arguments worked out and validated per model:

  • GPT-OSS 20B and 120B
  • Ministral 3 8B and 14B
  • Qwen3-Coder 30B and Next 80B
  • Nemotron 3 Nano and Super
  • Mistral Medium 3.5
  • Llama 4 Scout
  • Gemma 4 31B

Starting one of those is just its name; the model, the GPU fit, and the serving arguments are already decided:

fuzzball workflow catalog start "Nemotron 3 Nano"

Invoked bare like that, the CLI walks you through the entry's settings with the preset's value filled in for each, so pressing Enter through the list gets you the validated configuration. Pass --values instead and it skips the prompts, with anything you don't name left at its preset:

fuzzball workflow catalog start "Nemotron 3 Nano" --values Volume=my-models,MaxReplicas=8

That keeps the downloaded checkpoint on a persistent volume rather than pulling it again on every start, and lets the pool grow to eight replicas under load.

Each model is packed with a built-in LiteLLM Model Gateway that puts a single OpenAI-compatible endpoint in front of the workflow. This allows the workflow to scale up and down based on demand, keeping a single endpoint for clients and agents. Or disable the workflow's gateway and run a single gateway for all your models: this lets you have a single endpoint that can discover every model-serving workflow you have access to.

Fuzzball also comes with two coding agents in the catalog, OpenCode and hermes-agent, which run in the cluster and can detect your model gateways automatically.

Finally, a rag catalog entry runs a document retrieval service built on haiku.rag. It can ingest PDF, DOCX, PPTX, HTML or plain text documents and expose hybrid vector and full-text search exposed as MCP tools on a workflow endpoint.

Autoscaled pools and pool endpoints

To support the catalog work above, service endpoints now work against an autoscaled pool, not just a single-instance service.

A pool's endpoint keeps one URL for the life of the workflow and forwards each request to whichever replica is ready at that moment. Scale up and new replicas join the rotation as they pass readiness; scale down and one leaves when its drain period starts, after in-flight requests finish. Callers never learn how many replicas exist.

Pools can also idle at zero. A pool with replicas.min: 0 keeps its endpoint while nothing is running, and a request that arrives gets a 503 with Retry-After: 120 and starts a replica. Retry after the cold start and you're served. The request has to be authenticated, so a scope: public endpoint will never wake its pool, which is deliberate, because otherwise anyone on the internet could turn your GPUs on. A burst of requests starts one replica, not one per request, and every wake shows up in fuzzball workflow events.

If you'd rather do your own load balancing, per-replica: true gives every live replica its own endpoint alongside the pool's.

Endpoints can now carry arbitrary key-value annotations that make endpoints discoverable to clients. (This is how the LiteLLM gateway finds the pools it should front.) And a service can combine multinode with autoscaler.replicas, so each replica is a gang-scheduled group of ranks serving from rank 0.

Expanded support for node-local volumes

Fuzzball can be configured to make node-local storage available as local volumes. In the past, this has only supported ephemeral volumes; but Fuzzball v4.3 adds support for persistent local volumes, and improves handling of ephemeral local volumes as well.

Every volume on a local provisioner (declaring local: true, only supported with the hostpath provisioner today) has exactly one holding node, which is recorded during fuzzball volume create, or specified with fuzzball volume create --node. Fuzzball places every stage that mounts the volume on the volume's holding node.

An ephemeral local volume is created at the start of the workflow, and later workflow stages that mount the volume are placed on the volume's holding node as well.

Placement is capacity-aware: nodes without room for the requested size are skipped, and targeting a node by hand that has insufficient space fails.

Local volumes do not currently support concurrent multi-node workflow stages. This means multinode (e.g., MPI) jobs, concurrent task arrays, and multi-node scale-up services need a shared volume to work today.

That last case is worth noticing if the pool endpoints above just caught your eye: autoscaled serving and node-local volumes are mutually exclusive by construction. If you point the vLLM entry's Volume at a persistent volume, it wants a shared provisioner -- the entry's MaxReplicas defaults to 4.

Before you upgrade

Three changes need attention:

  • Catalog repository URIs are now restricted to HTTPS and publicly routable hosts by default. Repositories that don't qualify will stop syncing and need re-adding over HTTPS. Operator-managed deployments can relax this with spec.fuzzball.workflowCatalog.restrictions.allowPrivateAddresses.
  • Catalog entry IDs change on the first sync after upgrade, because the ID now derives from the catalog source as well as the entry's own id. Saved links and scripts referring to entries by UUID need updating.
  • Ingress routes on the kong and nginx classes now require TLS and redirect plaintext callers with a 308. Other classes are unchanged.

v4.3.0 also includes all fixes already released through v4.2.3.

Full release notes, including the complete fix list and the security fixes in this release, are at docs.ciq.com/fuzzball.

Subscribe to our newsletter

Related posts

Making computing serve the science: Wolfgang Resch’s journey to Fuzzball

Making computing serve the science: Wolfgang Resch’s journey to Fuzzball

Your GPUs are throttled by the kernel you inherited

Your GPUs are throttled by the kernel you inherited

Fuzzball 4.2: AI agents that drive Fuzzball, and workflows that submit workflows

Fuzzball 4.2: AI agents that drive Fuzzball, and workflows that submit workflows

Rebuild the system, solve the problem: Howard Van Der Wal's work with NASA

Rebuild the system, solve the problem: Howard Van Der Wal's work with NASA

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

9

Enterprise products

Spanning the kernel to the orchestrator

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert