Can I self-host Stable Audio 3.0 and fine-tune it on my own library?

Yes. Stable Audio 3.0 Small, Small-SFX, and Medium download as open weights and run on your own hardware, and the repository ships LoRA fine-tuning so you can adapt a model to your own audio library. Training needs roughly 20 to 50 described clips and a CUDA GPU or Apple Silicon Mac. Large is API only.

Key takeaways

  • Self-hosting and fine-tuning are separate decisions, and a studio can do the first without ever doing the second.

  • A LoRA fine-tune produces a .safetensors file of roughly 50 to 200 MB that loads on top of an unmodified base checkpoint.

  • Stability AI's documented starting point for a training set is around 20 to 50 audio clips with matching text descriptions, and more material generally helps.

  • Stable Audio 3.0 Large sits outside the open-weight release and is not supported by the local training code.

Self-hosting and fine-tuning are two separate decisions

Running the weights locally and adapting them to your catalog solve different problems. Self-hosting answers where generation happens, which decides your cost curve and whether prompts and audio cross a network boundary. Fine-tuning answers what the model sounds like when it generates.

Teams often arrive at the second question through the first. A post house that has already pulled Stable Audio 3.0 Medium onto a workstation for confidentiality reasons then has a trained-on-your-own-material option sitting right there, because the same repository that serves inference also ships the training script. Our comparison of the self-hosted and API deployment routes for generative audio covers that decision in full.

Which Stable Audio 3.0 models accept a fine-tune

Three of the four models in the family can be fine-tuned on your own hardware, and they are the same three released as open weights in the Stable Audio 3.0 collection on Hugging Face.

Model Parameters Inference hardware Max length Training VRAM Fine-tunable locally
Small-Music 433M CPU 120s ~2.5 GB, or ~2 GB reduced Yes
Small-SFX 433M CPU 120s ~2.5 GB, or ~2 GB reduced Yes
Medium 1.4B GPU (CUDA) 380s ~6.5 GB, or ~5.5 GB reduced Yes
Large 2.7B API 380s Not applicable No

Figures from the Stable Audio 3.0 LoRA training documentation and the Stable Audio 3.0 repository README, retrieved October 2026. The reduced column assumes --base_precision bf16 with the lora-xs adapter.

Inference and training have different hardware floors, and the gap catches people out. Small and Small-SFX generate on a CPU with no GPU at all, which is the headline figure most comparisons quote. Training a LoRA on either of them still calls for a CUDA GPU under the main training script.

What a LoRA checkpoint actually is

Fine-tuning a Stable Audio 3.0 model does not rewrite the model. During training both the diffusion model and the conditioner are frozen, with gradients switched off, and the optimizer receives only the small low-rank matrices injected into the Linear and Conv1d layers. When the run saves, only those tensors are written out, with the adapter configuration embedded in the file metadata.

What lands on disk is a .safetensors file in the region of 50 to 200 MB, loaded on top of any base checkpoint at generation time. Your base weights stay byte-identical. Several consequences follow for a team managing a catalog of house styles: a LoRA swaps without a reinstall, several can be loaded at once, and a bad training run costs you a file rather than a model. The same local setup supports generating stems and inpainting regions of a track, covered separately.

What you need before training on your own catalog

Audio with words attached is the real prerequisite. Stability AI's documentation puts the floor at roughly 20 to 50 clips paired with matching text descriptions, and notes that more is better. A sample library with no metadata needs a captioning pass before any of it reaches the training script, and the quality of those captions shapes what the finished LoRA responds to.

The software side is one install flag, uv sync --extra lora, and one script. Stability's recommended starting configuration trains from the medium-base checkpoint at rank 16 with the DoRA-rows adapter for 1,000 steps, and the documentation is explicit that these are reasonable defaults rather than optimal settings, since behavior varies with dataset size, style and hardware.

Rights in the training material are a question for your own legal counsel. A back catalog you recorded, a library you licensed and a folder of material someone else owns may sit in very different positions, and this page does not attempt to tell you which is which. The Stability AI Terms and Community License require compliance with applicable laws and regulations, and prohibit using Stable Audio models to infringe, misappropriate, or violate intellectual property or other legal rights.

Fitting training into the VRAM you have

Medium trains inside a mid-range consumer card. Standard settings peak around 6.5 GB, and two flags bring that down to roughly 5.5 GB: --adapter_type lora-xs, which trains a much smaller set of parameters, and --base_precision bf16, which halves the memory held by the frozen base weights while keeping the LoRA parameters in fp32. Stability describes the quality difference as usually small.

Apple Silicon has its own route. The repository ships LoRA training on Apple Silicon written in pure MLX with no PyTorch dependency, so a composer on an M-series Mac can train without renting a GPU instance. Small trains at roughly 2.5 GB standard and around 2 GB with the same two flags.

One practical setting is worth knowing before a first run on a small dataset. Passing --exclude seconds_total keeps the adapter off the duration conditioner, which the documentation recommends to prevent conditioner hijacking when there is not much material to learn from.

Choosing an adapter type

Eight adapter variants trade expressiveness against parameter count, and the default is a sensible place to start.

Adapter Trainable parameters per layer When to use it
lora rank * (fan_in + fan_out) General-purpose fine-tuning
dora-rows rank * (fan_in + fan_out) + fan_out Documented default, tends to generalize well
dora-cols rank * (fan_in + fan_out) + fan_in Per-input-feature magnitude instead
bora rank * (fan_in + fan_out) + fan_in + fan_out Independent row and column magnitudes
lora-xs rank² Maximum parameter efficiency and lowest VRAM
dora-rows-xs, dora-cols-xs, bora-xs rank² plus magnitude vectors Frozen SVD bases with DoRA or BoRA magnitude

Rank sets how much the adapter can express, at 16 by default. The -xs family freezes the low-rank bases from an SVD of the original weight and trains only a small core matrix, plus magnitude vectors in the DoRA and BoRA hybrids, which at rank 8 means 64 trainable parameters per layer against thousands for standard LoRA. Computing that SVD at startup is slow, and --svd_bases_path lets you compute it once and reuse it across runs.

Running more than one LoRA at a time

Stacking is built in rather than bolted on. Multiple checkpoints load together through PyTorch's parametrization API, each gets its own index, and each carries independent controls: a strength value on the diffusion backbone, a separate strength on the text conditioner, a noise interval, and a layer filter.

The interval control is the one with no obvious equivalent elsewhere. Setting a LoRA active only in the late, low-noise part of sampling confines its influence to fine detail, while the early steps shape global structure. A house-style LoRA at full strength across the whole sampling run can therefore sit alongside a texture LoRA at half strength in the final steps only.

Merging is the other option. Weighted LoRA deltas can be baked directly into the base weights, with an application weight per checkpoint, after which the parametrizations switch off and the effect is permanent in that copy of the model.

What the license says about a fine-tune you ship

Ownership of your outputs runs the same whether you fine-tune or not. Under the Stability AI Community License, as between you and Stability AI, you own outputs the models generate,, subject to applicable law and the Acceptable Use Policy. The license is free to organizations generating under $1M in annual revenue, counted in aggregate across you and your affiliates and regardless of where the revenue comes from.

A LoRA you train or distribute brings the license's distribution conditions with it: providing a copy of the license to the recipient, retaining the attribution notice in a NOTICE text file, and displaying "Powered by Stability AI" prominently in a related surface such as a user interface, website or product documentation. Read the license text before you publish a checkpoint, and take advice on anything ambiguous for your situation. Stable Audio LoRAs are subject to the same Community License and Acceptable Use Policy as the base models.

One restriction bears directly on fine-tuning plans. The license prohibits using the models, derivative works or their outputs to create or improve any foundational generative AI model other than Stability's own, so harvesting generations as training data for a separate house model is prohibited. Teams weighing that against other vendors can see how Stable Audio compares with other generative audio platforms on licensing and self-hosting.

Where a fine-tune will not help

Fine-tuning moves style. Several constraints sit below that level and survive any amount of training on your own material.

Key and pitch are not conditionable. Tempo goes in the prompt, the released models expose no key conditioning, and generation is non-deterministic unless a seed is set, so separately generated parts still need the method of deriving adaptive layers from a single bed. Vocals are out of scope across the family: the models produce instrumental music and sound effects, with vocal textures appearing occasionally and no intelligible singing. Output arrives at 44.1 kHz stereo, so a 48 kHz pipeline still needs its resample step. Large remains unavailable for local training.

Frequently asked questions about self-hosting and fine-tuning Stable Audio 3.0

Can I fine-tune Stable Audio 3.0 on my own sample library? Yes, assuming you have sufficient IP rights to the samples in question. The open-weight models ship with LoRA fine-tuning documentation and a training script, so Small-Music, Small-SFX and Medium can be adapted to a specific style or catalog. Training needs audio files paired with text descriptions, starting from around 20 to 50 clips, and produces an adapter file rather than a modified base model.

Do I need a GPU to fine-tune Stable Audio 3.0? Yes, for the main training script, which calls for a CUDA GPU. Medium peaks at roughly 6.5 GB of VRAM during training, or about 5.5 GB with the lora-xs adapter and bf16 base precision. CPU inference on the Small models does not extend to CPU training.

Can I train a LoRA on a Mac? Yes, through the repository's MLX route, which implements LoRA training in pure MLX with no PyTorch dependency for Apple Silicon machines. Training speed depends on the machine and the dataset. The standard CUDA path remains the documented default for everything else.

How many audio files do I need to train a LoRA? Around 20 to 50 clips is the documented minimum, with better results expected from more material. Every clip needs a matching text description. A small set also makes the --exclude seconds_total flag worth using, which Stability recommends to stop the duration conditioner dominating the adapter.

Do I have to caption every clip in my training set? Yes. The training data is audio paired with text, and the captions teach the model which words should summon which sound. An uncaptioned sample library needs a description pass first, and the vocabulary you choose there becomes the vocabulary your finished LoRA responds to.

Does fine-tuning change the base model? No. Base weights are frozen throughout training and only the injected low-rank parameters are updated and saved. The result is a separate .safetensors file of roughly 50 to 200 MB that loads on top of an unchanged checkpoint, so a LoRA can be swapped, stacked, turned down or removed without reinstalling anything.

Can I fine-tune Stable Audio 3.0 Large? No, not through the open-weight release. Large is the 2.7B-parameter model, reachable through the Stability AI API, and the repository's model table lists it as unsupported by the local code. Teams needing the family's highest musicality on their own terms should talk to Stability AI directly about options.

Do I own the model I fine-tune, and the audio it generates? Fine-tunes are subject to the same Community License as the core models. Under the Community License, as between you and Stability AI, you own outputs you generate from the Core Models and from fine-tunes, subject to applicable law and the Acceptable Use Policy. The license applies free to organizations under $1M in annual revenue, and above that figure an Enterprise license applies.

Can I distribute or sell a LoRA I trained? LoRAs you train are subject to the same Community License as the core models, and distribution may carry conditions including supplying the license to recipients, retaining the attribution notice in a NOTICE file, and displaying "Powered by Stability AI". Check the license text and your own counsel before publishing, particularly on anything involving the material you trained on.

Can I use Stable Audio 3.0 output to train my own music model? No. The license prohibits using the models, derivative works or their outputs to create or improve any foundational generative AI model other than Stability's. A platform planning to collect user generations as training data for a model of its own needs a different arrangement with Stability AI.

Next
Next

Is AI-generated music copyright-safe?