Stable Audio 3.0 Prompt Guide

Stable Audio 3.0 is built for the iterative nature of audio production. So our models give you several different ways to work with audio:

  • Music composition

  • Samples and sound effects

  • Audio-to-audio

  • Inpainting and continuation

You can generate full songs, individual instruments and stems, samples and sound effects, or use existing audio as the starting point for a new idea. Here are a few ways to approach each one, with examples you can try in your own work.

All Stable Audio models are trained on fully licensed data. You own your outputs, and can distribute and commercialize them freely – and you have as many downloads as you’d like. We can’t wait to see what you make.

1. Music composition

Start with the basics of what you want to hear:

  • Genre: What style of music are you making?

  • Instruments: What is playing, and how should it sound?

  • Mood and energy: What feeling should the track create?

  • BPM: What tempo does the track need?

A useful order is: start with the genre, name the main instruments and rhythmic elements, then describe the mood, arrangement, and production character. Then add details about performance, texture, arrangement, or production when they're important.

Note: Stable Audio 3.0 is designed primarily for instrumental work. It can sometimes create non-lexical vocal-like textures, but it is not intended to generate intelligible vocals.

Try style references

It’s helpful to reference an era, location, recording style, or musical context to describe what you're after. For example:

  • 90s garage rock instrumental with a grunge influence, poppy distorted guitars, frantic drums, and tube-distorted bass

90s garage rock
  • Detroit-influenced techno with a stripped-back arrangement, metallic percussion, deep sub bass, and raw analog synths 130 BPM

Detroit-influenced techno

You can also describe the context in which you'd hear the music:

  • An ambient electronic instrumental for the final scene of a slow-burning sci-fi film

An ambient electronic instrumental for the final scene of a slow-burning sci-fi film
  • A euphoric house track that feels like the last song of a long night at a club

A euphoric house track that feels like the last song of a long night at a club

Experiment with generating full mixes vs. instruments

The models allow for generating different track types, including full mixes and stems.

If you need an individual part, like a bass line, guitar texture, or percussion layer, you can try prompting for a stem instead.

Use tags to specify what kind of track you want

Start with TrackType: Instrument, then say what instrument you want and how it should be played, recorded, or processed.

TrackType: Instrument is an optional tag. Tags are simple labels that help tell the model what kind of result you want. For example:

  • TrackType: Music signals a full musical track.

  • TrackType: Instrument signals an isolated instrument or stem.

  • TrackType: SFX signals a sound effect or one-shot.

  • Format: Duo signals that you want two instruments together.

  • Genre: Funk adds a direct genre label.

You do not need tags in every prompt. Use them when you want to give Stable Audio a clearer signal about the format of the output.

Music example: The last track in the DJ set

TrackType: Music, a triumphant and stylish UK bass-flavoured tech-house tune that evokes feelings of the last tune played in a DJ set. The pumping four-to-the-floor kick is supported by an 808 bass that is syncopated. There are gliding emotional synth leads that build sections to their climax. Playful stabs and chops support the rhythm of the drums in sections. There is a beautiful gospel house piano that plays in the drop, giving the track a euphoric feeling

A triumphant and stylish UK bass-flavoured tech-house tune

Music example: A 70s TV-theme hip-hop groove

A funky hip hop instrumental with live recorded instrumentation that has the vibe of a 70's TV show theme. The rhythm guitar strums lightly in the background, the lead guitar occasionally delivers jagged flanger-effected chords and phrases, and abstract sounds punctuate the beat with character. A close-mic'd electric piano plays hard with nostalgic supporting chords, flute glides over the beat adding an exciting texture, and the drums are full of swing and old school flavour. The production has a warm, textured sound associated with analogue gear and tape

A funky hip hop instrumental with live recorded instrumentation that has the vibe of a 70s TV show theme

Stem example: Sombre finger-picked acoustic guitar

TrackType: Instrument, a sombre solo acoustic guitar track with cavernous reverb and delicate finger picking

A sombre solo acoustic guitar track

Stem example: Dynamic Latin percussion duo

TrackType: Instrument, Format: Duo, a dynamic Latin drums and percussion track

A dynamic Latin drums and percussion track

Stem example: Vintage-studio electric guitar lead

TrackType: Instrument, an epic solo electric guitar track, live-recorded in a vintage studio

An epic solo electric guitar track

For stems, you can add details like technique, room sound, mic perspective, effects, performance energy, or recording character.

2. Samples and sound effects

For sound effects and samples, describe the sound as precisely as possible. Focus on:

  • The source: What object, instrument, or synth is making the sound?

  • The action: How is the sound being triggered, and how long does it last? (e.g., slamming shut with fast decay, or a massive suck-back followed by a supersonic crack)

  • The production/characteristics: Where is the mic placed, and what's the room character? How is the sound processed? (e.g., recorded in a dead room, captured with a vintage ribbon mic, or processed with a dark wooden reverb)

TrackType: SFX is another optional tag. It tells Stable Audio that you want a sound effect rather than music. This can help when you are making hits, risers, foley, or transitions.

Example: Distorted wooden-drawer slam

A blunt, powerful "thud" made by slamming a wooden desk drawer shut. It has a pronounced low-mid body, making it feel heavy, and is given a touch of analog distortion for aggressive character

A blunt, powerful thud made by slamming a wooden desk drawer shut

Example: Tape-stop sub-bass effect

A massive sub-bass note that mimics a vinyl record or tape machine being turned off. The pitch and speed drop simultaneously, causing the high-end harmonics to "smear" and thicken as the sound grinds to a halt at a sub-sonic frequency

A massive sub-bass note that mimics a vinyl record or tape machine being turned off

Example: Fast-decay hi-hats

A classic, natural recording of 14-inch hi-hats played tightly closed. Zero ring and a very fast decay.

A classic, natural recording of 14-inch hi-hats played tightly closed

3. Audio-to-audio: Generate variations of existing audio

Audio-to-audio lets you start with an existing piece of audio. You upload or select a source file, then write a prompt that explains what you want Stable Audio to turn it into.

You can give three inputs for audio-to-audio:

Input What it is
Seed audio The file you start with; for example, a violin phrase, beatbox recording, or previous Stable Audio generation
Prompt (optional) The new direction you want Stable Audio to take the audio; for example, “heavy metal guitar” or “lo-fi hip-hop beat”
Noise level (optional) How far Stable Audio should move away from the input

Use noise level to tell the model how much to change

Use init_noise_level to control how much the original audio is modified. Lower values usually preserve more of the source. Higher values create a more substantial variation.

Note: “Noise level” in this context does not refer to volume or decibels. In a diffusion model, noise is the random signal used during generation.

A lower noise level keeps more of the original performance, including details such as rhythm and melody. A higher level changes the input more substantially, allowing Stable Audio to take on more of the new prompt’s style.

In general, keep the value here between 0.3 and 0.9, with 0.3 being almost no change and 0.9 significantly modifying your original audio.

Timbre transfer

Change the instrument while generally preserving the notes.

Prompt examples:

  • Input: violin_solo.wav | Prompt: TrackType: Instrument, electric guitar | Noise level: 0.82

Violin solo
Electric guitar
  • Input: beatbox.wav | Prompt: Lofi hip hop beat, chillhop | Noise level: 0.5

Beatbox
Lofi hip hop beat

Style transfer

Shift the overall style of a track to create variations.

Prompt example:

  • Input: trance.wav | Noise level: 0.6

  • Prompt: (none)

Trance
Generated trance

4. Inpainting and continuation

Inpainting and continuation let you keep the parts of a track that are working while generating new audio where you need it.

Inpainting: Re-create a specific section

Use inpainting when part of a track needs a new take. Select a section of audio, write a prompt, and Stable Audio regenerates that section while preserving the rest. Try using it to:

  • Replace an awkward transition.

  • Rework a section that is too busy.

  • Change a passage without losing the rest of the arrangement.

  • Introduce a different instrument or energy level later in a track.

The audio around the section you are changing matters a lot. If you mask only a few seconds, Stable Audio has a lot of context and may make only a subtle adjustment. If you want a bigger change, regenerate a larger portion first, then refine it with a smaller selection.

Example: Bring a guitar solo into a classical piece

Input audio Prompt Section to replace
classical.wav — 107 seconds long An electric guitar solo live-recorded in a vintage studio with an orchestra 30 seconds to 107 seconds
An electric guitar solo live-recorded in a vintage studio with an orchestra

This keeps the first 30 seconds as context, then regenerates the remaining 77 seconds in a direction that introduces the guitar solo.

Note: If it's too far from the input audio, the prompt does not work. Also, if the section to replace is too small compared to the rest of the audio, it will likely ignore the prompt. 

Continuation: Extend the length

Use continuation when you want to make a piece longer. Start at the end of the existing audio, choose how long you want the continuation to run, and generate the next section.

Your prompt should fit the audio around it. If you are extending a progressive house track, a prompt for a build and drop is more likely to create a smooth result than a completely unrelated sound.

Tip: Stable Audio 3.0 enables variable-length generation, which means you can output the exact length you want, up to about six minutes for Stable Audio 3.0 Large.

Example: Extend a bird-call recording

Input audio Prompt New length
bird_call.wav — 3 seconds long No prompt needed Extend from 3 seconds to 100 seconds
Birdsong (no prompt)

This tells Stable Audio to continue the existing audio from its endpoint until the file reaches 100 seconds total.

bird_call.wav—3 seconds long TrackType: Music, Lofi hiphop chill beat with melodic synth stabs and pitch-shifted bird song that builds Extend from 3 seconds to 55 seconds
Birdsong continuation

5. Try using the plugin and web app

The new Stable Audio plugin and updated web app are now in beta, giving you two ways to bring Stable Audio 3.0 into an active production workflow.

In the Stable Audio plugin

Use the Stable Audio plugin when you want to generate directly inside your favorite DAW. Prompt for material with a clear role in your arrangement: a bass part, percussion layer, guitar texture, transition, sound-design element, or alternate section.

Then treat what you generate like any other audio in the session:

  • Arrange and edit it against the rest of the track.

  • Process it with your effects chain.

  • Layer it with recorded instruments or samples.

  • Chop it into new phrases or textures.

  • Print the take and keep building.

In the Stable Audio web app

Use the enhanced web app to go beyond the initial generation to give iterative direction, try variations, mix and edit, work with multiple tracks, and extend the length. 

You can also make tape and splice edits, such as reverse, stutter, freeze, cut, copy, and paste, without generating new audio. For a full walkthrough of mixing, editing, and exporting in the web app, read our guide to making your first mix with Stable Audio.

There’s no perfect prompt, only your process

A good prompt gives Stable Audio a clear creative direction, but there is no single perfect way to write one. Stable Audio will not produce the same result every time. That variation can help you find unexpected ideas. 

Start with a prompt, and then follow your own creative process to experiment and keep shaping the sound until you have something worth taking into the next stage of your production.

Previous
Previous

Make your first mix with Stable Audio

Next
Next

Six ways to create smarter, not harder in Brand Studio