Skip to content

Stable Audio 3

Stable Audio 3 is a rectified flow diffusion model for stereo music production, sound effects, and audio-conditioned inpainting.

Two tasks, and they use different packages:

Task Where Packages
Stable Audio 3 Create The music packages
Make a sound effect Sound The SFX packages

Stable Audio ships both package sets in one spec, so Miso splits the family into two cards on the Models screen. Installing the music model does not give you the effects model, and the other way round.

Folded together, the SFX packages were invisible on the card, which is why the split exists.

No lyric singing. It can produce vocal choir textures or ambient chants when prompted, but it does not sing words.

If you want vocals, use ACE-Step or one of the other song writers.

Variant For
stable-audio-3-small-music Fast music, up to 120 seconds
stable-audio-3-small-sfx Short sound effects and Foley
stable-audio-3-medium Higher fidelity music
Option Default Does
text English description: genre, instruments, mood, production texture
duration_seconds Length in seconds
num_inference_steps 8 Flow diffusion steps
guidance_scale 1.0 CFG scale
negative_prompt Unwanted instruments or qualities
sampler pingpong Also euler and dpmpp-2m
audio_input_kind init_audio or inpaint_audio
init_noise_level Strength for audio-to-audio, 0.0 to 1.0
inpaint_mask_start_seconds, inpaint_mask_end_seconds Region boundaries for inpainting

Describe the sound, not the song structure. There are no sections to mark.

For music:

chill lo-fi hip hop beat with dusty vinyl crackle, warm rhodes keys,
mellow sub bass, lazy swing drums

For effects, describe the event:

heavy wooden door closing in a stone hallway

6 to 8 GB for the small variants, around 12 GB for medium.