Stability AI shipped Stable Audio 3 on May 20, 2026, and the first question everyone asks is whether it retires version 2.5. It doesn't. The two now do different jobs. The new family brings open weights you can run on your own hardware, tracks that stretch past six minutes, and editing tools like inpainting. Stable Audio 2.5 stays the faster, zero-setup way to get a clean instrumental in under a minute. Below you get the specs side by side, the license and API cost in plain numbers, prompt tips that hold up on both, and a walkthrough for running 2.5 on PicassoIA today.

Stable Audio 3 at a Glance
Stable Audio 3 is not one model. It is a family of four built on a new architecture: a semantic-acoustic autoencoder that squeezes audio roughly 4096 times and lets the model generate any length down to the second. Three of the four ship as open weights on Hugging Face. The biggest one is API only.
The Four Models
| Model | Parameters | Max length | Access |
|---|
| Small SFX | 459 million | About 2 minutes | Open weights |
| Small (music) | 459 million | About 2 minutes | Open weights |
| Medium | 1.4 billion | 6 min 20 sec | Open weights |
| Large | 2.7 billion | Over 6 minutes | API or enterprise hosting |
All four take a text prompt. The family also handles audio-to-audio, which reshapes an existing clip while keeping its timing, and inpainting, which regenerates one stretch of a track while the rest stays untouched. Inpainting comes in three flavors: a single segment, several segments at once, and causal continuation that extends a track from where it ends. Small and Medium accept LoRA fine-tuning, and ComfyUI support arrived on launch day.

💡 Worth knowing: none of the four models sing. There are no vocals and no lyrics anywhere in the family. If a track needs a voice, you add it with a separate model, and the last sections show how.
What Changed Since 2.5
Stable Audio 2.5 tops out at 190 seconds per generation, about three minutes. Stable Audio 3 Medium doubles that, at 6 minutes 20 seconds. Training also got cleaner: the dataset is reported at about 1.28 million licensed recordings, roughly 806,000 tracks from AudioSparx and about 473,000 Creative Commons files from Freesound. Copyrighted material was screened out, which matters if you plan to publish what you generate.
Specs Side by Side
Here is the short version in one table. Prices and lengths come from public launch reports and the model page on PicassoIA, so check the official pricing page before you budget a large project.
| Feature | Stable Audio 2.5 | Stable Audio 3 |
|---|
| Max length | 190 seconds | 6 min 20 sec (Medium) |
| Weights | Hosted only | Open for Small SFX, Small, Medium |
| Run on your machine | No | Small on CPU, Medium on GPU |
| Vocals | Instrumental focus | None |
| API price per run | About $0.20 (20 credits) | $0.26 (26 credits) |
| Editing tools | Text prompt on PicassoIA | Text, audio-to-audio, inpainting, continuation |
| Fine-tuning | No | LoRA on Small and Medium |
| Setup | None, runs in a browser | Download weights or call the API |

Where 2.5 Still Wins
Speed and simplicity. The example runs on the Stable Audio 2.5 page produced 90-second tracks in about six seconds each at the default 8 steps. You open a page, type a prompt, press generate. No Python environment, no GPU drivers, no license form.
Price per run. At roughly 20 credits against 26, 2.5 is the cheaper API call when a short loop is all you need.
Predictable iteration. Reuse a seed and you get the same audio back, so you can change one word in the prompt and hear exactly what that word did.
Where 3 Pulls Ahead
Length. A six-minute Medium track is enough for a short film cue or a podcast bed without stitching clips together.
Control. Inpainting lets you fix bar 40 without rerolling the whole song. Continuation extends a track from where it stops. Audio-to-audio reshapes a clip you already have.
Ownership of the workflow. Open weights mean you can run offline, keep unreleased material off third-party servers, and fine-tune with LoRA on your own sound library. A game studio, for example, can train Small SFX on its own foley recordings so every footstep matches the house style.
Honest quality check. One review scored the family 7.8 out of 10 and named the strengths: ambient drones, foley, sound effects, instrumental loops and atmospheric beds. The same review noted that the music can feel correct without feeling interesting, and that rhythmic genres tend to flatten. Treat 3 as a strong tool for beds and effects, not a replacement for a band.
Open Weights: What You Can Run
Open weights change who can use this. Before, audio models at this quality lived behind an API. Now a laptop and a download are enough for the two small models.

Small Models on a CPU
Small SFX and Small run on a CPU with no discrete graphics card. Launch reports put generation times under half a second on an H200 server GPU, and the point of the design is full music composition on-device, offline, without the short sample limits older open models had. Use Small SFX for impacts, ambience, footsteps and whooshes. Use Small for loops and short instrumental pieces up to two minutes.
Both small models also take LoRA fine-tuning. A LoRA is a small add-on trained on your own clips, such as your band's drum recordings or a game's foley library, and it nudges every generation toward that character. Because the add-on is tiny, the job is much lighter than retraining a full model.
Medium is the heavier option at 1.4 billion parameters. Setup notes list a CUDA GPU with Flash Attention 2 as the supported path, and the reported speed on an H200 is under two seconds per track. One reviewer measured a few seconds on a MacBook Pro M4, so check the repository for your exact hardware before you plan around it.
License Rules in Plain Words
The open models use the Stability AI Community License.
- Under $1 million in annual revenue: free to download, run, fine-tune and use commercially once you register for the Community License. You own your outputs, to the extent the law allows.
- Over $1 million: you need an Enterprise License, which also adds legal indemnification.
- Large: never downloadable. It is reachable through the Stability API, through fal.ai, or through enterprise self-hosting.

💡 Read before you ship: "you own the output" is not the same as "the output is protected by copyright." In many countries purely machine-made audio gets weaker protection than human-made work. If a track is central to your brand, layer your own performance or edits on top.
Using the API
Not everyone wants to install anything. The hosted route is the shortest path from idea to file.

Cost per Generation
Stability prices in credits, and one credit equals one cent. A Stable Audio 3 generation costs 26 credits, or $0.26. Stable Audio 2.5 is reported at 20 credits, about $0.20. A small gotcha: the free starter balance of 25 credits is one credit short of a single Stable Audio 3 run, so budget a small top-up even for testing.
| Volume | At $0.26 per run | At $0.20 per run |
|---|
| 10 tracks | $2.60 | $2.00 |
| 100 tracks | $26.00 | $20.00 |
| 1,000 tracks | $260.00 | $200.00 |
When the API Beats Local
Pick the API when you need the Large model, because it is the only way to reach it short of an enterprise contract. Pick it too when you have no GPU, when your volume is bursty, or when a product needs audio on demand without you maintaining servers. Pick local weights when you generate thousands of clips, work offline, or need to fine-tune.
A typical hosted workflow looks like this:
- Write the prompt with genre, instruments, mood and tempo.
- Set the length you need, in whole seconds.
- Send the request and wait for the audio file.
- Listen, then change one variable and run it again.
- Trim, fade and loop the winner in your editor.
The hosted Stable Audio 2.5 schema is small: a text prompt, a duration, a step count, a CFG scale and an optional seed. Those are the same dials you will meet in most Stable Audio endpoints, so what you practice on 2.5 carries over.
Prompt Tips That Work
Prompts for both versions reward specificity. A vague request like "chill music" returns generic wallpaper. A prompt that names genre, instruments, mood and tempo returns something you can use.

The Five Part Formula
Write prompts in this order, separated by commas:
- Genre: boom bap hip hop, ambient house, post-rock, cinematic synthwave
- Instruments: solemn piano, SP-1200 drums, 808 kick, strings, sine wave bass
- Mood: peaceful, euphoric, melancholic, tense
- Tempo: a number, such as 90 BPM or 125 BPM
- Texture or scale: warm, dusty, wide, minor tonality
| Weak prompt | Strong prompt |
|---|
| Relaxing music | Soulful boom bap hip hop instrumental, solemn piano, SP-1200 drums, sine wave bass, peaceful, 90 BPM |
| Dance track | Ambient house, 808 drum machine, claps, shaker, synth bass, euphoric, 125 BPM |
| Epic movie music | Post-rock, guitars, drum kit, bass, strings, uplifting, flowing, sentimental, 125 BPM |
| Door sound | Heavy wooden door slam in a stone hall, long echo tail, no music |
For longer tracks, add timing cues to the end of the prompt: "soft pads for the first 20 seconds, drums enter after 20 seconds, slow fade at the end." Long Stable Audio 3 tracks benefit most from this, since a six-minute piece needs an arc. On 2.5, keep timing hints short and judge the result by ear.
For sound effects, swap the music formula for three details: the object, the surface it touches, and the space around it. "Boots on wet gravel, empty quarry, distant wind" usually beats "footsteps", because the model can place the sound in a room.
💡 Shortcut: two or three words can also work. Short genre prompts such as "Lo-fi Funk" or "Cinematic Synthwave" gave usable 90-second tracks in the 2.5 examples. Add detail only when the first result drifts.
Mistakes That Waste Credits
- Asking for vocals. Stable Audio 3 has no vocals or lyrics, and 2.5 is aimed at instrumentals and sound design. Use a song model for singing.
- Piling on adjectives. Ten mood words cancel each other. Pick two.
- Skipping the tempo. Without a BPM the rhythm wanders. A number anchors it.
- Rerolling blindly. Change one thing at a time and keep the seed fixed, so you know which edit helped.
- Expecting a hit chorus. These models are best at beds, loops, ambience and effects.
Fix Weak Transitions
Generated music can show weak transitions, repetition, unstable rhythm, background noise and abrupt endings. On Stable Audio 3 you patch the bad section with inpainting instead of starting over. On 2.5, generate a little longer than you need, trim the weak ending in any audio editor, and add a short fade.
At the time of writing, the PicassoIA catalog lists Stable Audio 2.5, not Stable Audio 3. It is still the quickest way to test a prompt style without installing anything, and the habits transfer.

Step by Step
- Open the Stable Audio 2.5 page on PicassoIA.
- Paste a prompt built with the five part formula above.
- Set the duration in seconds. The default is 190, the maximum.
- Leave steps at 8 for a fast first pass.
- Leave CFG scale at 1, then raise it if the result ignores your instruments.
- Press generate and listen. Most 90-second tracks finish in around six seconds.
- Copy the seed you like, change one word in the prompt, and generate again.
Settings Worth Changing
| Setting | Default | What it does |
|---|
| Prompt | None | Describes genre, instruments, mood and tempo |
| Duration | 190 | Length of the track in seconds |
| Steps | 8 | More steps trade speed for sharper detail |
| CFG scale | 1 | Higher values follow the prompt more closely |
| Seed | Random | Reuse it to reproduce a result |
Try short durations first. A 30-second loop costs the least time to judge, and you can extend a winning prompt to full length afterward.
Add Vocals, Voiceovers and Transcripts
Because neither version sings, a finished project usually needs two or three tools. PicassoIA groups them in one place.

Songs With Vocals
When the brief calls for lyrics and a singer, switch models. Music 2.5 generates full songs with vocals. Lyria 3 Pro creates full-length songs, and Music from ElevenLabs composes tracks from a text prompt. A common workflow: draft the instrumental bed in Stable Audio, then compare it against a vocal model to see which one fits the video.
Voiceovers and Transcripts
For narration over your instrumental, Speech 2.8 HD gives studio-quality voiceovers, Gemini 3.1 Flash TTS offers 30 voices across 70+ languages, and V3 from ElevenLabs is a natural-sounding option for longer reads.
To turn recordings back into text, run them through GPT 4o Transcribe or Gemini 3 Pro. Transcripts help with captions, show notes, and checking that a voiceover says exactly what the script says before it goes under the music.
Make Your First Track Today
So which one should you pick? It depends on the job:
| Job | Best pick | Why |
|---|
| Footsteps, impacts and ambience pack | Stable Audio 3 Small SFX | Free, offline, runs on a CPU |
| Loop for a video or reel | Stable Audio 3 Small or Stable Audio 2.5 | Fast and cheap, and 2.5 needs no setup |
| Six-minute background bed | Stable Audio 3 Medium | Longest open model, needs a GPU |
| Audio on demand inside a product | Stable Audio 3 Large (API) | Top tier, no servers to maintain |
| Quick test of a prompt idea | Stable Audio 2.5 on PicassoIA | Results in seconds, nothing to install |
The fastest way to hear the difference in prompts is to try them. Open Stable Audio 2.5 on Picasso IA, paste one of the prompts from the table above, and generate a 30-second loop. Then change the tempo, swap one instrument, and listen again. Within ten minutes you will have a feel for how this family responds, and when you are ready for voices or transcripts, the speech and transcription models sit one click away on the same platform.