Generate musicGenerate speechTranscribe audio

Stable Audio 3 vs 2.5: Open Models, API and Prompt Tips

Stable Audio 3 brings open weights, six-minute tracks and inpainting, while Stable Audio 2.5 stays the quick, zero-setup option. See the specs side by side, what the Community License allows, what the API costs per track, and the prompt formula that gets cleaner music from both.

Stable Audio 3 vs 2.5: Open Models, API and Prompt Tips
Cristian Da Conceicao
Founder of Picasso IA

Stability AI shipped Stable Audio 3 on May 20, 2026, and the first question everyone asks is whether it retires version 2.5. It doesn't. The two now do different jobs. The new family brings open weights you can run on your own hardware, tracks that stretch past six minutes, and editing tools like inpainting. Stable Audio 2.5 stays the faster, zero-setup way to get a clean instrumental in under a minute. Below you get the specs side by side, the license and API cost in plain numbers, prompt tips that hold up on both, and a walkthrough for running 2.5 on PicassoIA today.

Overhead view of an oak desk with a laptop, headphones, notebook and MIDI controller

Stable Audio 3 at a Glance

Stable Audio 3 is not one model. It is a family of four built on a new architecture: a semantic-acoustic autoencoder that squeezes audio roughly 4096 times and lets the model generate any length down to the second. Three of the four ship as open weights on Hugging Face. The biggest one is API only.

The Four Models

ModelParametersMax lengthAccess
Small SFX459 millionAbout 2 minutesOpen weights
Small (music)459 millionAbout 2 minutesOpen weights
Medium1.4 billion6 min 20 secOpen weights
Large2.7 billionOver 6 minutesAPI or enterprise hosting

All four take a text prompt. The family also handles audio-to-audio, which reshapes an existing clip while keeping its timing, and inpainting, which regenerates one stretch of a track while the rest stays untouched. Inpainting comes in three flavors: a single segment, several segments at once, and causal continuation that extends a track from where it ends. Small and Medium accept LoRA fine-tuning, and ComfyUI support arrived on launch day.

Four condenser microphones of increasing size lined up on a wooden table

💡 Worth knowing: none of the four models sing. There are no vocals and no lyrics anywhere in the family. If a track needs a voice, you add it with a separate model, and the last sections show how.

What Changed Since 2.5

Stable Audio 2.5 tops out at 190 seconds per generation, about three minutes. Stable Audio 3 Medium doubles that, at 6 minutes 20 seconds. Training also got cleaner: the dataset is reported at about 1.28 million licensed recordings, roughly 806,000 tracks from AudioSparx and about 473,000 Creative Commons files from Freesound. Copyrighted material was screened out, which matters if you plan to publish what you generate.

Specs Side by Side

Here is the short version in one table. Prices and lengths come from public launch reports and the model page on PicassoIA, so check the official pricing page before you budget a large project.

FeatureStable Audio 2.5Stable Audio 3
Max length190 seconds6 min 20 sec (Medium)
WeightsHosted onlyOpen for Small SFX, Small, Medium
Run on your machineNoSmall on CPU, Medium on GPU
VocalsInstrumental focusNone
API price per runAbout $0.20 (20 credits)$0.26 (26 credits)
Editing toolsText prompt on PicassoIAText, audio-to-audio, inpainting, continuation
Fine-tuningNoLoRA on Small and Medium
SetupNone, runs in a browserDownload weights or call the API

Two black studio monitors facing the camera in a treated room

Where 2.5 Still Wins

Speed and simplicity. The example runs on the Stable Audio 2.5 page produced 90-second tracks in about six seconds each at the default 8 steps. You open a page, type a prompt, press generate. No Python environment, no GPU drivers, no license form.

Price per run. At roughly 20 credits against 26, 2.5 is the cheaper API call when a short loop is all you need.

Predictable iteration. Reuse a seed and you get the same audio back, so you can change one word in the prompt and hear exactly what that word did.

Where 3 Pulls Ahead

Length. A six-minute Medium track is enough for a short film cue or a podcast bed without stitching clips together.

Control. Inpainting lets you fix bar 40 without rerolling the whole song. Continuation extends a track from where it stops. Audio-to-audio reshapes a clip you already have.

Ownership of the workflow. Open weights mean you can run offline, keep unreleased material off third-party servers, and fine-tune with LoRA on your own sound library. A game studio, for example, can train Small SFX on its own foley recordings so every footstep matches the house style.

Honest quality check. One review scored the family 7.8 out of 10 and named the strengths: ambient drones, foley, sound effects, instrumental loops and atmospheric beds. The same review noted that the music can feel correct without feeling interesting, and that rhythmic genres tend to flatten. Treat 3 as a strong tool for beds and effects, not a replacement for a band.

Open Weights: What You Can Run

Open weights change who can use this. Before, audio models at this quality lived behind an API. Now a laptop and a download are enough for the two small models.

A man listening on earbuds with a laptop in a sunlit living room

Small Models on a CPU

Small SFX and Small run on a CPU with no discrete graphics card. Launch reports put generation times under half a second on an H200 server GPU, and the point of the design is full music composition on-device, offline, without the short sample limits older open models had. Use Small SFX for impacts, ambience, footsteps and whooshes. Use Small for loops and short instrumental pieces up to two minutes.

Both small models also take LoRA fine-tuning. A LoRA is a small add-on trained on your own clips, such as your band's drum recordings or a game's foley library, and it nudges every generation toward that character. Because the add-on is tiny, the job is much lighter than retraining a full model.

Medium is the heavier option at 1.4 billion parameters. Setup notes list a CUDA GPU with Flash Attention 2 as the supported path, and the reported speed on an H200 is under two seconds per track. One reviewer measured a few seconds on a MacBook Pro M4, so check the repository for your exact hardware before you plan around it.

License Rules in Plain Words

The open models use the Stability AI Community License.

  • Under $1 million in annual revenue: free to download, run, fine-tune and use commercially once you register for the Community License. You own your outputs, to the extent the law allows.
  • Over $1 million: you need an Enterprise License, which also adds legal indemnification.
  • Large: never downloadable. It is reachable through the Stability API, through fal.ai, or through enterprise self-hosting.

Two hands signing a printed agreement next to an audio interface

💡 Read before you ship: "you own the output" is not the same as "the output is protected by copyright." In many countries purely machine-made audio gets weaker protection than human-made work. If a track is central to your brand, layer your own performance or edits on top.

Using the API

Not everyone wants to install anything. The hosted route is the shortest path from idea to file.

A developer typing at a standing desk with studio headphones around the neck

Cost per Generation

Stability prices in credits, and one credit equals one cent. A Stable Audio 3 generation costs 26 credits, or $0.26. Stable Audio 2.5 is reported at 20 credits, about $0.20. A small gotcha: the free starter balance of 25 credits is one credit short of a single Stable Audio 3 run, so budget a small top-up even for testing.

VolumeAt $0.26 per runAt $0.20 per run
10 tracks$2.60$2.00
100 tracks$26.00$20.00
1,000 tracks$260.00$200.00

When the API Beats Local

Pick the API when you need the Large model, because it is the only way to reach it short of an enterprise contract. Pick it too when you have no GPU, when your volume is bursty, or when a product needs audio on demand without you maintaining servers. Pick local weights when you generate thousands of clips, work offline, or need to fine-tune.

A typical hosted workflow looks like this:

  1. Write the prompt with genre, instruments, mood and tempo.
  2. Set the length you need, in whole seconds.
  3. Send the request and wait for the audio file.
  4. Listen, then change one variable and run it again.
  5. Trim, fade and loop the winner in your editor.

The hosted Stable Audio 2.5 schema is small: a text prompt, a duration, a step count, a CFG scale and an optional seed. Those are the same dials you will meet in most Stable Audio endpoints, so what you practice on 2.5 carries over.

Prompt Tips That Work

Prompts for both versions reward specificity. A vague request like "chill music" returns generic wallpaper. A prompt that names genre, instruments, mood and tempo returns something you can use.

A hand writing in a notebook beside a small analog synthesizer

The Five Part Formula

Write prompts in this order, separated by commas:

  1. Genre: boom bap hip hop, ambient house, post-rock, cinematic synthwave
  2. Instruments: solemn piano, SP-1200 drums, 808 kick, strings, sine wave bass
  3. Mood: peaceful, euphoric, melancholic, tense
  4. Tempo: a number, such as 90 BPM or 125 BPM
  5. Texture or scale: warm, dusty, wide, minor tonality
Weak promptStrong prompt
Relaxing musicSoulful boom bap hip hop instrumental, solemn piano, SP-1200 drums, sine wave bass, peaceful, 90 BPM
Dance trackAmbient house, 808 drum machine, claps, shaker, synth bass, euphoric, 125 BPM
Epic movie musicPost-rock, guitars, drum kit, bass, strings, uplifting, flowing, sentimental, 125 BPM
Door soundHeavy wooden door slam in a stone hall, long echo tail, no music

For longer tracks, add timing cues to the end of the prompt: "soft pads for the first 20 seconds, drums enter after 20 seconds, slow fade at the end." Long Stable Audio 3 tracks benefit most from this, since a six-minute piece needs an arc. On 2.5, keep timing hints short and judge the result by ear.

For sound effects, swap the music formula for three details: the object, the surface it touches, and the space around it. "Boots on wet gravel, empty quarry, distant wind" usually beats "footsteps", because the model can place the sound in a room.

💡 Shortcut: two or three words can also work. Short genre prompts such as "Lo-fi Funk" or "Cinematic Synthwave" gave usable 90-second tracks in the 2.5 examples. Add detail only when the first result drifts.

Mistakes That Waste Credits

  • Asking for vocals. Stable Audio 3 has no vocals or lyrics, and 2.5 is aimed at instrumentals and sound design. Use a song model for singing.
  • Piling on adjectives. Ten mood words cancel each other. Pick two.
  • Skipping the tempo. Without a BPM the rhythm wanders. A number anchors it.
  • Rerolling blindly. Change one thing at a time and keep the seed fixed, so you know which edit helped.
  • Expecting a hit chorus. These models are best at beds, loops, ambience and effects.

Fix Weak Transitions

Generated music can show weak transitions, repetition, unstable rhythm, background noise and abrupt endings. On Stable Audio 3 you patch the bad section with inpainting instead of starting over. On 2.5, generate a little longer than you need, trim the weak ending in any audio editor, and add a short fade.

Run Stable Audio 2.5 on PicassoIA

At the time of writing, the PicassoIA catalog lists Stable Audio 2.5, not Stable Audio 3. It is still the quickest way to test a prompt style without installing anything, and the habits transfer.

Over-the-shoulder view of a woman with headphones watching a waveform on a laptop

Step by Step

  1. Open the Stable Audio 2.5 page on PicassoIA.
  2. Paste a prompt built with the five part formula above.
  3. Set the duration in seconds. The default is 190, the maximum.
  4. Leave steps at 8 for a fast first pass.
  5. Leave CFG scale at 1, then raise it if the result ignores your instruments.
  6. Press generate and listen. Most 90-second tracks finish in around six seconds.
  7. Copy the seed you like, change one word in the prompt, and generate again.

Settings Worth Changing

SettingDefaultWhat it does
PromptNoneDescribes genre, instruments, mood and tempo
Duration190Length of the track in seconds
Steps8More steps trade speed for sharper detail
CFG scale1Higher values follow the prompt more closely
SeedRandomReuse it to reproduce a result

Try short durations first. A 30-second loop costs the least time to judge, and you can extend a winning prompt to full length afterward.

Add Vocals, Voiceovers and Transcripts

Because neither version sings, a finished project usually needs two or three tools. PicassoIA groups them in one place.

A singer in a denim jacket at a condenser microphone inside a vocal booth

Songs With Vocals

When the brief calls for lyrics and a singer, switch models. Music 2.5 generates full songs with vocals. Lyria 3 Pro creates full-length songs, and Music from ElevenLabs composes tracks from a text prompt. A common workflow: draft the instrumental bed in Stable Audio, then compare it against a vocal model to see which one fits the video.

Voiceovers and Transcripts

For narration over your instrumental, Speech 2.8 HD gives studio-quality voiceovers, Gemini 3.1 Flash TTS offers 30 voices across 70+ languages, and V3 from ElevenLabs is a natural-sounding option for longer reads.

To turn recordings back into text, run them through GPT 4o Transcribe or Gemini 3 Pro. Transcripts help with captions, show notes, and checking that a voiceover says exactly what the script says before it goes under the music.

Make Your First Track Today

So which one should you pick? It depends on the job:

JobBest pickWhy
Footsteps, impacts and ambience packStable Audio 3 Small SFXFree, offline, runs on a CPU
Loop for a video or reelStable Audio 3 Small or Stable Audio 2.5Fast and cheap, and 2.5 needs no setup
Six-minute background bedStable Audio 3 MediumLongest open model, needs a GPU
Audio on demand inside a productStable Audio 3 Large (API)Top tier, no servers to maintain
Quick test of a prompt ideaStable Audio 2.5 on PicassoIAResults in seconds, nothing to install

The fastest way to hear the difference in prompts is to try them. Open Stable Audio 2.5 on Picasso IA, paste one of the prompts from the table above, and generate a 30-second loop. Then change the tempo, swap one instrument, and listen again. Within ten minutes you will have a feel for how this family responds, and when you are ready for voices or transcripts, the speech and transcription models sit one click away on the same platform.

Share this article