Been waiting for an audio model that fits right into my existing workflow without eating up my GPU, and Stable Audio 3.0 looks like it finally delivers. It’s Stability AI’s new family of music models, all about artistic tinkering, and they claim the training data is fully licensed, so no worries about commercial use.
What really hits the spot for me is the tier system: the Small version for sound effects and music can run straight on CPU, no big GPU needed, and it can crank out almost two minutes of audio. If you want longer, more structured tracks, go for Medium—with a GPU, it stretches to over six minutes. Compared to their earlier Stable Audio Open, which only did like 10 to 40 seconds, this is a massive leap.
The workflow is the same old story: update ComfyUI, pick the Stable Audio 3.0 template under the Audio category in the left sidebar, download the models and put them in the right folder if you’re running locally, then write your prompt, set the duration, and hit run. You can generate sound effects, single instruments, atmospheres, and one-shot samples. What I value most is being able to get two minutes locally on CPU—slapping some ambient background music on a video or making UI sounds for a small game without needing a separate machine just for audio.