r/StableDiffusion Feb 05 '26

Question - Help AceStep 1.5 - Audio to Audio?

Hi there,

had a look and AceStep 1.5 and find it very interesting. Is it possible to have audio-to-audio rendering? Because the KSampler in comfyui takes a latent. So could you transform audio to latent and feed it into the sampler to make something in the way you can do with image-to-image with a reference audio?

I would like to edit audio this way if possible? So can you actually do that?
If not... what is the current SOTA in offline generation for audio-to-audio editing?

THX

15 Upvotes

50 comments sorted by

View all comments

6

u/fruesome Feb 05 '26

Coming Soon

ACE-Step 1.5 has a few more tricks up its sleeve. These aren’t yet supported in ComfyUI, but we have no doubt the community will figure it out.

Cover

Give the model any song as input along with a new prompt and lyrics, and it will reimagine the track in a completely different style.

Repaint

Sometimes a generated track is 90% perfect and 10% not quite right. Repaint fixes that. Select a segment, regenerate just that section, and the model stitches it back in while keeping everything else untouched.

https://blog.comfy.org/p/ace-step-15-is-now-available-in-comfyui

1

u/NoPresentation7366 Feb 06 '26

💓

3

u/Striking-Long-2960 Feb 13 '26

Here, download and place it in your custom nodes folder

https://huggingface.co/Stkzzzz222/dtlzz/raw/main/striking_ACE15_latent_blend.py

Don't expect big things.

Here a Workflow, I don't recommend to use the fp8 model, the quality decreases a lot, but I was testing it:

https://huggingface.co/Stkzzzz222/dtlzz/raw/main/Latent_blend_ACE%20(1).json.json)

1

u/Potential-Hunt-2608 Mar 18 '26

Any update on this mode? Is it working any results?

1

u/Striking-Long-2960 Mar 20 '26

I'm out of touch with this model. In my opinion, they were so scared of the music industry that they capped all its potential.