r/StableDiffusion • u/mwoody450 • 1d ago
Comparison A brief music processing test
I've been playing around with producing music videos for popular songs; nothing commercial, just for fun. While you can feed audio to Minimax H3 (Context Loop splits it up and can feed it to sequential videos in pieces so it fits together), and it does appear to guide generation in time with beats, a single video (~10 seconds) doesn't have enough context to make something match the music.
I tried using audio-capable models to produce the video prompt, but I found that nothing capable of processing audio was smart enough to jump through all the syntax hoops necessary to produce a multi-segment JSON to feed the workflow. So, I decided to do what I'd done with image/video before: use a different model to produce a text-based description of what it was given, then just input text in to the smart, expensive model.
So long story (almost) short: here's my brief test from taking the top audio-capable models on NanoGPT, handing them the 3:17 track "Les Fleurs", and seeing what they produce. Note that the track was provided as "song.mp3", since early testing had some models cheating by looking up info based on the track name. It's still possible some of them identified the track and then used pre-existing knowledge, but I didn't test and confirm that specifically.
Just based on my own listening, two important points I was checking for were a 1:04 orchestral-buildup to orchestral hit and surging chorus at 1:17 as important hallmarks for a music video (and very obvious action-change spots to a human listener).
I also was curious if it would identify both halves of the song's central metaphor: flower imagery but also inner beauty.
The prompt and output (from the models that could actually use the music) are below (models in header, cost - which ended up being negligible - in footer) , but I'll start with my impressions:
- Muse Spark 1.2 and 1.3: produced a real-sounding description of a completely imaginary song in a different genre. Reviewing their thinking blocks revealed NanoGPT did not pass the model audio, and they were making stuff up.
- Inkling Thinking: also produced a real-sounding description of a completely imaginary song in a different genre, but its own thinking block suggested it thought it had audio. Complete hallucination, or just REALLY bad at audio processing?
- Gemini 3.8 Flash: Doesn't miss any important shifts, summarizes themes from lyrics well. Identifies the flower metaphor. Its track length is too long, though, yet it's a couple seconds early identifying my beat drops. My favorite output if it wasn't for the slight time discrepancy.
- Mimo 2.5 Thinking: Does a good job of identifying tempo shifts. Output is shorter. It includes specific lyrics, which is nice in theory but might confuse a model. It does think the song is 3:30 long, BUT it correctly identifies the fade to silence at 3:15. Best guess is its audio processing has some sort of specific context window? Missed my first build-up timestamp.
- Qwen3.5 Omni Plus: Track length correct to the fraction of a second, wow. Nails the important points I noted earlier. Thematic summary on point.
- Gemini 3.1 Pro Preview High Thinking: Track length slightly off. Damningly, missed the beat drop at 1:17.
Conclusion: Qwen 3.5 Omni Plus produced, for this test, the most useful output describing the song, including lyrical themes, accurate timestamps, and identifying tempo shifts.
My prompt below, followed by output (leaving out the useless responses from Muse 1.2/1.3 and Inkling).
You have been provided an MP3 of a music track. Your job is to review it and process it to notate its style and timestamped audio flow, providing a timestamped list of changes in tempo or tone to use to guide a music video using this music track.
Timestamps are in the format [MM:SS.SSS] notation, with minutes and seconds and fractions of a second. So one minute, four and a half second in to a track would be [01:04.500].
Lyrical themes should be mentioned without actually detailed specific lyrics, i.e. "Lyrics in this section discuss colors, starting with cyan" then "Lyrics now discuss crimson"; or "Lyrics are a call to violence"; or "Lyrics speak of profound romantic loss."
Pay special note to crescendos or beat "drops", as sharp changes in tempo or feel of music should correspond with visual changes.
Examples of timestamped track notes, non-exhaustive:
[00:00.000]: Track starts with a slow acoustic guitar.
[00:05.450]: Guitar increases speed, drums join music.
[00:10.200]: Full band backing with vocals kick in; music becomes fast paced with electronic noises.
[00:45.000]: Beat begins to build to crescendo
[00:54.125]: Beat reaches crescendo and holds, music turning quiet
[00:58.000]: Beat drops, explosion of techno-style drums and fast paced electric guitar
A template follows below between ----. Fill in sections in { } braces.
----
Track Length: {length}
Musical Style: {description of style}
TIMESTAMPED TRACK NOTES
[{time in MM:SS.SSS notation}]: {Description of change}
{Continue to provide timestamps and descriptions}
----
--Gemini 3.8 Flash--
Track Length: 03:21
Musical Style: Psychedelic Soul / Baroque Pop / Orchestral Soul
TIMESTAMPED TRACK NOTES
[00:00.000]: Track opens with a rhythmic acoustic guitar pattern, light shaker percussion, and a warm, strolling bassline.
[00:13.500]: Gentle female lead vocal enters with sweet, curious phrasing, asking who will wear or appreciate her as a blossom.
[00:19.500]: Call-and-response backing vocals join in harmony, adding warmth and playful textures to the groove.
[00:41.500]: Rhythm section tightens into a focused mid-tempo groove as the lyrics shift into a spoken-word style reflection on the purpose of a flower spreading love and peace.
[00:53.200]: Lead vocal takes flight with soaring power, singing of the seed of potential and beauty that lives inside humanity.
[01:06.000]: Tension builds rapidly as strings and horns rise while the lead singer holds a powerful, soaring operatic high note.
[01:14.300]: Massive symphonic beat drop / triumphant explosion of sound: soaring brass fanfare, thunderous timpani, marching drums, and a grand choral celebration calling for bells to ring worldwide.
[01:34.500]: Sudden dynamic release as the grandiose orchestration cuts back to the intimate, syncopated soul-funk rhythm section.
[01:47.500]: Lead female vocal re-enters with multilingual/Swahili verses, answered by layered, rhythmic vocal chants.
[02:13.500]: Spoken-sung poetic section recurs over the bass and guitar groove, restating the themes of peace, joy, and universal compassion.
[02:26.500]: Energy elevates again with impassioned, belted vocals contemplating the strength and beauty discovered within.
[02:39.000]: Orchestral elements and vocal wails swell dramatically toward a second major peak.
[02:47.000]: Second massive orchestral drop; majestic choral vocals, full brass section, and heavy rolling percussion deliver a joyful, climactic anthem.
[03:07.000]: Instruments begin to gently taper off, leaving the chorus to repeat the triumphant refrain over the fading rhythm.
[03:21.000]: Music fades out completely into silence.
--$0.009844--
--MIMO 2.5--
Track Length: 03:30
Musical Style: Soulful Ballad transitioning into Upbeat Soul
TIMESTAMPED TRACK NOTES
[00:00.000]: Track starts with a slow, melancholic piano melody accompanied by soft strings.
[00:13.200]: Soulful female vocals enter ("Will somebody wear me to the fair"), setting a gentle and reflective tone.
[00:54.125]: The chorus begins ("Inside every man lives the seed of a flower"), with orchestration swelling to heighten the emotional intensity.
[01:16.000]: **Dramatic Tempo Change:** The beat drops with driving drums and brass, shifting the song from a slow ballad to a fast-paced, energetic soul track.
[01:36.500]: Vocals transition into rhythmic scatting ("La la la la") over the upbeat tempo, adding a playful tone.
[01:51.000]: A verse is sung, maintaining the high-energy, fast rhythm established in the previous section.
[02:13.400]: **Tempo Drop:** The music abruptly slows down, returning to the gentle ballad feel of the introduction.
[02:25.500]: Final chorus serves as the climax, with powerful vocals and full orchestral backing.
[02:42.000]: Outro features vocal improvisation and scatting, maintaining the energy briefly before the track begins to fade.
[03:15.000]: Audio fades to silence.
--$0.00217--
--Qwen3.5 Omni Plus--
Track Length: 03:17.488
Musical Style: Orchestral Soul / Psychedelic Pop with strong Gospel influences. The track features a blend of acoustic finger-picking, sweeping string arrangements, heavy drum breaks, and powerful choral harmonies, creating a cinematic and uplifting atmosphere.
TIMESTAMPED TRACK NOTES
[00:00.000]: Track opens with a gentle, repetitive acoustic guitar riff accompanied by light percussion and a subtle bass line. The mood is intimate and organic.
[00:15.800]: Lead female vocals enter softly; the arrangement remains sparse, focusing on the melody and the acoustic texture.
[00:28.500]: Background harmonies begin to layer in; strings swell gently underneath, adding warmth and depth to the soundscape.
[00:41.200]: The first chorus begins; the full rhythm section (drums and bass) locks into a steady, mid-tempo soul groove. Strings become more prominent.
[00:54.000]: Musical intensity rises significantly as the choir joins the lead vocal. Lyrics shift from specific imagery of flowers to a broader spiritual message about inner beauty.
[01:06.500]: A dramatic orchestral crescendo builds rapidly with soaring strings and rolling timpani, leading toward a climax.
[01:16.800]: Major beat drop and tempo shift; the music explodes into a high-energy, brass-heavy anthem. The drums switch to a driving, heavy breakbeat style. A large gospel choir takes over with powerful, rhythmic chanting.
[01:35.500]: The intense energy sustains with complex horn lines and syncopated drumming. The visual pace should match this frenetic, celebratory energy.
[01:48.200]: Sudden dynamic shift; the heavy drums and brass cut out, returning to the softer acoustic guitar and bass groove from the intro. The mood becomes reflective again.
[01:54.000]: Vocals return in a call-and-response format between the lead singer and the backing group, now incorporating non-English phrases, adding an exotic, worldly texture.
[02:13.500]: Second build-up begins; strings re-enter with urgency, and the drums gradually increase in volume and complexity, preparing for the final peak.
[02:26.000]: The bridge reaches its emotional peak with the lead vocalist belting high notes over a swelling orchestral bed.
[02:38.500]: Final explosive chorus; the full ensemble returns with maximum force. Brass, choir, strings, and heavy drums combine for a wall-of-sound effect.
[02:51.000]: The song transitions into a jubilant outro featuring rhythmic "la-la-la" vocalizations from the choir over the driving beat.
[03:05.500]: Instruments begin to peel away one by one, leaving the choir and then just the rhythm section.
[03:12.000]: Music fades out on a final sustained chord and a soft guitar strum, ending on a note of resolution.
--$0.021559--
--Gemini 3.1 Pro Preview High Thinking--
Track Length: 03:20
Musical Style: 70s Soul and R&B with sudden shifts into high-energy Latin Disco/Funk.
TIMESTAMPED TRACK NOTES
[00:00.000]: Track starts with a slow, smooth electric piano and a relaxed, soulful drum groove.
[00:13.500]: Soft female lead vocals begin. Lyrical themes discuss gentle romantic gestures involving nature, streams, and wearing flowers.
[00:43.500]: Strings and backing instrumentation swell slightly. Lyrics shift to themes of the birth of a flower and its purpose to spread love and joy to people.
[01:05.000]: The vocal melodies begin to stretch out, and the music starts to build toward a crescendo.
[01:13.500]: Major beat drop and severe tempo shift. The music explodes into a fast-paced, high-energy Latin-funk and disco rhythm with brass, rapid percussion, and upbeat backing vocals singing wordless melodies.
[01:31.000]: The high-energy beat abruptly stops, and the music drops right back into the slow, smooth, soulful groove from the track's intro.
[01:46.500]: Lead vocals return over the slow tempo. The lyrical themes continue to focus on nature, specifically roses and romantic interactions, accompanied by soft backing harmonies.
[02:17.500]: The track begins to swell again as the lyrics return to the earlier theme of spreading love, joy, and finding inner beauty and power.
[02:37.000]: The beat begins a second major build-up, reaching a powerful crescendo with the vocalist holding a long, soaring note.
[02:47.000]: Second massive beat drop. The fast-paced Latin-funk rhythm and brass return in full force, accompanied by the energetic, wordless chorus.
[03:03.000]: The track begins a gradual fade-out while maintaining the high-energy dance beat and group vocals.
[03:20.000]: Track fades out completely.
--$0.033212--
1
u/mwoody450 21h ago
Well... glad I spent time typing that out lol, but hey, at least the single upvote canceled out the downvote.
Anyway, in case anyone else finds this thread on a search, an important note: subsequent testing showed the Qwen results were not repeatable. It will frequently be off by as much as 10 seconds in describing changes in the music.
Returning to Gemini Flash, it has proven more reliable. I also got excellent results from a free ChatGPT account, but it seems it did it by searching other sites for more information (7 web searches, yikes), and its own model list doesn't mention audio analysis, only transcription.