Speech and audio
Music Source Separation and Stem Quality
Apply separation to music while handling stems, bleed, phase, stereo image, evaluation, and creative rights.
By the end you can
- Define music source separation and stem quality as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish broad stems, fine stems, and target extraction without treating them as interchangeable
- Trace the workflow from define the stem ontology through evaluate intended use
- Evaluate music source separation and stem quality using signal distortion and interference estimates with caveats and evidence from difficult deployment slices
Comparison
Four stems, eleven stems, or just one
Ask a separator to pull a finished record apart. The first thing to settle is how many pieces you want back. Broad stems means four: vocals, drums, bass, and accompaniment. They are the easiest to produce. The ease is paid for later. Once the accompaniment is a single lump, there is very little left to edit. Fine stems go further, down toward the individual instruments. Target extraction goes the other way. It asks for one specified source only, and leaves everything else where it is.
Fine stems are not a vague ambition. There is a corpus and a written ontology for them. MoisesDB, published in 2023, provides 240 tracks across twelve genres. Its design is one sentence: “For each song, we provide its individual audio sources, organized in a two-level hierarchical taxonomy of stems.” The stem level of that taxonomy has eleven entries: Bass, Bowed Strings, Drums, Guitar, Other, Other Keys, Other Plucked, Percussion, Piano, Vocals and Wind. Its baselines report SDR at four, five and six stems rather than at four alone.
Three requests, three different amounts of work. And three different tests — the part that governs the rest of this lesson. Four stems and eleven stems are not the same question asked at different resolutions. They are different questions. A system that does well on one of them has said nothing about the other two.
Broad stems
Vocals, drums, bass, and accompaniment; easier but less editable.
- Decision focus: Define the stem ontology
- Useful evidence: Signal distortion and interference estimates with caveats
- Watch for: Phase-incoherent stereo outputs
- Best used when its assumptions are documented for music source separation and stem quality
Fine stems
Individual instruments or performers; harder and often ambiguous.
- Decision focus: Preserve musical structure
- Useful evidence: Perceptual artifact and preference ratings
- Watch for: Transient smearing that harms percussion
- Best used when its assumptions are documented for music source separation and stem quality
Target extraction
Removes or isolates one known source using a query or reference.
- Decision focus: Choose separation strategy
- Useful evidence: Stereo image, transient, and loudness preservation
- Watch for: Removing shared effects that define the production
- Best used when its assumptions are documented for music source separation and stem quality
Example
The highest SDR of the three finished last with the listeners
One of those tests answering for the others is on record, with names and numbers, from the Sound Demixing Challenge 2023. The organisers did not stop at the scoreboard. They put the finalists in front of seven professional assessors: award-winning singers, songwriters, composers, music producers, sound engineers and an educator. Those assessors made 583 A/B comparisons on ten MoisesDB songs.
The two rankings came out inverted. A TrueSkill rating over the comparisons placed kimberley_jensen first at μ 24.793, then μ 24.362, then μ 24.011. The SDR ordering of the same three systems ran the other way: 9.18 dB, 9.26 dB, 9.97 dB. The system the professionals liked best held the lowest objective score of the three. SAMI-ByteDance's leading 9.97 dB placed third perceptually. The report says it without hedging: “Nevertheless, we find that the model by kimberley_jensen achieved first place in the final ranking, despite being third in the leaderboard obtained using SDR.”
Nothing in the dB column said so, and nothing in it could. The three finalists sat inside a single decibel of one another on the metric. That span bought the opposite of the verdict that mattered to the people who would have to work with the stems.
- That is the decision this lesson is about: whether to apply separation to music at all. Then how to handle stems, bleed, phase, stereo image, evaluation and creative rights once you have.
- The evidence anyone would think to ask for is signal distortion and interference estimates, reported with their caveats. In the challenge those estimates ranked the three finalists in the reverse of the order seven professional listeners put them in.
- One failure runs underneath results like that: stereo outputs that are no longer in phase with one another. It is audible on recombination and nearly invisible to a per-source energy ratio.
- The practical response is to stop testing on the passages that separate easily. Put a dense chorus, a quiet intro, shared reverb and doubled instruments in front of the system before believing anything it scores.
The original multitrack is not recoverable
Before asking why the metric and the assessors disagreed, be clear about what a separator is attempting. Music separation estimates sources — vocals, drums, bass, accompaniment — from a mastered mixture. Those categories come from how records are made. Nothing in the audio is sitting there waiting to be found. Shared effects, stereo processing, distortion and instruments that move together leave the original stems non-identifiable. The mixture does not contain enough to determine them. No amount of sophistication in the separator changes that, which is why how advanced a system is settles so little.
What that leaves is worth carrying through the rest of the lesson. A separated stem may be creatively useful without matching any original multitrack recording. It is an estimate, and it is allowed to be a good one on its own terms.
Which terms, though. Signal metrics, listener preference, editability, stereo coherence, and whether expressive detail survived are five separate questions. The challenge finalists answered them in different orders. One system holding 9.97 dB and third place among professionals is two of the five answered in opposite directions. No single verdict covers both facts. So the match has to be made first. If the stem categories you defined do not fit the use the audio is going to be put to, the separation metrics cannot support the decision at all.
Which of those questions the release depends on is a choice somebody has to make and record.
Case
1,541 submissions scored on the same fifty tracks
A field that cannot recover originals can still compare estimates, and music separation does this better than most. It has a shared test set and a public scoreboard. That is what makes its numbers comparable from one system to the next. MUSDB18 holds 150 stereo tracks at 44.1 kHz, split 100 for training and 50 for test, and every track carries drums, bass, vocals and other. The Music Demixing Challenge 2021 put 1,541 submissions from 61 teams through that same fifty-track test and ranked them by signal-to-distortion ratio. The best system trained only on MUSDB18-HQ reached 7.328 dB. The best one allowed extra data reached 8.326 dB.
Look at what those four names are, though. Drums, bass, vocals and other is a stem definition like any other — broad stems, fixed in place so that everyone is scored against the same one. Sixty-one teams competed on how closely they could hit a boundary somebody else had already drawn. A dB gap of 7.328 to 8.326 is a statement about that. It is not a statement about whether the boundary suited the job in front of you.
Change the training material rather than the boundary and the scoreboard moves much further than a tenth of a decibel. In the Sound Demixing Challenge 2023 the best model restricted to the deliberately corrupted bleeding data scored 6.58 dB global SDR. On the unrestricted Standard leaderboard the figure was 9.97 dB. Same task, same kind of number. What changed was what the models were allowed to learn from.
Example
Bleed, measured at −7 to −12 dB
Arguments about those boundaries lean on four words, and two of them are easy to mistake for the same complaint. Bleed is what one microphone hears of an instrument it was never pointed at. It is in the recording before any model touches it. Phase coherence is what keeps a stereo image from collapsing when the stems are recombined, and it is a property of what comes out the far end. Stem and target extraction name the output either way, all the sources or one of them.
Bleed is the one of the four that has been given a number. For the Sound Demixing Challenge 2023 the organisers built SDXDB23_Bleeding by copying each stem into another at −7 to −12 dB through a low-pass or band-pass filter. Then they measured what training on it cost: “training on SDXDB23_LabelNoise degrades the average separation quality by 1.42dB, while training on SDXDB23_Bleeding degrades it by 0.83dB”. Contamination a producer would call inaudible is worth 0.83 dB of baseline separation quality. That is eight times the 0.1 dB that whole papers have been argued over.
Each of the four terms may ask for its own proof, in its own unit, and leave the decision with somebody else. Which is how a system clears one of the four and gets described as having cleared the set.
- A stem is a grouped audio source used in mixing: the vocals, say, or the drums.
- Bleed is energy from one instrument or microphone turning up inside another source — synthesised in SDXDB23_Bleeding by copying each stem into another at −7 to −12 dB through a low-pass or band-pass filter.
- Phase coherence is the consistency of timing relationships that a stable stereo reconstruction depends on.
- Target extraction is isolating one specified source rather than all of them.
Example
One stem, four incompatible standards
That vocabulary problem is a symptom of a harder one. The same stem goes to four places that want incompatible things from it. A channel that is a clear pass in one of them is a clear failure in another. None of the four is measured the way the challenge scoreboard measures.
The archive case is the one with the least room for bluffing, and there is a released example of it done in the open. On 26 October 2023 Apple Corps announced that a John Lennon vocal had been separated from his piano on a late-1970s home demo, using WingNut Films' MAL de-mixing: “Peter Jackson and his sound team, led by Emile de la Rey, applied the same technique to John’s original home recording, preserving the clarity and integrity of his original vocal performance by separating it from the piano.” The result came out as “Now And Then” on 2 November 2023. It won Best Rock Performance at the 67th GRAMMY Awards in February 2025. Note what was claimed and what was not. A vocal separated from a piano, presented as a new recording. Never as a recovered multitrack.
- Karaoke can care more about removing the vocal intelligibly than about archival fidelity, so a stem that lost the singer's breaths can still be the right stem there.
- Remixing lives or dies on clean transients, stereo phase and preserved effects, and losing any one of the three makes a stem unusable however clean it sounds.
- Music transcription is undone by leakage, which makes it write down false notes and rhythm events that nobody played — the same bleed that cost the SDX 2023 baseline 0.83 dB.
- Archive access will accept an estimate — Apple Corps described a vocal separated from a piano, not a recovered original multitrack — but the estimate must never be labelled as one.
Visual
Shared reverb belongs to no single stem
Four incompatible standards cannot be met by one number. The usual route to one number is treating four decisions as a single act of running a separator. Naming the stems, modelling them, deciding on that basis, and verifying the result are four choices. The path below holds them apart instead of folding separation quality and intended use into the same score.
The first of the four is the one that goes missing. Shared reverberation belongs to no single stem — it is on the voice, in the room, and over everything the room touched. Somebody has to say where it is going before anything gets measured. A written ontology is what records that decision. MoisesDB's two-level taxonomy has eleven stem-level entries, down to Bowed Strings, Other Keys and Other Plucked, precisely so that the awkward material has a named home rather than a default one. Where no such definition exists, the effects, the breaths and the room go wherever the model puts them. Nobody can say afterwards whether that was a choice.
1. Define the stem ontology
State instrument groups, backing vocals, effects, bleed, silence, and unknown content.
2. Preserve musical structure
Handle stereo phase, transients, tempo, shared reverb, and long-range sections.
3. Choose separation strategy
Use waveform, spectrogram, hybrid, conditioned, or multi-stage models.
4. Evaluate intended use
Test remixing, transcription, karaoke, restoration, or search separately.
Draw the stem boundaries once, shared reverb included, and the intended-use evaluation judges the stems that definition produced rather than the definition itself.
Steps
Evaluate one song for three uses
So write that definition down, then try to break it. Take one song, run it through three uses, and expect three different verdicts. That shape is the exercise. Its point is to leave another team able to challenge a phase-incoherent stereo output rather than accept a clean-sounding solo on trust.
Three lines are enough. What does defining the stem ontology assume — where did the shared reverb go, and what has been swept into accompaniment? Then one counterexample, a passage where that assumption misleads you; the quiet intro and the dense chorus are the obvious places to go looking. Then what evaluating against the intended use actually does with the result when karaoke, remix and the archive come back disagreeing.
If the panel sounds expensive, note the scale that was enough to overturn a scoreboard: seven assessors and 583 A/B comparisons across ten songs. That is a week of work, not a research programme. And it produced a ranking the dB column had exactly backwards.
1. Select difficult passages
Include dense chorus, quiet intro, shared reverb, and doubled instruments.
2. Create three briefs
Define karaoke, remix, and archival-restoration expectations.
3. Score separately
Use listening panels and task-specific edits rather than one global metric.
4. Document rights
Track source license, performer rights, output use, and redistribution limits.
Stem quality is only half the verdict; license, performer rights, intended use and redistribution limits decide whether the same stems clear one use and fail the next.
Example
SDR, SIR, SAR — and what their own authors' successors say about them
Three verdicts need more than one kind of number. The first step is to stop calling the number “signal distortion and interference estimates” and name it. The measures come from a single paper of 2006, by Vincent and two colleagues, which the field has cited 3,203 times. Its method is one sentence: “In each case, we decompose the estimated source into a true source part plus error terms corresponding to interferences, additive noise, and algorithmic artifacts.” One energy ratio per error term: SDR, SIR, SAR.
The field then over-read them, and the accusation comes from inside. The team that proposed SI-SDR in 2019 wrote that SDR “has generally been improperly used and abused, especially in the case of single-channel separation, resulting in misleading results”. They describe the habit precisely: “In recent years, hundreds of papers have been relying on this toolkit to evaluate their proposed methods and compare them to previous works, often arguing that differences on the order of 0.1 dB proved the effectiveness of a method over others.”
The mismatch with what people hear is systematic rather than anecdotal. A 2021 review correlated twelve objective measures against human ratings: “We use perceptual scores from 7 listening tests about audio coding and 7 listening tests about source separation as ground-truth data for the correlation analysis.” In the separation domain the perceptual 2f-model came out best, at Pearson ρ = 0.86. SDR-type measures fell in the bottom third. So report the ratios properly. Say what the unit is, which mixes it was measured on, how wide the interval is, and what the playback chain was. And do not report them alone.
- The evidence for the core task is SDR, SIR and SAR as the 2006 paper defined them: an energy ratio per error term, quoted with its caveats rather than as a bare figure.
- Next to it belongs evidence of how the system behaves perceptually, because across fourteen listening tests the 2f-model reached ρ = 0.86 in the separation domain and signal-only measures ranked in the bottom third.
- The robustness slice is what survived the separation — the stereo image, the transients, and the loudness — none of which a per-source energy ratio decomposes.
- Over the working life of the stems, the evidence is task success: whether the remix, the transcription or the edit came out right. That is the question 583 A/B comparisons were asking and the dB column was not.
A gap of 0.1 dB has been used to prove a method better; 0.83 dB is what audible bleed costs, and 9.18 against 9.97 dB separated the three SDX finalists in the wrong order.
Key idea
Recombine the stems and hear the collapse
Four shortcuts close this lesson, and every one of them looks reasonable in the session where it is taken. The first is shipping stereo outputs that are no longer in phase, because the soloed stem sounded fine. The second is transient smearing, inaudible on a sustained pad and ruinous on percussion. The third is removing the shared effects that define the production — the room the singer was standing in. The fourth is distributing derived stems without settling rights and consent. That last one is not an audio question at all, and it will stop the release anyway.
They share one habit, and so does the fifth: reading the dB column as the verdict. Each judges a stem where judging is easiest — soloed, alone, in one energy ratio. The verdict actually falls somewhere else: recombined, edited, re-balanced, in front of somebody who has to work with it. Seven assessors reversed a three-way SDR ranking in the Sound Demixing Challenge 2023. And nearness to an original multitrack was never available as a standard anyway, since separation cannot recover one. The five questions from earlier have to be answered separately, against the intended use, every time.
Judge a separation by what happens when the stems are recombined, not by how clean any one of them sounds soloed.
Key takeaways
- In the Sound Demixing Challenge 2023, seven professional assessors made 583 A/B comparisons on ten MoisesDB songs. They put kimberley_jensen first at μ 24.793 — the lowest SDR of the three finalists, 9.18 dB. SAMI-ByteDance's leading 9.97 dB placed third. The metric ranked them in reverse.
- A separated stem may be creatively useful without matching any original multitrack recording. Shared effects, stereo processing, distortion and correlated instruments leave the originals non-identifiable. Apple Corps described a Lennon vocal separated from a piano, never a recovered multitrack.
- Stem work starts by naming which sources will be produced, shared reverb included. MoisesDB's two-level taxonomy fixes eleven stem-level entries across 240 tracks and twelve genres, and its baselines report SDR at four, five and six stems rather than four alone.
- Broad stems, fine stems and target extraction answer related but different questions, so a result from one of them says nothing about the other two.
- Bleed is measurable damage. SDXDB23_Bleeding copies each stem into another at −7 to −12 dB, and training on it cost the baseline 0.83 dB of separation quality. The best bleeding-restricted model reached 6.58 dB global SDR against 9.97 dB unrestricted.
- SDR, SIR and SAR come from a 2006 paper: one energy ratio per error term. The team behind SI-SDR calls 0.1 dB comparisons misleading. Across fourteen listening tests, a 2021 review found signal-only measures in the bottom third, against ρ = 0.86 for the perceptual 2f-model.