How does a model like Demucs pull a voice out of a finished track?
Stems · 3 min read
It does not cut the voice out of the file. A network trained on songs whose parts were recorded separately listens to the finished mix and writes a new estimate of each part: the vocal you get is a prediction, and the instrumental is the sum of the other predictions.
Why the voice cannot just be subtracted
A finished track is a sum: vocal plus drums plus bass plus everything else, at every moment. Where the parts sit at different frequencies, the split is easy. Where they overlap (a sung note and a synth lead in the same octave, a kick and a bass note on the same beat), the mix alone does not hold the answer. The model fills that gap with what it learned from several hundred songs whose parts were kept apart, so it does best on tracks like those.
How it listens
The model reads the mix twice: as the raw wave, and as frequencies over time. Each fails in its own way (hollow drum attacks from frequencies alone, crunchy static on vocals from the wave alone), so one network compares the two before writing the stems.
A sound it could have recognised from two bars earlier has to be recognised without that context.
How clean the stems are
In the authors' tests, drums and bass come out cleanest; "other", everything that is not vocals, drums or bass, worst. The error is bleed (a hi-hat whisper left under the vocal) or artifacts (warble, smeared transients, static that were in no part). Where two parts share frequencies, a wrong split shows up twice: bleed in one file, a hole in the other. The test songs come from multitrack libraries and their mixes are plain sums of their stems, while a commercial release has a limiter after the sum. Neither the papers nor we have measured what either difference does to club music.
What we run
Our stem tool sends your file to fal.ai, where Demucs runs; nothing is separated on our servers. You get six stems by default (guitar and piano added) or four, each a WAV in one ZIP; for an instrumental, sum every stem but the vocal. We set fal's storage expiry to 60 minutes on the upload and the stems.
The six-stem model has no published score, and its authors' README reports "a lot of bleeding and artifacts" for piano: treat that stem with suspicion.
What this does not tell you
- How our stems sound. No listening test, no score on our own output. Every quality figure is the authors', for a model we do not run.
- What the error sounds like, or where. Faint bleed and a short crunchy artifact of the same energy score the same, and a median can hide the chorus where the vocal falls apart.
- Club music. Nobody we can cite has measured separation on techno, house or minimal.
How we measured: the published scores, our conversions and the details
Method. No model run and no listening test of our own. The architecture and every quality figure come from the three Demucs papers (waveform, hybrid, transformer) and the facebookresearch/demucs repository, the dataset facts from the MUSDB18-HQ release, the metric from BSS Eval and the SiSEC 2018 campaign. The only numbers we computed are conversions of those published figures (SDR into an energy ratio, STFT sizes into milliseconds, the segment stride), with content-social/drafts/scripts/I4_sdr_energy.py; I4_figures.py draws the figures. Which model enguetedl.ch runs comes from our own code.
The terms behind the plain words. "Read as a wave" is the waveform branch, "read as frequencies" the spectrogram branch, and "one trained network" is Hybrid Transformer Demucs (version 4), whose transformer layers let the two branches consult each other repeatedly before the stems are written. Combining the two branches (version 3) gained 1.4 dB of SDR over the waveform-only model. The default six-stem model is htdemucs_6s, the four-stem option htdemucs; the error figure in the body is for htdemucs_ft, the best released four-stem model. The four sources, drums, bass, vocals and "other", are a benchmark convention, not a property of music. "Several hundred songs" is MUSDB18-HQ's training tracks plus the 800 songs below.
SDR, the signal-to-distortion ratio of BSS Eval, is the energy of the true stem divided by the energy of everything wrong with the prediction, in decibels. Every 10 dB means the error carries a tenth of the energy; the body's "1 in N" is that ratio, rounded (table further down).
MUSDB18-HQ. 150 full-length tracks, about 10 hours, 100 for training and 50 for testing. Each ships as the stereo mix plus four stems, uncompressed 44.1 kHz WAV. 100 come from DSD100 (drawn from the Mixing Secrets multitrack library), 46 from MedleyDB, 2 from a Native Instruments stems pack, 2 from a remix competition. Licensed for education and research only. The mix is, by construction, the plain sum of its four stems. The four-source split dates from 2015.
The 800 extra songs. The authors started from 3,500 songs by 200 artists, sorted each stem into one of the four sources by the name the producer gave the file ("vocals2", "fx", "sub"), which they call noisy and sometimes ambiguous, then used an earlier model to discard every song whose stems did not separate cleanly into their labels. 800 survived.
Training. Stems are remixed across songs, pitch-shifted and tempo-stretched; the loss is the distance between predicted and true waveforms. Switching remixing off cost 0.7 dB of SDR (8.70 against 8.00). Training from scratch on musically coherent remixes did worse, likely, the hybrid paper's author writes, because the model then leans on melodic structure.
The two branches. Waveform: five convolution layers, each shrinking time by four, so one step stands for 1,024 samples; it descends from the 2019 Demucs. Spectrogram: STFT over 4,096 samples (92.9 ms at 44.1 kHz), hop 1,024 (23.2 ms), 2,048 bins 10.77 Hz apart; convolutions along the frequency axis, with a frequency embedding because "music is not the same at every frequency". The hop equals the waveform branch's reduction, so the two line up step for step. On the way out, the spectrogram branch is inverted to a waveform and added to the waveform branch's output: that sum is the stem.
| Approach | Characteristic artifact | Where it shows most |
|---|---|---|
| Spectrogram, reusing the mix's phase | hollow attacks | drums, bass |
| Waveform only | crunchy static noise | vocals |
Listening test, Hybrid Demucs (version 3), out of 5: absence of bleeding 3.04 against 2.37 for the waveform-only model. Hybrid Demucs vocals: 2.55 for freedom from artifacts and 2.88 for freedom from bleed, where the true stems scored 4.18 and 4.43. The true stems scored 4.12 for quality, not 5. The transformer paper reports SDR only, no listening test.
The transformer. Hybrid Transformer Demucs (Rouard, Massa, Défossez, Meta AI, ICASSP 2023) keeps the outer four layers of each branch and replaces the innermost ones with a five-layer cross-domain transformer encoder. Its layers alternate self-attention (within one branch, every position against every other in the excerpt) and cross-attention (each branch reads the other's representation). It needs data: on MUSDB18-HQ alone it scored 7.52 dB, below the older hybrid model's 7.64; with the 800 extra songs 8.80, against 8.34 for the older model on the same data. Excerpts of 7.8 s instead of 3.4 s: 8.17 to 8.70 dB.
Segments. The transformer versions accept at most 7.8 s at a time. A track is cut into segments overlapping by 25% with a linear transition, so a new segment starts every 5.85 s and neighbours share 1.95 s. The bar ticks in the figure are our arithmetic: 4 beats at 128 BPM = 1.875 s, so 7.8 s is about 4.2 bars.
The released versions.
| Name | What it is |
|---|---|
htdemucs |
one Hybrid Transformer model, MUSDB plus 800 songs, the repository's default |
htdemucs_ft |
four copies, each fine-tuned for one source; about 4 times slower, "might be a bit better" |
htdemucs_6s |
experimental six-source version: adds guitar and piano |
| Sparse HT Demucs | the paper's best model (9.20 dB); not released, it needs custom GPU code |
Reported SDR. MUSDB18-HQ test set, dB, Table 1 of the transformer paper. "Extra songs" is training data beyond MUSDB.
| Model | Extra songs | All four | Drums | Bass | Other | Vocals |
|---|---|---|---|---|---|---|
| Ideal ratio mask (oracle) | n/a | 8.22 | 8.45 | 7.12 | 7.85 | 9.43 |
| Demucs v2 (2019) * | 150 | 6.79 | 7.58 | 7.60 | 4.69 | 7.29 |
| Hybrid Demucs (v3) | 800 | 8.34 | 9.31 | 9.13 | 6.18 | 8.75 |
| HT Demucs | none | 7.52 | 7.94 | 8.48 | 5.72 | 7.93 |
| HT Demucs | 800 | 8.80 | 10.05 | 9.78 | 6.42 | 8.93 |
| HT Demucs, fine-tuned per source | 800 | 9.00 | 10.08 | 10.39 | 6.32 | 9.20 |
| Sparse HT Demucs, fine-tuned (not released) | 800 | 9.20 | 10.83 | 10.47 | 6.41 | 9.37 |
* Evaluated on MUSDB18, the compressed version, which lacks content above 16 kHz.
The ideal ratio mask is an oracle: it rescales the mix's spectrogram using the true stems, which no real method has. SiSEC 2018 computes it as one of three standard oracle references for how well that kind of filtering can do; it is a yardstick, not the best any method could reach. HT Demucs scores above it on drums and bass and below it on vocals and other.
Across the rows: drums and bass separate best, vocals close behind, "other" worst by about 3 dB. The repository quotes 9.00 dB for htdemucs_ft, no per-model figure for plain htdemucs, none for htdemucs_6s. The released htdemucs was trained on 10-second excerpts, while the paper reports 3.4, 7.8 and 12.2 s, never 10.
SDR as an energy ratio (script output). Error energy as a share of the true stem is 10^(-SDR/10).
| Figure | SDR (dB) | Error energy, share of the true stem |
|---|---|---|
| HT Demucs fine-tuned, vocals | 9.20 | 0.120, about 1 part in 8 |
| HT Demucs fine-tuned, drums | 10.08 | 0.098, about 1 in 10 |
| HT Demucs fine-tuned, bass | 10.39 | 0.091, about 1 in 11 |
| HT Demucs fine-tuned, other | 6.32 | 0.233, about 1 in 4 |
| Ideal ratio mask, vocals | 9.43 | 0.114, about 1 in 9 |
| Demucs v2, vocals | 7.29 | 0.187, about 1 in 5 |
How the score is computed. BSS Eval first finds the linear filter that best maps the prediction onto the true stem and counts what that filter explains as correct, so a stem that is right but slightly quieter or differently equalised is not punished. Since SiSEC 2018 the filter is fixed per track (BSS Eval version 4). The score is computed on 1-second windows, the median taken per song, and the reported figure is the median over the 50 test songs. The 2021 Music Demixing Challenge used a simpler SDR with no matching filter, over the whole track: under it, a perfect stem 1 dB too quiet scores 19.3 dB, and 3 dB too quiet 10.7 dB (script output). Figures from the two definitions cannot be compared.
What SDR does not measure. BSS Eval scores are in the end squared-error criteria, and listeners have in some cases preferred a method the scores rank lower. The true stems are not perfect either: MUSDB's errata list tracks with other instruments bleeding into the vocal stem, and one whose synth vocals sit in "other". The score describes 50 songs from multitrack libraries: a fair comparison between models on those 50, not a forecast for a given track. The 800 extra songs are described only as diverse in genre.
The six-stem model. Released on 7 December 2022, when version 4 reached PyPI. One checkpoint of the same design with six outputs: vocals, drums, bass, guitar, piano, other; "other" gets narrower, nothing is added to the vocal. No description of where its guitar and piano stems came from. The README also says: "Note that the piano source is not working great at the moment." No update has been published; the repository is no longer maintained, and its author's fork is not actively maintained either. The authors do not say why piano fails; the four-source stems were labelled by file names the authors call noisy, and the six-source labelling is not described.
What we send to fal.ai. Only the model, the stem list and the output format (fal-ai/demucs), so fal's defaults apply to the rest: one pass, 25% segment overlap. Hardware, precision and package version are not stated in its documentation. Summing the non-vocal stems is also what the Demucs command line does by default when asked for "vocals and no vocals". Whether four or six stems gives a better vocal has not been published.
More limits. No paper row is exactly the released htdemucs, and htdemucs_6s has no score at all. The test songs come from multitrack libraries; the 800 extra songs are described only as diverse in genre. Some research models score higher on some sources; this page explains one family and does not rank tools.
Related
- Why does our mastering engine return the same file every time?
- Which audio files does a CDJ actually load?
- What does a 320 kbps MP3 throw away compared to the lossless file?
Sources
- Alexandre Défossez, Nicolas Usunier, Léon Bottou, Francis Bach. "Music Source Separation in the Waveform Domain." arXiv:1911.13254, 2019 (revised 2021). https://arxiv.org/abs/1911.13254
- Alexandre Défossez. "Hybrid Spectrogram and Waveform Source Separation." Proceedings of the ISMIR 2021 Workshop on Music Source Separation (MDX 2021). arXiv:2111.03600. https://arxiv.org/abs/2111.03600
- Simon Rouard, Francisco Massa, Alexandre Défossez. "Hybrid Transformers for Music Source Separation." ICASSP 2023. arXiv:2211.08553. https://arxiv.org/abs/2211.08553
facebookresearch/demucs, README anddocs/training.md, read 2026-10-03. https://github.com/facebookresearch/demucs- Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, Rachel Bittner. "MUSDB18-HQ, an uncompressed version of MUSDB18." Zenodo, 2019. doi:10.5281/zenodo.3338373. Dataset page with errata: https://sigsep.github.io/datasets/musdb.html
- Emmanuel Vincent, Rémi Gribonval, Cédric Févotte. "Performance measurement in blind audio source separation." IEEE Transactions on Audio, Speech, and Language Processing 14(4), 2006, pp. 1462-1469.
- Fabian-Robert Stöter, Antoine Liutkus, Nobutaka Ito. "The 2018 Signal Separation Evaluation Campaign." LVA/ICA 2018. arXiv:1804.06267. https://arxiv.org/abs/1804.06267
- fal.ai,
fal-ai/demucsAPI reference, read 2026-10-03. https://fal.ai/models/fal-ai/demucs/api
We build enguetedl.ch; the stem tool described above is ours, and it runs on fal.ai.