Why does our mastering engine return the same file every time?
Mastering · 3 min read
Because it is a fixed chain of measurements, EQ, gain and limiting, with nothing random and no trained model in it. Same track, same settings, same engine version: the same file, byte for byte. We mastered three test mixes twice each, and every pair matched.
What "the same" rests on
- No randomness. The one place the chain uses noise, the synthetic reference, draws it from a fixed seed.
- No trained model. Its tonal curves and loudness targets were computed once, offline, from a library of club records, and are read like a table.
- No hidden state. Nothing updates itself from uploads. Every setting that can change the sound is hashed into one engine ID, so a change shows up as a new ID.
The target: your reference, or noise shaped like one
With a reference of your own, the engine matches to that file. Without one, it measures the shape of your sound, not its level, picks the closest of four fixed centres derived from a reference library, and renders the target as 60 seconds of noise with that centre's tone and loudness.
Noise works because the matching stage reads only a reference's average tone, the level of its loudest passages and its peak. Rhythm and arrangement are averaged away, so noise with the right averages is as good a reference as a record.
Matching, then three guards
Matching is Matchering 2.0, an open-source library with one change of ours: it brings your level and tonal balance toward the target. Then:
- Stereo safety. Your width comes back above 150 Hz; below it the bass goes mono, so it cannot cancel on a mono club system.
- Makeup gain. Mono bass costs some loudness, so the engine adds back 95% of the gap, never more than 6 dB.
- True-peak limiter. FFmpeg's limiter, 8x oversampled so it sees peaks between samples, aimed at our -0.15 dBTP ceiling. Out comes a 16-bit WAV.
So the ceiling is a target with a small tolerance, not a wall. The test renders behind Club loudness, measured all held it; this mix did not quite.
What the engine does not do
- No compressor, saturation, exciter or widener, and never wider than your mix.
- It does not fix a mix, or change arrangement, tempo, key or length.
- It does not learn from your uploads, and does not guarantee a loudness number: the limiter has the last word.
What this does not tell you
- Deterministic is not the same as good. A chain that makes the same mistake every time is deterministic too.
- Our check is small. 3 mixes, default settings, two renders each, one machine.
- Same file means same machine and same engine. Another CPU or library build can round the last bit differently, and a library update underneath can change the output without moving the engine ID.
How we measured, and every setting of the chain
Method. What we describe: the chain as it ships, read from the engine's source and its generated parameter sheet (engine config 74340ab1…). Every parameter value below is copied from that sheet or from the engine source, not from memory. What we measured: 3 unmastered test mixes (one each from techno, minimal and tech house), each mastered twice through the shipped automatic chain at default settings, on one machine, both renders compared by SHA-256 hash. Measurement tool: our analyzer, which computes integrated loudness and true peak per ITU-R BS.1770 with libebur128. The script content-social/drafts/scripts/I14_determinism.py prints every number in the table; I14_figures.py draws the figures.
Determinism check. All three pairs came back identical.
| Mix | Renders | Same bytes | SHA-256 (first 12) | Delivered format | Loudness | True peak |
|---|---|---|---|---|---|---|
| 1 (techno) | 2 | yes | b4f823703b9b | 16-bit, 44.1 kHz | -8.98 LUFS | -0.29 dBTP |
| 2 (minimal) | 2 | yes | 611e1cd605eb | 16-bit, 48 kHz | -9.22 LUFS | -0.35 dBTP |
| 3 (tech house) | 2 | yes | 684e2e4beedd | 16-bit, 44.1 kHz | -8.25 LUFS | -0.11 dBTP |
The last row of the second figure is not a render: it is the raw noise draw of the synthetic reference (60 s at 44.1 kHz, NumPy default_rng), hashed by I14_figures.py with seed 0 twice, then with seeds 1 and 2 standing in for a process that picks a new seed each run.
What "deterministic" rests on, in full.
- No randomness between runs: the synthetic reference's noise comes from seed
0, and the same seed gives the same noise every time. - No trained model: nothing in the chain runs a network or weights learned from data. The numbers that do come from data, a set of tonal curves and loudness targets, were computed once, offline, from a library of club records and are checked into the source as a small file.
- No hidden state: the engine changes only when we change its source. Every declared parameter that can move a sample is hashed into one engine ID, and the file of target curves carries a hash of its own, so a parameter change shows up as a new ID.
1. Gates before any processing.
- Length: up to 900 seconds (15 minutes), for the track and for a reference.
- A reference you upload is screened for lossy encoding. The engine reads its first 10 seconds and compares the energy above 16 kHz with the energy between 10 and 16 kHz. Below a ratio of 0.05 it looks suspicious, and the engine then checks for a cliff: if any single 1 kHz band between 6 and 21 kHz drops 15 dB or more below its neighbour, the file is rejected as a transcoded lossy file. A master that simply rolls off gently passes.
- That screen is a heuristic, and it errs both ways. In our own tests it refused some genuinely lossless files (a mastering low-pass near 19 to 20 kHz, or a narrow peak in the top end, reads as a codec cliff) and let some 320 kbps MP3s through, whose 20 kHz low-pass has the same shape. A spectrum cannot prove a file lossless; what it can catch is a cut-off low enough to make matching dull the top end. The check is under review.
- A reference must be at least -18 LUFS. Matching a quiet reference would chase a quiet target.
- Sample rate: the whole chain runs at 48 kHz if your file is 48 kHz or higher, and at 44.1 kHz otherwise. Higher rates are brought down to 48 kHz because the EQ analysis uses a fixed 4096-point FFT, and at 96 kHz each frequency bin would be twice as wide as the one the target curves were built at.
Your own track gets the length check and nothing else.
2. The target. With a reference of your own, the target is that file, and the dials below are not used on this path. Without one:
- It measures nine features of your track that describe the shape of the sound and none of its level: how the energy splits across sub, bass, mids and highs, how bright the spectrum is, how much of the signal is tonal versus percussive, and how dense the onsets are. Loudness, true peak and dynamics are left out on purpose, because those are exactly what mastering is about to change.
- It compares them with four fixed centres and picks the closest. Each centre comes with a tonal curve and a loudness target, derived once from the reference library (the loudness of that library is published in Club loudness, measured).
- It decides how far to move your tone. The default, Shaping "Light", looks at your tonal curve frequency by frequency against the range the middle 80% of that centre's records cover (10th to 90th percentile). Where your track is inside that range, nothing moves. Where it is outside, the target moves 75% of the way to the nearest edge of the range. "Full" moves your whole curve all the way to the centre's median curve. "None" keeps your own tonal curve as the target, so only the level changes.
- It applies your dials, if you moved any: Loudness shifts the target by 1.5 LU per step; Top end and Low end bend the target curve by 2 dB per step, as a smooth shelf centred on 4 kHz and on 250 Hz, with a transition one octave wide.
- It keeps your width. The stereo part of the target (the Side channel's curve) is your own track's, not the library's.
- It renders that target as sound: 60 seconds of stereo noise, from seed 0, filtered so its average spectrum matches the target curve, then set to the target loudness.
Why noise works: the matching stage reads only a reference's average spectrum (for the Mid and the Side channels), the level of its loudest passages and its peak. Rhythm, arrangement and phase are averaged away before they reach any decision.
3. Reference matching. Matchering 2.0 is a GPLv3 library, run on our servers unchanged except for one swap. In order, it:
- Cuts both files into pieces of up to 15 seconds and keeps the loudest ones, so a quiet intro does not decide the level.
- Matches the RMS level of your loudest pieces to the reference's.
- Computes the average spectrum of those pieces for Mid and for Side, separately, with a 4096-point FFT, and divides reference by track. The result, smoothed with LOWESS, is the EQ curve: a linear-phase filter that boosts where the reference has more and cuts where it has less.
- Re-checks the level up to 4 times after the EQ and corrects it.
- Runs its own limiter, which catches sample peaks only (at about -0.016 dBFS) and is therefore not the last word on peaks.
The swap: Matchering smooths the EQ curve on a logarithmically spaced frequency grid; we smooth it on a grid spaced by the ERB scale (equivalent rectangular bandwidth). Compared with the logarithmic grid, it is coarser below roughly 400 Hz and finer above. The output of this stage is a 24-bit intermediate file.
4. Stereo safety. Matching Mid and Side separately can widen a track more than its mix did. Three steps undo or guard against that:
- Your width comes back: above 150 Hz, the Side-to-Mid level of the result is set back to what your uploaded mix had.
- Mono bass: below 150 Hz the Side channel is removed, through a linear-phase crossover (a 1001-tap filter).
- A correlation throttle: above 150 Hz, the correlation between left and right is tracked over 0.4-second windows. If it falls below +0.2, the Side channel is turned down in proportion, reaching mono at 0.0.
5. Makeup gain. Removing Side below 150 Hz removes energy, typically 1.1 to 1.6 LU, which nothing after this point would give back. So the engine measures, on this track, the loudness gap between the reference and the current result, and adds 95% of that gap as plain gain. The gain is never negative and never more than 6 dB. It is applied right before the limiter, so the limiter sees it and clamps it.
6. True-peak limiter. FFmpeg's alimiter running on a copy of the signal oversampled 8 times (352.8 kHz for a 44.1 kHz track, 384 kHz for 48 kHz), so it sees the peaks that fall between samples.
- Ceiling: -0.15 dBTP, measured as true peak. This is our setting, not an FFmpeg default. It is a target with a small tolerance, not a guarantee to the hundredth of a dB: in our three renders, two landed under it and one measured -0.11 dBTP, 0.04 dB over and still under 0 dBTP. The setpoint below was tuned on 18 other mixes, the renders behind Club loudness, measured; this mix was not one of them.
- Setpoint: the limiter is aimed 0.25 dB under the ceiling. Its own peak estimate reads optimistic against a BS.1770 true-peak meter, by 0.11 dB on average and 0.24 dB at worst in our tests, and that margin is what absorbs it.
- No levelling: the limiter's automatic level control is off, so it cannot move the loudness the earlier stages set.
- Release: its sustained-limiting compensation is on, which reduces audible pumping when it works hard.
7. 16-bit delivery. A 16-bit PCM WAV, at the rate the chain ran at (44.1 or 48 kHz). The stages before it work at 24-bit or in floating point; the drop to 16 bits happens once, at the very end. We used to deliver 24-bit, until a DJ reported that their CDJs would not load those files. We first blamed the bit depth. The manufacturers' manuals say otherwise: every CDJ and XDJ model we checked lists 24-bit WAV at 44.1 and 48 kHz. The more likely cause is the header: above 16 bit, the tool we encode with writes the WAV's format tag as the "extensible" variant, and DJs report some Pioneer players refusing that tag. That is not settled until the reporting DJ's deck is tested with both headers. Delivery stays 16-bit either way, and a 16-bit file at 44.1 or 48 kHz carries the plain header. The full story is in Why are our mastered files 16-bit WAV and not 24-bit?. 16 bits leave about 96 dB between full scale and the quantisation floor (6.02 dB per bit), more than any club system or club master uses. The conversion is a plain requantisation: no dither and no noise shaping is added, which is the FFmpeg resampler's default.
What the engine does not do, in full. The stages that change the sound are gain, linear-phase EQ, the Side-level and mono-bass stages above, and two limiters. We measured a compressor stage as a way to give "density" its own dial; it moved loudness range up on some tracks and down on others, and was dropped. The engine moves the overall balance and level toward a target; it does not hear a clashing bassline or a timing problem. It aims at the reference's (or target's) loudness, but a track with strong peaks can come back quieter than the target, because reaching it would mean shaving more peak than the ceiling allows. A different result for the same track and settings needs a change to the engine itself.
More limits.
- Our internal gate renders across more parameter states, but those runs are not what this note reports.
- Our production renders all run from one fixed container image; a render on another machine can differ from it in the least significant bit even when the chain is the same.
- A parameter change moves the engine ID; a change to the code or to a library underneath (Matchering, FFmpeg, NumPy) can change the output without moving it. "Same settings" includes the engine version.
- A 44.1 kHz file and a 48 kHz file of the same music are different inputs. The synthetic reference is rendered at the chain's rate, so its noise is a different (seeded, repeatable) realisation at each rate.
- Nothing here is a listening test. Every number above is a measurement; none of it says how a master sounds on your system.
Related
- Why are our mastered files 16-bit WAV and not 24-bit?
- What does mastering actually change in a track?
- Which audio files does a CDJ actually load?
- How is loudness measured, and why is LUFS not dB?
- How does a model like Demucs pull a voice out of a finished track?
- Club loudness, measured
Sources
- ITU-R BS.1770, Algorithms to measure audio programme loudness and true-peak audio level. https://www.itu.int/rec/R-REC-BS.1770
- Matchering 2.0.6, sergree, GPLv3 (source of the matching stage: piece analysis, RMS matching, FIR EQ curve, sample-peak limiter). https://github.com/sergree/matchering
- FFmpeg Filters Documentation,
alimiterandaresample. https://ffmpeg.org/ffmpeg-filters.html#alimiter - FFmpeg Resampler Documentation,
dither_method(default: none). https://ffmpeg.org/ffmpeg-resampler.html - B. R. Glasberg and B. C. J. Moore, "Derivation of auditory filter shapes from notched-noise data", Hearing Research 47 (1990), 103-138 (the ERB scale). https://doi.org/10.1016/0378-5955(90)90170-T
- W. S. Cleveland, "Robust Locally Weighted Regression and Smoothing Scatterplots", Journal of the American Statistical Association 74 (1979), 829-836 (the LOWESS smoothing Matchering applies to the EQ curve). https://doi.org/10.1080/01621459.1979.10481038
- libebur128 (implementation of EBU R 128 / ITU-R BS.1770 used for every measurement here). https://github.com/jiixyj/libebur128
We build enguetedl.ch; the mastering engine and the analyzer above are ours.