What does a 320 kbps MP3 throw away compared to the lossless file?
Formats · 3 min read
A 320 kbps MP3 made with LAME throws away everything above about 20 kHz and stores the rest a little wrong, on purpose. In our 10-track test the loudness did not move, and the true peak rose by a median 0.35 dB.
What it does not do at 320 kbps is cut at 16 kHz. That number belongs to lower bitrates.
A wall at 20 kHz
The MP3 standard fixes only the decoder, so this is one encoder: LAME 3.100, the one inside ffmpeg and most free tools. Its source code picks a lowpass from the bitrate: 20,500 Hz at 320 kbps, 17,000 Hz at 128. The famous 16 kHz cutoff sits between 112 and 128 kbps.
Little music lives up there: above 16 kHz these sources carry a median 0.016% of their power.
Below the wall: mostly precision
Little is missing here. The last kilohertz below it can thin: by 5.5 dB on one of our tracks, a quarter of a dB at the median. The rest is precision: the encoder estimates which detail the ear cannot hear next to louder sound nearby (masking) and stores each band only as precisely as it needs to hide its error there.
That error, source minus decode, reads a median 29 LU below the music, and 79% of its energy falls between 4 and 20 kHz, where the music has only 2.3% of its own.
Loudness stays, peaks move
Loudness moved at most 0.01 LU and the stereo width did not narrow: a loudness report cannot tell the two files apart.
The peaks can: the decode is the source plus that error, so on a track already at the ceiling, anything added moves the peaks.
A float decoder keeps those overs; a player that outputs fixed-point samples has nowhere to put them and clips. For the whole corpus, see Club loudness, measured.
Our own check
Our mastering step has a check meant to turn away lossy reference tracks. It is known to be imperfect and is under review: on these files it let 4 of 9 320 decodes through and rejected 1 of the 10 lossless sources.
What this does not tell you
- Whether anyone hears it. No listening test here; Does WAV sound better than a 320 MP3 on a club system? collects the blind tests.
- Other encoders, other material. One encoder, one decoder, 10 club tracks at 44.1 kHz. With n = 10 the ranges matter more than the medians.
- What happens inside a frame. Pre-echo, per-frame stereo switching and smeared transients average away over a whole track.
How we measured the round trip, with every number
Method.
- Files: 10 lossless club tracks from our reference library, one from each of 10 genre folders, the first file of each folder in alphabetical order (no picking by ear or by number). All FLAC, 44.1 kHz, 16-bit, 4.8 to 9.2 minutes long.
- Encoding: ffmpeg 6.1.1 with libmp3lame (LAME 3.100). Two settings:
-b:a 320k(constant bitrate 320 kbps, LAME's default joint stereo and default lowpass) and-q:a 0(variable bitrate, "V0"). - Decoding: ffmpeg's MP3 decoder to 32-bit float WAV, so nothing above full scale is clipped on the way back.
- Loudness and peaks: our analyzer, built on libebur128 (ITU-R BS.1770): integrated loudness, true peak, sample peak, plus its spectral rolloff, high-band share and side/mid balance. Source and both decodes, one file at a time.
- Spectrum: a long-term average spectrum (Welch, 8192-point Hann window, both channels averaged) of source and decode. The cutoff is the first 100 Hz band above 10 kHz where the decode keeps less than 1% of the source's power (-20 dB). This is a spectral fact, not a loudness reading, so it does not go through the loudness meter. The first two figures are the medians of the same spectra over the 10 tracks (100 Hz bands for the first, third-octave bands against each track's total power for the second); the shaded area in the first is the full range.
- Residual: source minus decode, sample by sample (the decodes came back at exactly the source length on all 10), measured with the same analyzer.
- n = 10. One encoder, one decoder, one sample rate. The script
content-social/drafts/scripts/I5_mp3_roundtrip.pyprints every number below;I5_figures.pydraws the figures.
The standard. MPEG-1 Audio Layer III is ISO/IEC 11172-3. It fixes the bitstream and the decoder: "The encoder algorithm is not standardized, and may use various means for encoding such as estimation of the auditory masking threshold, quantization, and scaling". Its psychoacoustic models sit in an annex that is "for information only".
How the encoder chooses. Layer III splits the signal in two stages: 32 equal subbands, each followed by an MDCT, which gives 576 frequency lines per block. At 44.1 kHz that is about 38 Hz per line. For transients it switches to short blocks, so quantisation noise does not smear audibly ahead of a drum hit (pre-echo). Then it spends a fixed budget of bits. In Painter and Spanias's words, "Masking refers to a process where one sound is rendered inaudible because of the presence of another sound". The encoder estimates that masking threshold and quantises each band coarsely enough to save bits, but finely enough to keep the error under the threshold, as far as the budget allows.
LAME's default lowpass by bitrate, from lame.c:
| Bitrate (kbps) | 96 | 112 | 128 | 160 | 192 | 224 | 256 | 320 |
|---|---|---|---|---|---|---|---|---|
| LAME default lowpass (Hz) | 15,100 | 15,600 | 17,000 | 17,500 | 18,600 | 19,400 | 19,700 | 20,500 |
LAME's own manual says why it filters at all: it "helps reducing the amount of data to encode. This is important in MP3 due to a limitation in very high frequencies (>16Khz)".
Change in each 1 kHz band, decode minus source:
| Band | 320 CBR, median (range), dB | V0, median (range), dB |
|---|---|---|
| 15-16 kHz | +0.08 (+0.01 to +0.10) | +0.08 (+0.01 to +0.14) |
| 16-17 kHz | +0.29 (+0.04 to +0.43) | +0.19 (-0.51 to +0.46) |
| 19-20 kHz | -0.24 (-5.47 to +0.05) | -0.21 (-6.82 to +0.41) |
| 20-21 kHz | -16.94 (-26.58 to -12.86) | -0.71 (-7.85 to +0.30) |
| 21-22 kHz | -47.98 (-65.23 to -31.78) | -2.90 (-15.04 to +0.06) |
At 320 kbps the decode drops below 1% of the source's power at 20.1 to 20.2 kHz on all 10 tracks (median 20.15 kHz), about 400 Hz below LAME's nominal 20,500. Above 16 kHz the sources carry a median 0.016% of their total power (0.001% to 0.059%); above 19.5 kHz, a median 0.0014%.
The 16-17 kHz row is the other half of LAME's sentence. At 44.1 kHz the last scalefactor band of a long block covers the lines from 418 to 575 in LAME's band table, which is about 16.0 kHz to the top. LAME's manual says that keeping precision there means raising precision "in all the bands (not only in this one)", and that constant-bitrate mode simply ignores the noise in that band. The +0.29 dB is that noise: above 16 kHz the 320 file does not lose content, it gains a little error.
Residual. Measured as loudness, it sits a median 29.1 LU below the music (24.8 to 38.4 LU). A median 79% of its energy falls between 4 and 20 kHz (63% to 90%), where the music itself carries a median 2.3% of its energy (0.3% to 4.8%). A residual 29 LU down is not "29 LU of inaudible": it is quiet next to the music and shaped to hide under it.
Per track, 320 CBR, labelled by genre folder only:
| Track (genre folder) | Loudness change (LU) | True peak, source (dBTP) | True peak, 320 decode (dBTP) | Change (dB) | Samples above full scale in the decode | 20 kHz wall at (kHz) | Residual below music (LU) |
|---|---|---|---|---|---|---|---|
| deep-house | 0.00 | +1.00 | +1.02 | +0.02 | 31,744 | 20.1 | 30.0 |
| deep-organic | 0.00 | -0.36 | +0.59 | +0.95 | 50 | 20.2 | 25.8 |
| groove | 0.00 | -0.15 | +0.57 | +0.72 | 48 | 20.1 | 28.5 |
| hard-techno | 0.00 | 0.00 | +0.43 | +0.43 | 2,468 | 20.1 | 29.7 |
| house | +0.01 | +1.48 | +1.75 | +0.27 | 22,106 | 20.1 | 24.9 |
| hypnotic | -0.01 | -0.31 | -0.19 | +0.12 | 0 | 20.2 | 38.4 |
| melodic | 0.00 | -0.11 | +0.53 | +0.64 | 103 | 20.2 | 26.8 |
| minimal | 0.00 | +0.52 | +0.49 | -0.03 | 18,851 | 20.2 | 37.7 |
| tech-house | +0.01 | +0.01 | +1.38 | +1.37 | 1,594 | 20.1 | 24.8 |
| techno | 0.00 | +0.61 | +0.48 | -0.13 | 301 | 20.2 | 32.3 |
- True peak rose by a median +0.35 dB (-0.13 to +1.37).
- Five sources measured above 0.0 dBTP. After the round trip, nine did.
- Sample peak rose by a median +0.71 dB (+0.11 to +1.57). Nine of ten decodes contain samples above full scale, a median of 948 per track. The FLAC sources cannot: one of them had 2,102 samples sitting exactly at full scale, the others none.
- On two tracks true peak fell slightly. Taking out everything above 20 kHz can lower a peak as well as raise one.
What the analyzer's report sees. Integrated loudness changed by at most 0.01 LU in either direction. The spectral rolloff (the frequency below which 85% of the energy sits, a median 6.7 kHz on these sources) moved a median 37 Hz. The high-band share (4-20 kHz) did not change at one decimal place.
Stereo. By default LAME uses joint stereo: "the encoder can use (on a frame by frame basis) either L/R stereo or mid/side stereo", giving more bits to the mid channel. Over whole tracks the side-to-mid balance moved a median +0.01 dB (0.00 to +0.03) at 320 kbps. That is a long-term average; it says nothing about a single frame.
V0. It averaged a median 275.6 kbps (255.7 to 281.4). LAME's default variable-bitrate mode sets its V0 lowpass at 24 kHz and then caps it at half the sample rate; at 44.1 kHz that is no lowpass at all. The 20-21 kHz band lost a median 0.71 dB, and a -20 dB point existed on only 2 of 10 tracks (21.1 and 21.9 kHz). V0 still thins the very top on some tracks (21-22 kHz down as much as 15 dB), but it has no fixed wall. True peak rose a median +0.45 dB (0.00 to +1.29); nine of ten decodes measured above 0.0 dBTP and held samples above full scale (median 1,574). A smaller file with more top end is a trade-off of LAME's settings, not of MP3.
Our reference check. It is known to be imperfect and under review: beyond these 10 tracks it also turns away some genuinely lossless files, and a fix is pending. As it stands it runs on the first 10 seconds:
- If the energy above 16 kHz is less than 5% of the energy between 10 and 16 kHz, the file is a suspect.
- A suspect is cleared if no 1 kHz band between 6 and 21 kHz sits 15 dB or more below the band before it, the sign of a gentle natural roll-off rather than an encoder's wall.
| Sources | 320 CBR decodes | V0 decodes | |
|---|---|---|---|
| Rejected as lossy | 1 of 10 | 6 of 10 | 2 of 10 |
| Rejected, of the 9 whose source passed | 5 of 9 | 1 of 9 | |
| Steepest band drop, median (range) | 5.2 dB (2.9 to 20.4) | 19.7 dB (16.1 to 25.1) | 6.2 dB (2.9 to 20.3) |
The second step does its job: every 320 decode shows a drop of at least 16 dB, at the 20 kHz wall. But it is only consulted when the first step flags the file, and four 320 decodes had enough energy between 16 and 20 kHz in their first 10 seconds to pass step one. The one source it rejected (deep-house) is not an MP3 in disguise: the 320 encode removed 26.6 dB of its 20-21 kHz band, so that content was really there. Its first 10 seconds happen to have a steep step in the treble. A false positive.
More limits.
- Re-encoding. An MP3 made from an MP3 was not tested.
- A hardware true peak. BS.1770's estimate, not a real converter.
- Our check, in general. 10 tracks and 20 decodes is a sample, not a validation.
Limits, in full. Another encoder picks its own lowpass and its own masking model; the standard allows that. At 48 kHz the scalefactor bands sit at different frequencies. On re-encoding, the literature does not expect a second pass to respect the first one's masking: transcoding "is neither guaranteed nor likely to preserve perceptual noise masking". True peak here is the BS.1770 estimate (libebur128, 4x oversampling below 96 kHz); a real converter's reconstruction filter can land slightly elsewhere.
Sources
- ISO/IEC 11172-3:1993, Information technology - Coding of moving pictures and associated audio for digital storage media at up to about 1,5 Mbit/s - Part 3: Audio. Introduction 0.1 (encoder not standardised) and the annex list (C to H "for information only"). https://www.iso.org/standard/22412.html
- T. Painter, A. Spanias, "Perceptual Coding of Digital Audio", Proceedings of the IEEE, vol. 88, no. 4, pp. 451-515, April 2000. Section II-C (masking), Section VIII-A (MPEG-1 Layer III: hybrid filter bank, MDCT, pre-echo, transcoding). https://doi.org/10.1109/5.842996
- LAME 3.100 source,
libmp3lame/lame.c:optimum_bandwidth()(bitrate to lowpass table) and the variable-bitrate lowpass inlame_init_params(). https://sourceforge.net/p/lame/svn/HEAD/tree/tags/RELEASE__3_100/lame/libmp3lame/lame.c - LAME 3.100,
USAGE(sections "Low pass filter", "Modes", "Ignore scalefactor band 21"). https://sourceforge.net/p/lame/svn/HEAD/tree/tags/RELEASE__3_100/lame/USAGE - LAME 3.100 source,
libmp3lame/quantize_pvt.c,sfBandIndex(the Layer III scalefactor band table, ISO/IEC 11172-3 Table B.8). https://sourceforge.net/p/lame/svn/HEAD/tree/tags/RELEASE__3_100/lame/libmp3lame/quantize_pvt.c - LAME 3.100,
include/lame.h(vbr_default=vbr_mtrh). https://sourceforge.net/p/lame/svn/HEAD/tree/tags/RELEASE__3_100/lame/include/lame.h - enguetedl.ch mastering engine, reference check (
check_reference), thresholds as declared in the engine's settings. - ITU-R BS.1770, Algorithms to measure audio programme loudness and true-peak audio level. https://www.itu.int/rec/R-REC-BS.1770 ; libebur128, https://github.com/jiixyj/libebur128
Related
- Does WAV sound better than a 320 MP3 on a club system?
- How is loudness measured, and why is LUFS not dB?
- Which audio files does a CDJ actually load?
- Club loudness, measured
We build enguetedl.ch; the analyzer and the reference check above are ours.