Any short recording is exactly the sum of sinusoids at its DFT frequencies. Add them in from the lowest frequency upward, and listen to the sound assemble itself.
For a real signal, bin k and bin N−k of the DFT together make one real cosine of amplitude 2|X[k]|/N, so keeping bins 0…K is a sum of K+1 sinusoids — an ideal lowpass at k·Fs/N. The reconstruction here is computed with a masked inverse FFT, which is the same sum evaluated efficiently. The cutoff slider is spaced in mel, so it moves through the frequency range roughly the way hearing does. Audio never leaves your browser.