ReMix
- 4 Devlogs
- 26 Total hours
An audio spatialiser that makes music much more immersive by placing stems around the listener in 3D.
An audio spatialiser that makes music much more immersive by placing stems around the listener in 3D.
[ship log!] It’s moments like this for which I’m alive. This has literally brought me to tears of happiness 🥹. Used a stem separator (htdemucs_6s by Meta research) to separate a song into it’s parts (guitar, drums, piano, other, bass, vocals) and spatialises them individually, placing the stems around the listener to make the song sound more open; as if the members of a band are playing around the listener. It worked EXCEPTIONALLY well, far beyond what I expected, even for songs I used to stress test it which were almost guaranteed to sound bad, they sounded great. The spatialiser asks for song file path, room size, and distance of source. The materials of the room can be altered by changing the alpha coefficients. Here are a few songs I tried out and their review:
Bill Withers - Ain’t no Sunshine: a very safe song for first try, clear separation of instruments, no stereo widening or spatialisation, no special effects. Genuinely shocked me, made me smile uncontrollably when it played. Sounded perfect, spatialised well (everything was distinguishable). Truly sounded “outside of head” and around me. Exactly the kind of song this was built to fix; flat and closed songs sounding open.
Ye - This a Must: from my fav artist. A song I’ve liked but disliked the closed, flat feeling of vocals in. This didn’t work just as well as others. Lost a bit of thump initially, but reducing distance helped. This also did exactly what I wanted, though it still lacked some thump.
Olivia Rodrigo - stupid song: A song i love because of the layering of instruments and her vocals. I used this to see what would happen in a song with a lot of stereo widening and effects already present. This sounded beautiful here, and I was concerned about it having clashing effects, but that didn’t happen. amazing.
Playboi Carti - POP OUT: I put this in as a joke, there was no way this was supposed to sound tolerable. But it blew me away, and worked really well, dare I say better than the original too?
Others: Jeremy Zucker - i dont know you, Stevie Wonder - Sir Duke, Carti - TOXIC
Demo videos: https://youtu.be/GdZYN0wQXBQ
This is the most I’ve ever learned in such a short time. Over the past few days I designed a program that spatialises audio using several psychoacoustic cues. I spent many of those hours understanding the maths and how to apply the cues. I’d researched broadly beforehand, so the code itself wasn’t too complex. It’s probably the most interesting project I’ve worked on.
apply_falloff()
How sound loses energy with distance (inverse square law).
apply_eq(), apply_air_absorption()
Air absorbs higher frequencies more than lower ones. I expected this to be minor, but it needed a whole EQ system, whose maths was interesting and took ages to wrap my head around. It also taught me about filters, which I used in later functions.
Changes to one frequency interfered heavily with others, which I fixed with a peaking biquad filter (mostly AI-written). The attenuation constants come from a one-time-run file using the ISO standard, and are multiplied by distance and fed into the EQ.
early_reflections()
Initial reflections from nearby surfaces, ~5-35ms after the direct sound. By the Haas effect the brain doesn’t hear them separately, but they imply spaciousness. I first used IACC and ASW (below), then dropped them for directional cues (HRIR), so each of the 6 reflections comes from its surface’s direction. It accounts for room size and pressure reflection (from an absorption coefficient).
decorrelate()
Artificially decorrelates a signal (FIR comb filters) into L and R, lowering IACC (interaural cross-correlation) and increasing ASW (apparent source width). An IACC of 1 sounds like a point source; lower sounds wider. It worked against my goal (localisation, not a larger source), so I dropped it from the direct audio, then from early reflections. Scrapped, but an interesting concept.
room_reverb()
Probably the most important cue for space. Late reverb is thousands of colliding reflections, an unintelligible wash. I built it from feedback comb filters, combined and diffused through an allpass filter, and cut off at RT60 (time to decay 60 dB, from speed of sound, room size and absorption).
Problem: the tail was distinct and ruined localisation (which relies on level differences between ears), like in a hall. Most of the energy making it legible was in its first ~50ms, masking the early reflections. Cutting the first [early reflection length - 20ms] helped a bit. Switching to an FDN made it denser and decorrelated L/R by design, so I could drop decorrelate().
Direct-to-reverberant ratio simulates source distance relative to the room: closer to the listener than the walls means louder direct sound, closer to the walls means louder reverb. I set the direct amplitude to 1 and calculated a gain for the reverb.
The only cue for direction. It convolves left and right impulse responses for an azimuth and elevation onto the direct audio and early reflections. They come from a HRIR dataset (.SOFA, SADIE II) recorded on a binaural mannequin. I tried optimising this earlier and failed (older devlogs), and want to revisit it.
Attaching images of my diary notes because I think it’s super interesting.
I tried to identify patterns in the HRTF curves by azimuth. Wrote a function which animates it frame by frame (by azimuth). Visually, I could see it gradually moving (similar shape, moving little by little), though as it was apparent from the heat maps, it wasn’t regular and had phases of very slight changes and phases of chaotic movement.
(This was all done with right ear curves since SADIE II is symmetrical, which made things easier).
To understand this, I wrote another function which compared the similarity of 5 curves at a time, and then plotted that. This had some very dense regions with little variation most of the time, but was quite chaotic around 50-150°.
I tried writing a set of functions which would go through the curves, identify significant points (turning points and drops/rises, detected by changes in signs of differences and slope changes) then judge whether the curves of the angles following it are similar (same number of peaks, within a threshold of frequency and magnitude).
Similar curves would be grouped within a “phase” where each phase is a dictionary with:
'initial' (initial coordinates of peaks, list of tuples)'freq_changes' (nested list of changes in frequency for each frame of each peak)'mag_changes' (similar to freq_changes)The changes arrays would then be simplified; if frequency/magnitude changes are subtle enough that it’s imperceptible (determined by continuous bark scale/hardcoded decibel threshold), it would be reduced to zero.
To reconstruct it, a function would get the phase’s initial, add the corresponding differences to get the exact peaks, then interpolate using pchip to get the curve.
It looked like it could optimise everything significantly, but it failed miserably.
The error could be made better by having more peaks, but it doesn’t optimise anything significantly.
I’ll move on with directly applying the HRTF curves now, though I’ll come back to this later because it looks promising! Attaching the azimuth animation video and the similarity plot
Spent a few days researching psychoacoustics and sound localisation to understand all the cues (5 of them, will talk more about them later!). Started working with an HRIR dataset (a dataset containing sound measurements from left and right ears of a mannequin from different directions, showing how ear shape and angles affect sound. the single most important cue in sound localisation). First wrote a function to plot raw left and right impulses, then ITD (time difference b/w ears). Then used Fourier transform to break impulses into frequency domain graphs showing magnitudes of particular frequencies, then the ILD (level difference). I was stuck here because the graph didn’t match my expectations at all (a very interesting problem too, check my reddit post: https://www.reddit.com/r/DSP/comments/1weg1dk/why_are_hrirhrtf_magnitude_differences_between/), until a person on reddit explained it, after which I also switched to a symmetrical HRIR dataset (SADIE II instead of Meta SS2). To visualise it better, I also created a heatmap of left/right ears, ILD, and left/right averages. Understanding this through different types of graphs and analysing them, recognising quirks, understanding what they meant was very interesting. Attaching pictures of the heatmaps, raw impulses, ITD, ILD curves :)