stardance has been extended another month! the new deadline is october 31 :)

You are browsing as a guest. Sign up (or log in) to start making projects!

ReMix

  • 4 Devlogs
  • 26 Total hours

An audio spatialiser that makes music much more immersive by placing stems around the listener in 3D.

Ship #1 Pending review

The concept came to me when I realised how dramatic the difference between stereo and mono audio was. How the intruments separated, opened up, and details could he heard really well, while mono sounded like noise. I thought about what would happen if instruments were literally separated and opened up, as if placed around the listener like a band. ReMix separates the stems of a song and spatialises them individually to create perception of a room and placement of instruments around you.

Demo of output songs. (listen with headphones)

What was challenging

  • While the code itself was moderately challenging, understanding all the different psychoacoustic cues (covered in more detail in README and devlogs), how they work and what they do, and the maths behind them, required a ton of deliberation. Though it was very rewarding and interesting.
  • I also spent a lot of time trying to analyse the HRIR curves themselves, trying to make sense of what frequencies getting altered and where and why. I used a lot of graphs and intuition to try to understand them. I now have a deep understanding of them, which would be helpful for future versions, for optimising them (which I also attempted, though naïvely).
  • Near the end, making the audio sound natural was challenging too where I had to try tweaking some things (mainly room absorbence/size and reverb)

What I’m proud of

  • I love that the concept I had thought of in theory actually works. When I first played a song, I couldn’t stop grinning because of how good it sounded, and the trippy illusion of it coming from around me. I’ve listened to 8D audio (a simpler version of this where the audio as a whole moves, less realistic, no band-like feeling, distracting and somewhat gimmicky), but it coming from something I created myself, and using code and maths; not some software. And something I could actually experience, I think that makes me very proud. I also had my friends listen to it, which got a lot of positive reactions. In the end, I’m also happy about how deep of an understanding I have about psychoacoustics now.

What you need to know to test it out

  • headphones are a must pls the effect breaks completely on speakers.
  • You can try it out from the repo, but here’s a demo
  • The song spatialisation concept is a demo for the concepts behind it, which can be used in a lot of applications like video games, movies, producers who can’t afford a studio, etc.
  • This is a very initial version. Updates may make the effects even more realistic, or optimise the program to make it suitable for running in real time (or applied on hundreds of elements for games).
  • 4 devlogs
  • 26h
Try project → See source code →
Open comments for this post

1h 3m 19s logged

[ship log!] It’s moments like this for which I’m alive. This has literally brought me to tears of happiness 🥹. Used a stem separator (htdemucs_6s by Meta research) to separate a song into it’s parts (guitar, drums, piano, other, bass, vocals) and spatialises them individually, placing the stems around the listener to make the song sound more open; as if the members of a band are playing around the listener. It worked EXCEPTIONALLY well, far beyond what I expected, even for songs I used to stress test it which were almost guaranteed to sound bad, they sounded great. The spatialiser asks for song file path, room size, and distance of source. The materials of the room can be altered by changing the alpha coefficients. Here are a few songs I tried out and their review:

Bill Withers - Ain’t no Sunshine: a very safe song for first try, clear separation of instruments, no stereo widening or spatialisation, no special effects. Genuinely shocked me, made me smile uncontrollably when it played. Sounded perfect, spatialised well (everything was distinguishable). Truly sounded “outside of head” and around me. Exactly the kind of song this was built to fix; flat and closed songs sounding open.
Ye - This a Must: from my fav artist. A song I’ve liked but disliked the closed, flat feeling of vocals in. This didn’t work just as well as others. Lost a bit of thump initially, but reducing distance helped. This also did exactly what I wanted, though it still lacked some thump.
Olivia Rodrigo - stupid song: A song i love because of the layering of instruments and her vocals. I used this to see what would happen in a song with a lot of stereo widening and effects already present. This sounded beautiful here, and I was concerned about it having clashing effects, but that didn’t happen. amazing.
Playboi Carti - POP OUT: I put this in as a joke, there was no way this was supposed to sound tolerable. But it blew me away, and worked really well, dare I say better than the original too?
Others: Jeremy Zucker - i dont know you, Stevie Wonder - Sir Duke, Carti - TOXIC

Demo videos: https://youtu.be/GdZYN0wQXBQ

0
0
11
Open comments for this post

10h 26m 30s logged

This is the most I’ve ever learned in such a short time. Over the past few days I designed a program that spatialises audio using several psychoacoustic cues. I spent many of those hours understanding the maths and how to apply the cues. I’d researched broadly beforehand, so the code itself wasn’t too complex. It’s probably the most interesting project I’ve worked on.

The cues, and what I explored

Falloff: apply_falloff()

How sound loses energy with distance (inverse square law).

Air absorption: apply_eq(), apply_air_absorption()

Air absorbs higher frequencies more than lower ones. I expected this to be minor, but it needed a whole EQ system, whose maths was interesting and took ages to wrap my head around. It also taught me about filters, which I used in later functions.

Changes to one frequency interfered heavily with others, which I fixed with a peaking biquad filter (mostly AI-written). The attenuation constants come from a one-time-run file using the ISO standard, and are multiplied by distance and fed into the EQ.

Early reflections: early_reflections()

Initial reflections from nearby surfaces, ~5-35ms after the direct sound. By the Haas effect the brain doesn’t hear them separately, but they imply spaciousness. I first used IACC and ASW (below), then dropped them for directional cues (HRIR), so each of the 6 reflections comes from its surface’s direction. It accounts for room size and pressure reflection (from an absorption coefficient).

Decorrelation: decorrelate()

Artificially decorrelates a signal (FIR comb filters) into L and R, lowering IACC (interaural cross-correlation) and increasing ASW (apparent source width). An IACC of 1 sounds like a point source; lower sounds wider. It worked against my goal (localisation, not a larger source), so I dropped it from the direct audio, then from early reflections. Scrapped, but an interesting concept.

Reverb: room_reverb()

Probably the most important cue for space. Late reverb is thousands of colliding reflections, an unintelligible wash. I built it from feedback comb filters, combined and diffused through an allpass filter, and cut off at RT60 (time to decay 60 dB, from speed of sound, room size and absorption).

Problem: the tail was distinct and ruined localisation (which relies on level differences between ears), like in a hall. Most of the energy making it legible was in its first ~50ms, masking the early reflections. Cutting the first [early reflection length - 20ms] helped a bit. Switching to an FDN made it denser and decorrelated L/R by design, so I could drop decorrelate().

DRR (part of reverb)

Direct-to-reverberant ratio simulates source distance relative to the room: closer to the listener than the walls means louder direct sound, closer to the walls means louder reverb. I set the direct amplitude to 1 and calculated a gain for the reverb.

Apply IR

The only cue for direction. It convolves left and right impulse responses for an azimuth and elevation onto the direct audio and early reflections. They come from a HRIR dataset (.SOFA, SADIE II) recorded on a binaural mannequin. I tried optimising this earlier and failed (older devlogs), and want to revisit it.

Process audio

  1. Converts to mono, resamples to 48kHz if needed.
  2. Gets azimuth, elevation and distance from coordinates.
  3. Applies air absorption, falloff, then IR on the direct sound.
  4. Generates early reflections and late reverb (L and R).
  5. Sums everything, scales to prevent clipping, returns L and R arrays :D

Attaching images of my diary notes because I think it’s super interesting.

0
0
30
Open comments for this post

6h 30m 31s logged

Animating by Azimuth

I tried to identify patterns in the HRTF curves by azimuth. Wrote a function which animates it frame by frame (by azimuth). Visually, I could see it gradually moving (similar shape, moving little by little), though as it was apparent from the heat maps, it wasn’t regular and had phases of very slight changes and phases of chaotic movement.

(This was all done with right ear curves since SADIE II is symmetrical, which made things easier).

To understand this, I wrote another function which compared the similarity of 5 curves at a time, and then plotted that. This had some very dense regions with little variation most of the time, but was quite chaotic around 50-150°.

Peak Analysis Compression

I tried writing a set of functions which would go through the curves, identify significant points (turning points and drops/rises, detected by changes in signs of differences and slope changes) then judge whether the curves of the angles following it are similar (same number of peaks, within a threshold of frequency and magnitude).

Similar curves would be grouped within a “phase” where each phase is a dictionary with:

  • 'initial' (initial coordinates of peaks, list of tuples)
  • 'freq_changes' (nested list of changes in frequency for each frame of each peak)
  • 'mag_changes' (similar to freq_changes)

The changes arrays would then be simplified; if frequency/magnitude changes are subtle enough that it’s imperceptible (determined by continuous bark scale/hardcoded decibel threshold), it would be reduced to zero.

Reconstruction

To reconstruct it, a function would get the phase’s initial, add the corresponding differences to get the exact peaks, then interpolate using pchip to get the curve.

Results

It looked like it could optimise everything significantly, but it failed miserably.

  • Very inaccurate (error ~5dB mean, ~14dB max)
  • About 77% of phases were just 1 curve
  • The max size of a phase was 6 curves (compared to my expected 20-30)

The error could be made better by having more peaks, but it doesn’t optimise anything significantly.

I’ll move on with directly applying the HRTF curves now, though I’ll come back to this later because it looks promising! Attaching the azimuth animation video and the similarity plot

0
0
8
Open comments for this post

7h 59m 11s logged

Spent a few days researching psychoacoustics and sound localisation to understand all the cues (5 of them, will talk more about them later!). Started working with an HRIR dataset (a dataset containing sound measurements from left and right ears of a mannequin from different directions, showing how ear shape and angles affect sound. the single most important cue in sound localisation). First wrote a function to plot raw left and right impulses, then ITD (time difference b/w ears). Then used Fourier transform to break impulses into frequency domain graphs showing magnitudes of particular frequencies, then the ILD (level difference). I was stuck here because the graph didn’t match my expectations at all (a very interesting problem too, check my reddit post: https://www.reddit.com/r/DSP/comments/1weg1dk/why_are_hrirhrtf_magnitude_differences_between/), until a person on reddit explained it, after which I also switched to a symmetrical HRIR dataset (SADIE II instead of Meta SS2). To visualise it better, I also created a heatmap of left/right ears, ILD, and left/right averages. Understanding this through different types of graphs and analysing them, recognising quirks, understanding what they meant was very interesting. Attaching pictures of the heatmaps, raw impulses, ITD, ILD curves :)

0
0
22

Delete project?

Are you sure you want to permanently delete this project? This action cannot be undone.

All devlogs, followers, and associated data will be removed.

Followers

Loading…