Stem separate your own music
23 Feb, 2026
Disclaimer: in the context of this post, I'll use 'tracks', 'stems', and 'instruments' quite often.
To clarify, for me 'tracks' are the "architectural" concept of tracks you find in your hardware or software (e.g. your DAW tracks).
A 'stem' represents a conceptual group of tracks, semantically related to some purpose in the song (e.g. 'Drums' stem is composed by all drum-related tracks).
Finally, 'instruments' are obviously the concept of different instruments (either programmed or humanly played) used in the song.
TL;DR
Do you record your music from your hardware setup in stereo, but you're desperate to go back to adjusting your mixes once they're already done and recorded? I propose a simple, limited but effective, even if unusual approach (just jump to the "Stem Separation part").
Context
DAWs have "infinite" tracks
In modern DAWs, usually, tracks count is higher than the number of instruments you use. Given a certain number of instruments in your song, you are usually able (and tend to) create more tracks, which is ultimately leading you to use something called a "bus" to create a submix to handle stems. For example, considering Vocals, Drums, Bass, and Guitar you might have:
- Vocals (Main track)
- Vocals (Backing track)
- Vocals (Room reverb)
- Drums (Dry)
- Drums (Room reverb)
- Guitar (Take 1)
- Guitar (Take 2)
- Guitar (Take 3)
- Bass (Take 1)
- Bass (Take 2)
- etc ...
This translates to controlling your mix with high degrees of freedom. You control the individual tracks, and you also control the submixes. You're able to assign volumes, dynamics, track interaction freely. It's like having purchased all the ingredients you need to bake the cake, having them ready at your disposal.
Hardware doesn't
In my personal approach to music, which is also common to many other "DAWless" approaches, the case is very different. I work with Nanoloop2, which is a limited software running on limited hardware (the GBA isn't exactly made for music in the first place), which means that track count is limited, to FOUR.
Despite the limited track count, you can still make articulate songs using lots of instruments, because you'd usually be allowed to place different instruments in the same track, which is a sort of 'multiplexing'. You can put the kick and the bass together, the implied limitation being that you can't have both playing together as they share the track, but that is a feasible compromise in most of the music you expect to make with such devices, and for the case of chiptune it's also something acknowledged as the usual practice.
Balancing efforts and results is a struggle
When it comes to recording your music, this limitation typical of using hardware devices can become painful: except for certain (costly) devices, you're most likely going to have only a stereo (mix) output. If you want to record tracks you hit a process wall: you'll still need to record your stereo mix out, having to manually solo tracks and record each of them, as many times as the track count. Big oof.
Moreover, if you have a "multiplexed use" of your tracks, it gets worse. You won't be able to extract the 'bass' only or 'drums' only stems easily: you'll have to somehow mute individual instruments. In the case of nanoloop2 (picking up an extrema), there's no way to mute "instruments" (there is no concept of instruments at all, only notes) so I'd have to remove the sequenced notes individually to isolate instruments and make stems. On top of the annoyment, this isn't even a good solution because it alters the way the track would normally play! If I remove kick notes that are in a track, the co-interaction of kick and bass results in a different sounding arrangement (e.g. some bass notes won't be stopped if I remove the kick, unless I add manually stop notes). Not good.
Finally, if you're trying to record something that you are playing in a jam or improvisation session, you won't have any possibility to go back in time and record stem tracks exactly the way you have initially jammed. It'd have to be right the first time, straight from the hardware, and that means having hardware that somehow outputs multitracks.
Now, my main motivations are two:
- Lazyness: if I'm recording something I don't want my process to be long, painful and repetitive. This could of course be different for people who strive for perfection and total control in their process but that is not me. Usually what I'm trying to do is to recreate in the recording the perception of when I play live. That doesn't require me to record multiple takes of hundreds of tracks. I want to record a single jam and make it sound good!
- Cost efficiency: many people get excited at the idea of making hardware music (rather than software) but hardware comes with a cost, so most people will typically have a somehow limited hardware setup (some synths and a mixer, typically). Setting up a multitrack setup requires not only money and planning, but also forces you to choose certain devices, altering your workflow, and spend additional time on your computer when you want to record, which is frustrating.
So let's say we want to record a simple jam in stereo, but we also want to have the ability to adjust it in post processing, with a (comparatively) larger freedom that usually doesn't come with "just" editing your stereo mix. You'd like to readjust the stems in your mix, after the mix is already done in hardware.
Stem separation
And the solution, for us living in the big '26 is clearly there: just stem separate it!
Now here is where most audio engineers or anyone remotely interested in "aiming for perfection", are going to close the webpage or punch me through the computer screen (ouch). Yea, this wasn't for you guys, I told you.
You heard me right, though. Everyone is doing stem separation on other people's songs to make instrumental tracks and acapellas, DJ stems, sampling, you name it. Nobody is stem separating their own music. And why would they? If it's their own music, that doesn't give any advantage because it implies you'd have all the tracks already. Except YOU DON'T, if you recorded a stereo mix as in the scensarios cited above!
Most audio software includes stem separators nowadays, and they are quite powerful and precise. You find them in most DAWs. If you don't use DAWs, you also find one for free in Audacity new versions.
They usually are built on top of pre-existing neural network models (like Spleeter, Demucs, etc.) and are probably either vanilla versions or slightly re-trained/fine-tuned. They will usually extract four stems: vocals, drums, bass, other instruments. That is obviously nowhere near reconstructing tracks but we're already dealing with a powerful trick that we can take advantage of.
Note obviously that these tools, depending on the songs you feed in, will sometimes struggle. I found it especially hard to decouple bass and kick on some of my (techno) tracks, but they still do a preposterous job.
Parallel mixing with stems
Now, if you are a sound engineer who survived reading up to here, you'll definetely quit now, because what is coming might be even more horrifying.
My suggestion is: keep your original stereo track intact but lower it in volume. This ensures what you recorded stays present and consistent, and nothing is lost in the process: it's very likely the separated stems have lost some detail in the high frequenc range, due to forms of resasmpling being applied internally to the neural networks, which means if you simply mix your stems and remove the original, you inherently lost quality.
Instead, use each stem as parallel submix bus that reinforces the original stereo track in different ways (i.e. performing different mixing choices).
What you need to be careful about is making sure that stems are phase aligned with your original track. However, when using the stem separator in FL Studio, I noticed the tracks created were already phase aligned to the original.
In my personal experience, (with nanoloop2) I've:
- Bass: applied a slight distortion and EQ to the bass stem in order to move it up in the spectrum, leaving more room to the kick.
- Drums: muffled the mids while keeping highs and lows to enhance kicks, hats and the high end of snares (sort of a "modern" EQ). Applied a powerful compression sidechaining other tracks.
- Other instruments: I've slightly EQ'd to remove the low end, and split the stem in three tracks, slightly panned/delayed/reverbered in order to get additional stereo width.
- Vocals: None present.
While I am perfectly conscious that this is a far cry from actual sound engineering done right, the results for me have been acceptable and worked wonders considering the level of effort I wanted to put in my recording and publishing session.
You can listen to the A vs. B result here:
And here is the final release:
In the end, I felt like this was worth sharing so that other people can take advantage of this idea. It's highly possible that this approach has already been considered by others but I just decided to post it, considering I came up with it independently.
You should resample your stems
This is obviously an opening to a whole other topic but, remember the inaccuracies I mentioned you get with stem separation? I noticed that, for example, these inaccuracies mean that non-drums sounds might "bleed" onto the drum stem, creating interesting and well-glued grooves.
This offers an interesting possibility of resampling your stems creating interesting layering variations on your original sound!
