All writing

How to remove vocals from a song (or keep them) without uploading it

To remove vocals from a song without uploading it, run a separation model locally: UVR5, audio-separator or VocalDrop. Audacity's trick is weaker.

To remove vocals from a song without uploading it, run a source-separation model on your own computer. The model splits the track into two stems, the vocal and everything else, and you keep whichever you came for: the instrumental for karaoke, the vocal for an acapella. The free ways to do it are UVR5, which bundles the models in a desktop app, and python-audio-separator, which runs the same model files from a terminal. Audacity's older vocal-reduction effect needs no model at all, but it works on a different principle and gives weaker results.

The studio's own tool for this is VocalDrop, a free app for macOS (Apple Silicon) and Windows. It runs one of the same model families as the two free tools above, without a Python setup. This post covers all four options, what each one costs you in setup, and the steps in VocalDrop.

Removing vocals and keeping vocals are one operation

A separation model does not delete the voice from a song. It produces an estimate of the vocal, and the instrumental is what is left. A "vocal remover" and an "acapella extractor" are therefore the same run, saved differently. The difference that matters is which stem the model was trained to get right, which comes up again under bleed below.

The older technique, the one Audacity used to ship, is not a model. It relies on where the vocal sits in the stereo mix, which is why Audacity's own page calls it "typically less reliable".

Option 1: Audacity

The centre-channel trick

Lead vocals are often mixed dead centre, identical in both channels. Subtract one channel from the other and the centre cancels. Audacity's support page describes the manual version: choose Split Stereo to Mono from the track's menu, select one of the two mono tracks, apply Effect → Invert, and play.

The same page is candid about the result. It "will remove everything panned in the center, not just vocals", returns dual mono, and often leaves artefacts, especially where there are backing vocals or reverb.

Audacity also had a dedicated effect, Vocal Reduction and Isolation. Its manual page says it is no longer shipped from Audacity 3.5.0 onwards, and remains available as a downloadable Nyquist plugin. Its Analyze action tells you whether centre removal is likely to work before you wait for it. The documented limits are the useful part:

  • the input must be true stereo, not dual mono;
  • stereo reverb is not fully removed;
  • it removes everything in the centre equally, bass and solo instruments included;
  • for an acapella, the manual warns the result "will probably still have some music in it".

What it costs you: a few minutes, and nothing to install beyond Audacity. It is worth a try on a stereo recording with a dry, centred vocal. Where there are backing vocals, reverb or a centred bass, the limits above apply in full.

The OpenVINO Music Separation plugin

Audacity's support page now leads with an AI option: Intel's OpenVINO plugin, found under Effect → OpenVINO AI Effects → OpenVINO Music Separation. It runs Demucs v4, as its feature documentation states, with a 2-stem mode (instrumental and vocals) and a 4-stem mode, on the CPU, GPU or an Intel NPU.

What it costs you: platform. The support page says it is available only on Windows and Linux. The latest release, v3.7.1-R4.2, ships a Windows installer only, built for Audacity 3.7.1; on Linux you compile it yourself. There is nothing for the Mac. On quality, ZFTurbo's model table scores the stock four-stem HTDemucs4 at 8.24 vocal SDR on the MVSEP Multisong set, against 11.00 for BS-PolarFormer on the same set. I have not checked that the plugin's converted model scores the same as the stock checkpoint.

Option 2: UVR5

Ultimate Vocal Remover is a free desktop app built around these models. Its README says the bundles "contain the UVR interface, Python, PyTorch, and other dependencies", so there are no prerequisites. It is MIT-licensed.

What it costs you:

  • Download size. The v5.6 release lists UVR_v5.6.0_setup.exe at 1.69 GB. The RoFormer models come in later beta builds attached to the same release, the first dated November 2024. The newest Windows one, UVR_1_15_25_22_30_BETA_full.exe, is 2.08 GB, and the matching Apple Silicon dmg is 765 MB.
  • Windows. The README says to install to the main C:\ drive, as "installing UVR to a secondary drive will cause instability". AMD Radeon and Intel Arc owners get a separate DirectML download.
  • macOS. The README warns the app "may take up to 5-10 minutes to start for the first time". If Gatekeeper blocks it, its fix starts with sudo spctl --master-disable, which allows apps from any source system-wide. The README recommends switching that back on once UVR opens, and you should.
  • Choosing a model. UVR5 offers VR, MDX-Net, Demucs and RoFormer models, and the list is long. For vocals, start with BS-Roformer-Viperx-1296; this post on the UVR5 vocal models explains why, and what the scores in the filenames do and do not mean.

Option 3: python-audio-separator

python-audio-separator is a command-line tool that loads the same model files as UVR5 by filename. It is also MIT-licensed, and VocalDrop credits it for inference.

What it costs you: a Python environment and ffmpeg. The project requires Python 3.10 or newer, and its README says you may need to install FFmpeg separately (brew install ffmpeg on macOS). Then:

pip install "audio-separator[cpu]"     # NVIDIA: [gpu]; Windows AMD/Intel: [dml]
audio-separator song.mp3 --model_filename model_bs_roformer_ep_368_sdr_12.9628.ckpt

The [cpu] extra is also the one the README gives for Apple Silicon, where PyTorch models run on MPS. It describes the DirectML extra as experimental and community-supported. Without --model_filename you get the default, model_bs_roformer_ep_317_sdr_12.9755, which is viperx 1297 rather than 1296. --single_stem Vocals or --single_stem Instrumental writes only one side, and --output_format MP3 changes the default FLAC. Models download on first use to /tmp/audio-separator-models/ unless you set --model_file_dir.

Of the three, this gives you the most control, and it is the one to script. It is also the one where installing Python, PyTorch and ffmpeg is your job.

When the vocal still has music in it

Whichever tool you pick, a vocal stem can come out with the band still faintly audible. In audio-separator issue #7, a vocal file "still contains background music". The maintainer's tests on the reporter's file found the instrumental-trained model, UVR-MDX-NET-Inst_HQ_3, left a quiet backing "perhaps 30% of original volume" under the vocal. Vocal-focused models, UVR-MDX-NET-Voc_FT and Kim_Vocal_2, returned a clean vocal from the same song, and the maintainer's advice was that for clean vocals "you'll probably get better results using a vocal-focused model". So the first fix is a model trained for the stem you want, and a larger model if the mix is dense.

The second fix is a post-stage: a de-noise model run over the separated vocal, aimed at what the separation left behind.

From the studio: VocalDrop chains both fixes into one run: its Max model, viperx 1296, for the separation, then De-noise, a MelBand RoFormer pass that "clears hiss and background noise left behind after isolation", and optionally Enhance vocals, the Apollo restorer, for a vocal from a lossy source.

In UVR5 or audio-separator you can do the same by hand: aufr33's MelBand de-noise model is in both model lists, and you run it over the vocal file as a second pass.

Option 4: VocalDrop, step by step

VocalDrop packages the same idea as a native app, about 245 MB to download (v1.6.2: a 244.7 MB dmg, a 239.8 MB Windows installer). ffmpeg and the Python runtime are bundled. The builds are not code-signed yet, so the first launch needs the one-time step in the README. On NVIDIA and Apple Silicon, it runs BS-PolarFormer for Fast and BS-Roformer viperx 1296 for Max, on the GPU. On AMD and Intel graphics under Windows, Fast runs a smaller MDX model, Kim Vocal 2, through DirectML.

If you already run viperx 1296 in UVR5, Max will not sound better: it is the same checkpoint. What changes is the setup, and the steps around the separation.

1. Get the audio in

Drop a file or a folder, or paste a link into Fetch from a link. The app fetches from YouTube, SoundCloud, Vimeo, TikTok and other sites, using the bundled yt-dlp. Video works too: the audio is extracted, separated, and can be muxed back in sync. Whatever tool you use, you are responsible for having the rights to the audio you process.

2. Choose Fast or Max

Isolate vocals has two settings. Fast is "a single fast pass — great for clean, well-recorded sources". Max is "the heaviest model — best fidelity", and noticeably slower; on a machine without a supported GPU, the app says so in the hint itself. Use Fast for a clean studio recording, and Max for a dense mix.

The model for each mode downloads once on first use, about 100 MB for Fast and 600 MB for Max.

3. Keep the instrumental

By default VocalDrop writes the vocal. To remove vocals rather than keep them, switch on Keep stems ("Also save the instrumental and backing tracks"), and the instrumental is saved beside it as <name>_instrumental.

4. Export

Convert writes the result as MP3, WAV, FLAC, M4A, OGG, Opus or AIFF, and the README says metadata and cover art are preserved. The same chain can also trim silences, which is useful for spoken word and less so for music.

5. Or do it from a terminal

The app binary doubles as a command-line tool. On macOS, open the app once to accept its terms, then run /Applications/VocalDrop.app/Contents/MacOS/vocaldrop-app install-cli. On Windows, tick the command-line option in the installer. Then:

vocaldrop vocals --quality max --keep-stems song.mp3
vocaldrop run --vocals --denoise --convert wav *.mp3 --out stems/

--json prints machine-readable events for scripts.

What leaves the machine

In VocalDrop, separation, conversion and clean-up all run locally. The network is used in four cases, and the README lists them: the one-time engine and model download, fetching a link you paste, the update check, and anonymous diagnostics. Diagnostics are crash reports and coarse usage counts, "never file contents, never file paths". Both the update check and diagnostics can be switched off in Settings. An NVIDIA machine also fetches a one-time CUDA package of about 2.5 GB when acceleration is enabled.

Audacity's effect, UVR5 and audio-separator run locally too, apart from their own model downloads. Any of the four keeps the song off someone else's server.

The limits

  • Two stems, not four. VocalDrop splits the vocal from the instrumental (Keep stems saves the instrumental too), so backing vocals stay with the lead and drums stay with the bass. If you need drums or bass on their own, use a 4-stem Demucs model in UVR5 or the OpenVINO plugin.
  • No model hears your song in advance. Published scores come from test sets that do not contain your track. Run it through two models and listen.
  • No Intel Mac build. VocalDrop runs on Apple Silicon Macs and on Windows 10 and 11 (x64). On an Intel Mac, UVR5's x86_64 build is the option.

If you want the model without a Python setup, download VocalDrop and start with Fast on a clean recording. If you would rather choose the checkpoint yourself, the model post covers which file to load in UVR5 or audio-separator.

All writingNextSort IDM downloads by file type: categories, masks and the registry