VocalDrop lifts the vocal clean out of any track

A free AI vocal remover that lifts a studio-clean acapella out of any song, on your own CPU or GPU — your audio never leaves the machine.

Unsigned build: on macOS right-click then Open; on Windows More info, Run anyway.

All platforms

v1.5.0 · 23 August 2026

2

stems out of one pass — the isolated vocal and the instrumental together

All

GPU vendors run it: Metal, CUDA and DirectML, with a CPU engine where there is none

0

uploads — every model runs on your own machine, so the track never leaves it

Free

no tiers, no watermark and no account

Built like a studio tool given away free.

Fast and Max models

The RoFormer model family UVR5 uses, offered at two qualities. Fast for a quick pass, Max when the mix is dense.

Both stems, one pass

The isolated vocal and the instrumental land together in the Results feed — separation is one job, not two runs.

Video too

Drop a music video or a movie clip; the audio is extracted, separated, and the stem you chose is re-muxed back in sync.

Or just paste a link

YouTube, SoundCloud, Vimeo, TikTok and a thousand-odd more, fetched with the real title and artwork — no separate downloader in the way.

Two passes, depending on the mix.

VocalDrop is built on the RoFormer separation models, run locally. One model runs per mode — nothing is blended behind your back, so what you hear is what that model produced.

  1. Bring the audio in

    A file, a folder, or a pasted link. Video is de-muxed, separated, and can be re-muxed back in sync with whichever stem you kept.

  2. Fast — BS-PolarFormer

    A compact RoFormer that runs on the GPU — Metal on Apple Silicon, CUDA on NVIDIA — and turns a track around in seconds. On AMD/Intel graphics, Kim Vocal 2 runs on the GPU via DirectML instead. Enough for a clean source.

  3. Max — BS-Roformer 1296

    A single, much larger pass for the highest fidelity. Slower, and the one to reach for on a dense mix.

  4. Both stems land together

    The vocal and the instrumental come out of the same run, so a karaoke bed costs nothing extra.

How good, ranked by measured quality.

TierModelSizeVocal qualitySpeed on GPU
MaxBS-Roformer 1296 (viperx)609 MB★★★★★ ~12.96 SDR — the community benchmarkseconds on CUDA/Metal, ~10 min on CPU
FastBS-PolarFormer97 MB★★★★☆ ~11.5 SDR — 80% of Max at 10× speedseconds
Fast (AMD/Intel)Kim Vocal 2 (MDX)66 MB★★★☆☆ ~9.8 SDR — best ONNX modelseconds on DirectML
Post-stageApollo Vocal Restorer194 MBrestoration — repairs lossy/dull vocalsruns after isolation
Post-stageMelBand De-noise870 MBdenoising — strips hiss and bleedruns after isolation

Against the usual suspects — the ones that upload your audio.

Compared onVocalDropUVR5LALAL.AIMVSEPMoisesiZotope RX
Audio never uploadedyesyesno — uploadsno — uploadsno — cloudyes
Free, unlimitedyesyesno — per-minuteno — queue-limitedno — subscriptionno — $199+
SOTA separation (RoFormer)yesyespartialyespartialno
GPU on every vendoryes — Metal · CUDA · DirectMLno — CUDA onlyn/an/an/ayes
Automatic GPU setupyes — one-time, in-appno — manualn/an/an/an/a
Full chain: isolate → denoise → restore → silence → convertyesno — manual chainingno — separation onlyno — separation onlypartialno — manual
AI vocal restoration (Apollo)yesnonononono
Link ingestion (YouTube, 1000+ sites)yesnonononono
Video in → synced video outyesnononopartialno
One-click installyes — 315 MBno — 1.6 GB + Pythonn/an/ayesyes
No account requiredyesyesnonono — requiredyes
New models without app updateyes — remote catalognoyes — server-sideyes — server-sideyes — server-sideno
vocaldrop vocals song.mp3 --quality max

Every option the app exposes, with --json for pipelines.

Your audio never leaves the machine — and the diagnostics that do.

Most free vocal extractors are websites that upload your audio to someone else's server. VocalDrop runs the models locally on your own CPU or GPU, so the audio is never sent anywhere, and there is no track limit, watermark or subscription attached to the decision.

What does leave the machine, stated plainly: anonymous crash and usage diagnostics — never file contents, never paths, never personal data — the one-time engine and model download on first launch, a URL if you paste one in for the app to fetch, and the update check if you leave it on. That is the complete list.

Press

Writing about it? The kit is already open.

Screenshots and the icon at every size, boilerplate in two lengths, the fact sheet and the palette — free to publish without asking, and downloadable as one zip.

Open the VocalDrop press kit

No embargo · review copies on request

Questions

What model does it use?

The RoFormer family, the same architecture UVR5 ships, run through audio-separator — as a native desktop app, with no Python setup and no dependency roulette. Max is BS-Roformer 1296 on every machine (GPU on NVIDIA and Apple Silicon, CPU on AMD/Intel where DirectML cannot run it). Fast is BS-PolarFormer, or Kim Vocal 2 on AMD/Intel graphics where it runs on the GPU via DirectML. Two optional post-stages — MelBand de-noise and the Apollo vocal restorer — polish the result.

Is it really free?

Yes. Free to use, with no watermarks, no track limits, no feature paywall and no account. It is closed source; the public repository holds the downloads and the README.

Does it need the internet?

Only to download your model on first use — about 100 MB for Fast, 600 MB for Max. The Python engine itself ships inside the installer. After that separation is entirely offline (an NVIDIA machine also fetches its one-time CUDA package when you enable acceleration).

What about my files?

They are processed on your machine and never uploaded. VocalDrop does send anonymous crash and usage diagnostics — never file contents, file paths or personal data — and the only other network calls are the one-time model download, a URL you paste in yourself, and the optional update check.

Which Macs and PCs?

Apple Silicon Macs (M1 to M4), Windows 10 or 11 on x64, and Linux on x64. Intel-Mac builds are not provided. Separation is GPU-accelerated on all three: Metal on Apple Silicon, and on Windows and Linux the app detects your GPU at launch — an NVIDIA card gets an automatic one-time CUDA setup, AMD/Intel graphics run on DirectML, and everything else uses the optimized CPU engine.

Why does my Mac say it cannot open it?

The current builds are not code-signed or notarized yet, so Gatekeeper objects on first launch: right-click the app and choose Open. On Windows, SmartScreen wants More info, then Run anyway. Signed builds are on the roadmap.

Can I automate it?

Yes — a full CLI ships inside the app and mirrors everything the interface does, with --json output for pipelines. macOS also gets right-click Quick Actions in Finder.

The VocalDrop desktop app

The vocals come out, and nothing goes up.

Separation runs on your own CPU or GPU, so the track never leaves the machine and there is no queue to wait in. One download for macOS, Windows and Linux.

Free · macOS, Windows and Linux

Built by a one-person studio. If it has earned a place in your day, you can help fund the next release →

VocalDrop

studio-grade acapella, on your own hardware