All writing

BS-RoFormer, viperx 1296 and BS-PolarFormer: what the SDR numbers mean

BS-RoFormer viperx 1296 scores a 12.10 dB median vocal SDR on 40 tracks; the 12.96 in its filename is a label. BS-PolarFormer's 11.00 is from another test.

BS-Roformer viperx 1296 is a vocal checkpoint of the Band-Split RoPE Transformer, and UVR5 lists it beside a sibling called 1297. Scored independently by python-audio-separator, 1296 reaches a median vocal SDR of 12.10 dB over 40 tracks. The "12.9628" in its filename is a label the trainer attached, not a result you can compare with anything. BS-PolarFormer is a smaller relative with a different positional encoding, scored at 11.00 dB on a different test set, so it cannot be ranked against the 12.10.

The rest of this post explains where each model comes from, what SDR measures, and why one checkpoint can carry three different scores. It ends with how to run them, including in VocalDrop, the studio's free vocal remover, which packages several of them.

What BS-RoFormer is

BS-RoFormer comes from the paper Music Source Separation with Band-Split RoPE Transformer by Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong and Yun-Ning Hung (September 2023). The name describes the two ideas in it:

  • Band-split. The input spectrogram is cut into 62 non-overlapping frequency bands, narrow at the bottom and wider towards the top, and each band is projected into its own representation before the transformer sees it. The model can then treat a bass register and a sibilant register differently.
  • RoPE. Positions inside the transformer are encoded with rotary position embedding.

The paper reports that the system ranked first in the music source separation track of the Sound Demixing Challenge 2023 (SDX23). A smaller six-layer version, trained only on MUSDB18-HQ, scored 9.80 dB averaged across four stems on that dataset's test songs, and 10.66 dB on vocals.

viperx 1296 and 1297: one configuration, two epochs

The checkpoints UVR5 calls "BS-Roformer-Viperx-1296" and "BS-Roformer-Viperx-1297" are named by their files, as UVR's model list (download_checks.json, the file the app reads, per gui_data/constants.py) shows:

UVR5 nameFile
BS-Roformer-Viperx-1297model_bs_roformer_ep_317_sdr_12.9755.ckpt
BS-Roformer-Viperx-1296model_bs_roformer_ep_368_sdr_12.9628.ckpt

So "1296" and "1297" are the first four digits of the number in each filename. The two YAML configs UVR downloads for them (1296, 1297) are identical: width 512, 12 layers, 62 frequency bands, target instrument vocals. Only the epoch in the filename differs, 317 against 368. ZFTurbo's pretrained model list calls this the "viperx edition" and links the trainer's GitHub account. The 1296 checkpoint is 639,317,465 bytes, about 609 MB.

Neither filename says what data produced its number. Treat the number as part of the name.

BS-PolarFormer

BS-PolarFormer is not in UVR5's download list. It was published by ZFTurbo in Music-Source-Separation-Training release v1.0.20 on 22 March 2026, described as a "Version of BS Roformer based on PoPE (Polar Coordinate Positional Embeddings)". The release holds one checkpoint, model_bs_polarformer_float16.ckpt, 102,410,301 bytes (about 97 MB), and its config.

PoPE comes from Gopalakrishnan, Csordás, Schmidhuber and Mozer. Their argument is that RoPE entangles what a token is with where it sits, and PoPE separates the two.

The config shows what changed against viperx 1296. It keeps the 12 layers and the 62 bands, halves the width to 256, sets use_pope: True, and stores the weights in float16. That is why it is about a sixth of the size.

What SDR measures

SDR is signal-to-distortion ratio, in decibels. You need the true isolated vocal for the song, so it can only be measured on a test set where the stems are known. The Music Demixing Challenge 2021 paper defines the simple, "global" version:

SDR = 10 · log10( Σ‖s(n)‖² + ε  /  Σ‖s(n) − ŝ(n)‖² + ε )

Here s is the true stem, ŝ is the model's estimate, and ε is 10⁻⁷. In words: the energy of the real vocal divided by the energy of everything the estimate got wrong, whether that is drum bleed, smeared consonants or missing air. Higher is better, and because the scale is logarithmic, +3 dB means the error energy roughly halved.

The other common version is BSS Eval v4, shipped as the museval package. It splits each song into one-second chunks, scores each chunk, and takes the median per song; papers then report the median across songs. The BS-RoFormer paper's figures are computed this way. BSS Eval v4 also reports SIR, SAR and ISR, and it discards chunks where a source or an estimate is silent. The challenge paper found the two measures correlate strongly, but they do not produce the same number.

Why one checkpoint has three scores

Here is every published score I could find for the models above, with where each one came from:

ModelVocal SDRSourceTest set and method
viperx 129612.9628filenamenot stated
viperx 129612.10audio-separatormedian of 40 tracks, BSS Eval v4
viperx 129712.9755filenamenot stated
viperx 129711.77audio-separatormedian of 40 tracks, BSS Eval v4
viperx 129710.87ZFTurboMVSEP Multisong
BS-PolarFormer float1611.00ZFTurboMVSEP Multisong
Kim Vocal 210.18audio-separatormedian of 40 tracks, BSS Eval v4

The two test sets are different in kind.

python-audio-separator scores models with its own test script: museval in v4 mode, run on MUSDB18-HQ, with the median taken across tracks. The 40 tracks in the file run alphabetically from "A Classic Education - NightOwl" to "James May - All Souls Moon". Five of them, including "Actions - One Minute Smile", appear in the musdb package's validation list, which its README carves out of MUSDB18's 100-song training folder. That means these are not the 50 test songs papers report on, and any model trained on MUSDB18 may have heard them.

MVSEP Multisong is 100 one-minute excerpts across many genres, with the reference stems kept on MVSEP's server. Its quality checker computes an average SDR that it describes as similar to the Music Demixing Challenge's, which is the global formula above. MVSEP itself warns that "there is a small chance that some of the tunes were used to train some models".

Different songs, different formula, different averaging. So:

  • 12.10 against 11.00 tells you nothing about which is better. It is 1296 on one set and PolarFormer on the other.
  • Within one table, the comparison holds. On audio-separator's 40 tracks, 1296 beats 1297 by 0.33 dB on vocals (12.10 against 11.77), although the filenames rank them the other way round. 1297 is slightly ahead on the instrumental, 16.45 against 16.31. On ZFTurbo's Multisong table, the 97 MB PolarFormer (11.00) and viperx 1297 (10.87) are within 0.13 dB. That table has no row for 1296.
  • The same name can be several checkpoints. MVSEP's BS PolarFormer page lists a 62-band model at 11.75 and a 124-band model at 12.02, both on Multisong. The public float16 file has 62 bands and scores 11.00 on the same set, so it is not the checkpoint behind those numbers. Quote 11.00 for the file you can download.

The filename habit causes the same confusion elsewhere. The MelBand de-noise checkpoint by aufr33 is called denoise_mel_band_roformer_aufr33_sdr_27.9959.ckpt. That 27.9959 measures a different task altogether and says nothing about vocals.

Where Kim Vocal 2, MDX and Demucs sit

Kim Vocal 2 is an MDX-Net model, listed in UVR5 as "MDX-Net Model: Kim Vocal 2". It ships as a 66 MB ONNX file, so it runs wherever ONNX Runtime has a backend, DirectML on Windows included. On audio-separator's 40 tracks it scores 10.18, about 1.9 dB under viperx 1296. Neither number comes from Multisong.

Demucs is a different tool for a different job. Hybrid Transformer Demucs separates four stems (drums, bass, vocals, other), and its README reports 9.00 dB averaged across all four on the MUSDB-HQ test set. It was trained with 800 extra songs. That looks close to BS-RoFormer's 9.80, which is also a four-stem average on the MUSDB18-HQ test set, but BS-RoFormer reached it with no extra data. Neither number is a vocal-only score.

For vocals alone, ZFTurbo's Multisong table gives the 4-stem HTDemucs4 8.24, against 11.00 for PolarFormer. audio-separator lists htdemucs_ft at a 10.79 median, but only 13 tracks were scored, so it is not like-for-like with the 40-track rows. The Demucs repository also says it is no longer maintained.

Running these models yourself

In UVR5

viperx 1296 and 1297 are in UVR5's model downloader under the names above, with Kim Vocal 2 and the other MDX-Net models. UVR5 does not list BS-PolarFormer.

With python-audio-separator

python-audio-separator loads the same files by filename and downloads them on first use. Use [gpu] instead of [cpu] for NVIDIA; the [cpu] build is also the one its README gives for Apple Silicon:

pip install "audio-separator[cpu]"
audio-separator song.wav --model_filename model_bs_roformer_ep_368_sdr_12.9628.ckpt
audio-separator song.wav --model_filename Kim_Vocal_2.onnx

Each run writes a vocal file and an instrumental file. audio-separator -l --list_filter=vocals lists the vocal models it knows. BS-PolarFormer is not among them today.

BS-PolarFormer with Music-Source-Separation-Training

PolarFormer runs through ZFTurbo's own inference.py. Download the two release files first. The repository's bs_roformer model builds it, but PoPE needs the separate PoPE-pytorch package, which requirements.txt does not install:

pip install -r requirements.txt PoPE-pytorch
python inference.py \
    --model_type bs_roformer \
    --config_path model_bs_polarformer_float16.yaml \
    --start_check_point model_bs_polarformer_float16.ckpt \
    --input_folder input/ \
    --store_dir separation_results/

In VocalDrop

From the studio: VocalDrop is the studio's free vocal remover for macOS (Apple Silicon) and Windows. It packages models from this post, so there is no Python to set up.

Which model runs depends on the machine:

  • On NVIDIA and Apple Silicon, Fast is BS-PolarFormer and Max is BS-Roformer viperx 1296, both on the GPU.
  • On AMD and Intel graphics under Windows, Fast runs Kim Vocal 2 on the GPU through DirectML.

The model downloads once, about 100 MB for Fast and 600 MB for Max, and after that separation runs offline. The audio is never uploaded. Every run returns the vocal and the instrumental, and two optional post-stages follow: the MelBand de-noise model and the Apollo vocal restorer. The same work runs from a terminal:

vocaldrop vocals --quality max song.wav

So which is the best UVR5 vocal model?

Of the UVR5 models in this post, viperx 1296 has the best independent vocal score: a 12.10 median on 40 tracks, 0.33 dB ahead of its 1297 sibling and about 1.9 dB ahead of Kim Vocal 2 on the same tracks. UVR5 lists other vocal models this post does not cover, so that is a ranking of three, not of the whole list. PolarFormer is not in UVR5 at all. At 97 MB it sits within 0.13 dB of viperx 1297 on Multisong, but no published test scores it against 1296, so the gap between the two is unknown.

When you compare models, check three things before trusting a number:

  • the filename is not the score;
  • both models were scored on the same test set;
  • both were scored with the same formula.

For a living ranking on one set, the MVSEP Multisong leaderboard recalculates every three hours. The last check is your own: no test set contains your song, so run it through two models and listen.

If you would rather do that without setting up Python, VocalDrop packages models from this post for macOS and Windows, free, and the audio stays on your machine.

All writingNextIs Kotlin Multiplatform production-ready? Every status, dated