TuneyCompany/tuney-small
Viewer • Updated • 1 • 17
Weights for PHRASER: Loop-Precise Stem-Aware Music Structure Segmentation, trained on the Tuney-Small train split (426 tracks). This is the model behind the Tuney-Small rows of Tables 1 and 2 in the paper.
phraser_tuney_small.ckpt (PyTorch dict with state_dict, config, epoch)PhraserModel(dim_embed=128, num_heads=16), on frozen
MuQ-large features of four SCNet stems. The dim_embed stored in config is a stale
default; the network is built with 128.SHA256SUMS.from phraser.modules.phraser import PhraserModel
import torch
model = PhraserModel(dim_embed=128, num_heads=16)
sd = torch.load("phraser_tuney_small.ckpt", map_location="cpu", weights_only=False)["state_dict"]
model.load_state_dict({k.replace("module.", ""): v for k, v in sd.items()}, strict=False)
The end-to-end evaluation (separation, MuQ, decoding with onset snapping, scoring) is
python -m phraser.tuney_small.eval_ts phraser phraser_tuney_small.ckpt. See the repository README.
The model was trained on 426 instrumental, grid-built tracks. It has no vocal supervision, and the vocal output channel is untrained and must not be used. Accuracy on live or tempo-drifting music is lower.
CC BY-NC 4.0. Non-commercial use only. The model depends on MuQ weights, which are also licensed for non-commercial use only.