Download API.md from hoyso48/STEMMA-MOSS-Music: direct link, hf CLI and curl.
- Browser
- Download file 7.62 kB
-
https://huggingface.co/hoyso48/STEMMA-MOSS-Music/resolve/main/API.md
- Command line
-
hf download hf://hoyso48/STEMMA-MOSS-Music/API.md
-
curl -L -o API.md https://huggingface.co/hoyso48/STEMMA-MOSS-Music/resolve/main/API.md
Native API contract
Entry points
load_model(local_snapshot, device="cuda:0", dtype=torch.bfloat16, attention="sdpa")returns aSessionwith public.modeland.processor. Missing/unexpected tensors fail. MOSS uses vendored pinned source and instance-only compatibility patches; no remote trust or external base snapshot is required after packaging.Request(prompt, audios)holds one user turn and 0--5 ordered local files orWaveform(samples: float[N], sample_rate: int). Waveforms must be mono. Files are mean-downmixed. URLs, fullsong directories, native control tokens and multi-turn are intentionally outside this initial contract.prepare(requests)->PreparedBatch;generate(requests, profile=..., max_new_tokens=...)->list[str], one completion per request.extract_audio_embeddings(batch, grad=False)->list[AudioEmbedding]in sample/source order.extract_contextual_embeddings(batch, layers=(-1,), grad=False)-> per-source dictionaries of requested layer tensors.forward(batch, output_hidden_states=False, grad=False)returns native ModelOutput, logits[B,1,V], optional language hidden states[B,L,D]..model(**batch.inputs, ...)permits advanced custom calls/full logits.
All core functions document input/output shapes in source docstrings. No method quietly averages tokens, detaches gradients requested by grad=True, sorts sources or inserts silence between different slots.
Missing markers follow the paper-reference wrapper: MF single audio uses
<audio>\n{prompt} without an added label; MOSS and multiple audios use numbered
labels. Explicit markers and arbitrary labels remain unchanged.
MF retains even a subsecond trailing window. A window below the conv/pool token resolution can produce zero tokens. An instance-only RoTE patch avoids upstream out-of-range indexing for that unused row; waveform/mel/masks are not truncated. A whole source producing zero tokens is rejected explicitly.
Shapes and source mappings
B=requests; A=total sources; W=total independent MF windows; L=left-padded expanded prompt width; F=mel frames; U=all valid audio tokens; U_i=one source's valid tokens.
| Item | MF2601 | MOSS-Music |
|---|---|---|
| Input IDs / attention mask | [B,L] |
[B,L] |
| Mel / feature mask | [W,128,3000] / [W,3000] |
[A,128,F_max] / lengths [A] |
| Source grouping | slot lengths + windows/source | group sizes [B], token lengths [A] |
| Raw encoder/source | [U_i,1280] before RoTE |
[U_i,1280] after final encoder norm |
| Projected/source | [U_i,3584] after RoTE/projector |
[U_i,4096] after audio adapter |
| DeepStack/source | none | 3 raw [U_i,1280], 3 projected [U_i,4096] |
| Language hidden states | [B,L,3584] |
[B,L,4096] |
| Slot token_positions | [U_i] padded-row indexes |
[U_i] padded-row indexes; exclude time digits |
AudioSlot.sample_index/audio_index and AudioEmbedding identify source boundaries.
No output embedding padding is returned. PreparedBatch.inputs.attention_mask
marks language padding; slot positions specify actual audio tokens, which may be
noncontiguous because timestamps/control tokens appear inside native spans.
MOSS encodes complete-source mel, splits at 400 frames without crossing a source
boundary and uses reset positions plus padding attention masks per chunk. The
last short chunk is preserved. conv_chunksize=64 batches independent chunks;
it is not an audio duration. Valid flat final and DeepStack streams [1,U,1280]
are regrouped by explicit source metadata before injection into language layers.
MF uses separate 30-second windows with a 1,200-second cap per source, and resets
RoTE per source. Sources with 2/3 windows have timeline indexes [0,1,0,1,2].
Reference SoundFile/SciPy 16-kHz preprocessing preserves subsecond tails.
File decoding uses SoundFile float32 and an arithmetic channel mean. Resampling
uses scipy.signal.resample_poly with rates reduced by their greatest common
divisor and SciPy's pinned default filter; it does not use soxr/librosa defaults.
Resampling precedes the MF duration cap. Do not replace this with legacy MF
evaluation helpers that use soxr or discard a final subsecond 30-second window.
Raw/projected extraction does not run the language model or produce a pooled retrieval vector. Contextual extraction performs a complete causal language prefill, then selects only the audio-token positions at the requested layers (layer 0 is the decoder input; negative indexes follow Python indexing). An audio token can see preceding text/audio, not a question placed after it. For explicitly question-conditioned audio states, place the question before the audio markers. No pooling or L2 normalization is applied; callers choose those research policies.
Generation, caching and hooks
paper_reference is greedy with repetition=1 and output budget 16,384 by default;
generic_sampled is MF temperature .7/top_p .9 or MOSS 1/top_p .8, both top_k 50.
GenerationConfig is constructed explicitly, not inherited from training's saved
sampled/cache-disabled defaults. KV cache is explicitly enabled for generation,
disabled for embedding prefill. Prefix caching is not implemented here.
Native contexts are 32,768 (MF) / 40,960 (MOSS). Batched output budget conservatively uses the largest padded prompt width; near the limit, call with a single request for the paper's per-request remaining-context budget. Generic sampled batches share one seed-3407 RNG stream; caller RNG is restored afterwards. This differs from vLLM's independently seeded requests. Greedy cross-kernel text identity and paper-score reproduction have not been certified.
Use session.model.named_modules() and register_forward_hook to inspect/edit
forward execution. Projector hooks avoid ambiguity about DeepStack injection.
MOSS's temporary language-layer hooks add DeepStack after each early layer and
use prepend=True, before Transformers' persistent output recorder and normally
registered user hooks. Thus ordinary hooks and returned hidden states observe
post-injection boundaries on both the first and subsequent forwards. To observe
pre-injection outputs, instrument the editable layer forward itself; validate
the intended observation point for your research use.
Remove handles in finally. The same MOSS instance must not run concurrent calls.
Ordinary autograd is supported by grad=True and direct model calls. This wrapper is not a fine-tuning implementation: checkpointed training is rejected because temporary DeepStack hooks would otherwise be removed before backward recomputation. Long-prompt full hidden-state collection can be expensive; select a small request for exploratory embedding work. Multi-turn and efficient streaming extraction of only selected hidden layers are follow-ups, not silently claimed capabilities.
Optional vLLM reference
vLLM is not a pip dependency of this native package. The paper used locally
patched vLLM 0.23.1.dev0 at source commit 0fc695fc6d1d82e9a5ac6835ac8e4e1c83703665,
not stock vLLM. MF requires multi-audio processor registration, independent RoTE
timestamps and HF weight-name mapping; MOSS requires the compatible custom audio
model, Music weights/processor identities and regenerated position buffer.
The historical MOSS receipts did not seal the custom model source file itself;
an installed-version string does not prove its historical byte identity.
This release does not advertise stock-vLLM compatibility or ship a working vLLM
server. Stage 3 compares native features with the available audited local route.