File size: 7,619 Bytes
5750492
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
# Native API contract

## Entry points

- `load_model(local_snapshot, device="cuda:0", dtype=torch.bfloat16, attention="sdpa")`
  returns a `Session` with public `.model` and `.processor`. Missing/unexpected
  tensors fail. MOSS uses vendored pinned source and instance-only compatibility
  patches; no remote trust or external base snapshot is required after packaging.
- `Request(prompt, audios)` holds one user turn and 0--5 ordered local files or
  `Waveform(samples: float[N], sample_rate: int)`. Waveforms must be mono. Files
  are mean-downmixed. URLs, fullsong directories, native control tokens and
  multi-turn are intentionally outside this initial contract.
- `prepare(requests)` -> `PreparedBatch`; `generate(requests, profile=...,
  max_new_tokens=...)` -> `list[str]`, one completion per request.
- `extract_audio_embeddings(batch, grad=False)` -> `list[AudioEmbedding]` in
  sample/source order. `extract_contextual_embeddings(batch, layers=(-1,),
  grad=False)` -> per-source dictionaries of requested layer tensors.
- `forward(batch, output_hidden_states=False, grad=False)` returns native
  ModelOutput, logits `[B,1,V]`, optional language hidden states `[B,L,D]`.
  `.model(**batch.inputs, ...)` permits advanced custom calls/full logits.

All core functions document input/output shapes in source docstrings. No method
quietly averages tokens, detaches gradients requested by grad=True, sorts sources
or inserts silence between different slots.

Missing markers follow the paper-reference wrapper: MF single audio uses
`<audio>\n{prompt}` without an added label; MOSS and multiple audios use numbered
labels. Explicit markers and arbitrary labels remain unchanged.

MF retains even a subsecond trailing window. A window below the conv/pool token
resolution can produce zero tokens. An instance-only RoTE patch avoids upstream
out-of-range indexing for that unused row; waveform/mel/masks are not truncated.
A whole source producing zero tokens is rejected explicitly.

## Shapes and source mappings

B=requests; A=total sources; W=total independent MF windows; L=left-padded expanded
prompt width; F=mel frames; U=all valid audio tokens; U_i=one source's valid tokens.

| Item | MF2601 | MOSS-Music |
| --- | --- | --- |
| Input IDs / attention mask | `[B,L]` | `[B,L]` |
| Mel / feature mask | `[W,128,3000]` / `[W,3000]` | `[A,128,F_max]` / lengths `[A]` |
| Source grouping | slot lengths + windows/source | group sizes `[B]`, token lengths `[A]` |
| Raw encoder/source | `[U_i,1280]` before RoTE | `[U_i,1280]` after final encoder norm |
| Projected/source | `[U_i,3584]` after RoTE/projector | `[U_i,4096]` after audio adapter |
| DeepStack/source | none | 3 raw `[U_i,1280]`, 3 projected `[U_i,4096]` |
| Language hidden states | `[B,L,3584]` | `[B,L,4096]` |
| Slot token_positions | `[U_i]` padded-row indexes | `[U_i]` padded-row indexes; exclude time digits |

`AudioSlot.sample_index/audio_index` and `AudioEmbedding` identify source boundaries.
No output embedding padding is returned. `PreparedBatch.inputs.attention_mask`
marks language padding; slot positions specify actual audio tokens, which may be
noncontiguous because timestamps/control tokens appear inside native spans.

MOSS encodes complete-source mel, splits at 400 frames without crossing a source
boundary and uses reset positions plus padding attention masks per chunk. The
last short chunk is preserved. `conv_chunksize=64` batches independent chunks;
it is not an audio duration. Valid flat final and DeepStack streams `[1,U,1280]`
are regrouped by explicit source metadata before injection into language layers.

MF uses separate 30-second windows with a 1,200-second cap per source, and resets
RoTE per source. Sources with 2/3 windows have timeline indexes `[0,1,0,1,2]`.
Reference SoundFile/SciPy 16-kHz preprocessing preserves subsecond tails.

File decoding uses SoundFile float32 and an arithmetic channel mean. Resampling
uses `scipy.signal.resample_poly` with rates reduced by their greatest common
divisor and SciPy's pinned default filter; it does not use soxr/librosa defaults.
Resampling precedes the MF duration cap. Do not replace this with legacy MF
evaluation helpers that use soxr or discard a final subsecond 30-second window.

Raw/projected extraction does not run the language model or produce a pooled
retrieval vector. Contextual extraction performs a complete causal language
prefill, then selects only the audio-token positions at the requested layers
(layer 0 is the decoder input; negative indexes follow Python indexing). An audio
token can see preceding text/audio, not a question placed after it. For explicitly
question-conditioned audio states, place the question before the audio markers.
No pooling or L2 normalization is applied; callers choose those research policies.

## Generation, caching and hooks

`paper_reference` is greedy with repetition=1 and output budget 16,384 by default;
`generic_sampled` is MF temperature .7/top_p .9 or MOSS 1/top_p .8, both top_k 50.
GenerationConfig is constructed explicitly, not inherited from training's saved
sampled/cache-disabled defaults. KV cache is explicitly enabled for generation,
disabled for embedding prefill. Prefix caching is not implemented here.

Native contexts are 32,768 (MF) / 40,960 (MOSS). Batched output budget conservatively
uses the largest padded prompt width; near the limit, call with a single request
for the paper's per-request remaining-context budget. Generic sampled batches
share one seed-3407 RNG stream; caller RNG is restored afterwards. This differs
from vLLM's independently seeded requests. Greedy cross-kernel text identity and
paper-score reproduction have not been certified.

Use `session.model.named_modules()` and `register_forward_hook` to inspect/edit
forward execution. Projector hooks avoid ambiguity about DeepStack injection.
MOSS's temporary language-layer hooks add DeepStack after each early layer and
use prepend=True, before Transformers' persistent output recorder and normally
registered user hooks. Thus ordinary hooks and returned hidden states observe
post-injection boundaries on both the first and subsequent forwards. To observe
pre-injection outputs, instrument the editable layer forward itself; validate
the intended observation point for your research use.
Remove handles in `finally`. The same MOSS instance must not run concurrent calls.

Ordinary autograd is supported by grad=True and direct model calls. This wrapper
is not a fine-tuning implementation: checkpointed training is rejected because
temporary DeepStack hooks would otherwise be removed before backward recomputation.
Long-prompt full hidden-state collection can be expensive; select a small request
for exploratory embedding work. Multi-turn and efficient streaming extraction of
only selected hidden layers are follow-ups, not silently claimed capabilities.

## Optional vLLM reference

vLLM is **not** a pip dependency of this native package. The paper used locally
patched vLLM 0.23.1.dev0 at source commit `0fc695fc6d1d82e9a5ac6835ac8e4e1c83703665`,
not stock vLLM. MF requires multi-audio processor registration, independent RoTE
timestamps and HF weight-name mapping; MOSS requires the compatible custom audio
model, Music weights/processor identities and regenerated position buffer.
The historical MOSS receipts did not seal the custom model source file itself;
an installed-version string does not prove its historical byte identity.
This release does not advertise stock-vLLM compatibility or ship a working vLLM
server. Stage 3 compares native features with the available audited local route.