File size: 1,333 Bytes
7ff0331
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
This Stability AI Model is licensed under the Stability AI Community License,
Copyright (c) Stability AI Ltd. All Rights Reserved.

This is a Derivative Work. The Stability AI Materials were modified as follows:

  The model was converted from its original PyTorch format to Apple's Core AI
  `.aimodel` format, using `torch.export` and `coreai-torch`. The network was split
  into four separately-callable graphs (text conditioner, diffusion transformer,
  latent decoder, audio encoder) so that the sampling loop runs on the host.

  Two training-time behaviours were disabled, both inactive at inference:
    - `mask_noise` in the autoencoder (a training augmentation)
    - the softnorm bottleneck's decode-time dither

  The bottleneck dither was measured before removal: output with it disabled differs
  from stock by SNR 18.8 dB / cosine 0.9934, while two stock runs differ from each other
  by SNR 21.1 dB / cosine 0.9961. Removing it therefore lands within the model's own
  run-to-run variance, and makes generation deterministic.

  No weights were retrained, fine-tuned, quantized or otherwise altered in value.
  Graph weights are stored at float32, matching the source checkpoint.

Original model: stabilityai/stable-audio-3-medium
The T5Gemma text encoder is subject to the Gemma Terms of Use (see LICENSE_GEMMA.md).