CLAP larger_clap_general for Core AI: asset, tokenizer files, logit scales
Browse files
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
crate-clap-general-float32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,56 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: laion/larger_clap_general
|
| 4 |
+
tags:
|
| 5 |
+
- audio
|
| 6 |
+
- clap
|
| 7 |
+
- audio-text-retrieval
|
| 8 |
+
- zero-shot-audio-classification
|
| 9 |
+
- core-ai
|
| 10 |
+
- apple-silicon
|
| 11 |
+
- macos
|
| 12 |
+
library_name: swift-sample-search
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# crate-clap-general — CLAP for Core AI (macOS 27)
|
| 16 |
+
|
| 17 |
+
LAION's [`larger_clap_general`](https://huggingface.co/laion/larger_clap_general) (Wu et al.,
|
| 18 |
+
"Large-scale Contrastive Language-Audio Pretraining", ICASSP 2023; Apache 2.0) exported to an
|
| 19 |
+
Apple Core AI `.aimodel` for on-device sample-library search, similarity and zero-shot tagging.
|
| 20 |
+
Why this checkpoint and not `larger_clap_music`: the music checkpoint's Hugging Face conversion
|
| 21 |
+
collapses every clip to nearly one vector (audio–audio cosines 0.81–0.99, audio–text 0.01–0.04,
|
| 22 |
+
logit scale 1.03, measured in `transformers` itself), so it cannot search or tag anything. The
|
| 23 |
+
general checkpoint was trained on music too and behaves (a kick file scores 0.49 against
|
| 24 |
+
"a kick drum").
|
| 25 |
+
|
| 26 |
+
Read by [swift-sample-search](https://github.com/arraypress/swift-sample-search) and the
|
| 27 |
+
[`crate`](https://github.com/arraypress/swift-crate-cli) command-line tool.
|
| 28 |
+
|
| 29 |
+
## Files
|
| 30 |
+
|
| 31 |
+
| Path | What |
|
| 32 |
+
|---|---|
|
| 33 |
+
| `crate-clap-general-float32.aimodel/` | Two entry points: `audio` (log-mel `[1,1,1001,64]` → `[1,512]`, HTSAT + projection) and `text` (RoBERTa ids `[1,77]` + mask `[1,77]` → `[1,512]`). Float32, 797 MB. |
|
| 34 |
+
| `clap-support/vocab.json`, `merges.txt` | The checkpoint's own RoBERTa byte-level BPE files. |
|
| 35 |
+
| `clap-support/scales.json` | `exp(logit_scale_a)` and `exp(logit_scale_t)` from the checkpoint (38.66 and 14.29), the token length, the source model id. |
|
| 36 |
+
|
| 37 |
+
The host computes the log-mel exactly as `ClapFeatureExtractor` does (48 kHz, 64 Slaney mels
|
| 38 |
+
50–14000 Hz, n_fft 1024, hop 480, `repeatpad` for clips under ten seconds), tokenises with the
|
| 39 |
+
files above, and L2-normalises the outputs — `get_audio_features` / `get_text_features`.
|
| 40 |
+
|
| 41 |
+
## Fidelity
|
| 42 |
+
|
| 43 |
+
Measured against `transformers` 5.17 on the same inputs: tokenizer ids identical over 20
|
| 44 |
+
phrases; log-mel 116–156 dB PSNR; text embeddings 142 dB; audio embeddings 144–147 dB
|
| 45 |
+
(cosine 0.99999+); zero-shot probabilities within 1e-8. The export script and the fixtures are
|
| 46 |
+
in the library repo (`Tools/export_clap.py`).
|
| 47 |
+
|
| 48 |
+
## Use
|
| 49 |
+
|
| 50 |
+
```sh
|
| 51 |
+
hf download arraypress/crate-clap-general --local-dir models
|
| 52 |
+
crate model install models
|
| 53 |
+
crate index ~/Samples && crate search "punchy 808 kick"
|
| 54 |
+
```
|
| 55 |
+
|
| 56 |
+
Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.
|
clap-support/merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
clap-support/scales.json
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"logit_scale_audio": 38.66471862792969,
|
| 3 |
+
"logit_scale_text": 14.285714149475098,
|
| 4 |
+
"max_tokens": 77,
|
| 5 |
+
"model": "laion/larger_clap_general"
|
| 6 |
+
}
|
clap-support/vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
crate-clap-general-float32.aimodel/main.hash
ADDED
|
Binary file (32 Bytes). View file
|
|
|
crate-clap-general-float32.aimodel/main.mlirb
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7ff9511b7cd94860359eb9008cc50cb929e19a7c0fafe3aff6d434bb4ff59e0c
|
| 3 |
+
size 797349985
|
crate-clap-general-float32.aimodel/metadata.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"author" : "LAION (Wu, Chen, Zhang, Hui, Berg-Kirkpatrick, Dubnov) — CLAP; Core AI export by SampleSearch",
|
| 3 |
+
"creationDate" : "20260924T112203Z",
|
| 4 |
+
"assetVersion" : "2.0",
|
| 5 |
+
"description" : "CLAP laion\/larger_clap_general: audio = log-mel [1,1,1001,64] → embedding [1,512] (HTSAT + projection); text = RoBERTa ids [1,77] + mask → embedding [1,512]. L2-normalise in the host. Tokenizer files beside the asset.",
|
| 6 |
+
"producer" : "coreai-core 1.0.0b2",
|
| 7 |
+
"license" : "Apache-2.0"
|
| 8 |
+
}
|