File size: 2,507 Bytes
5673983 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 | ---
license: apache-2.0
base_model: laion/larger_clap_general
tags:
- audio
- clap
- audio-text-retrieval
- zero-shot-audio-classification
- core-ai
- apple-silicon
- macos
library_name: swift-sample-search
---
# crate-clap-general — CLAP for Core AI (macOS 27)
LAION's [`larger_clap_general`](https://huggingface.co/laion/larger_clap_general) (Wu et al.,
"Large-scale Contrastive Language-Audio Pretraining", ICASSP 2023; Apache 2.0) exported to an
Apple Core AI `.aimodel` for on-device sample-library search, similarity and zero-shot tagging.
Why this checkpoint and not `larger_clap_music`: the music checkpoint's Hugging Face conversion
collapses every clip to nearly one vector (audio–audio cosines 0.81–0.99, audio–text 0.01–0.04,
logit scale 1.03, measured in `transformers` itself), so it cannot search or tag anything. The
general checkpoint was trained on music too and behaves (a kick file scores 0.49 against
"a kick drum").
Read by [swift-sample-search](https://github.com/arraypress/swift-sample-search) and the
[`crate`](https://github.com/arraypress/swift-crate-cli) command-line tool.
## Files
| Path | What |
|---|---|
| `crate-clap-general-float32.aimodel/` | Two entry points: `audio` (log-mel `[1,1,1001,64]` → `[1,512]`, HTSAT + projection) and `text` (RoBERTa ids `[1,77]` + mask `[1,77]` → `[1,512]`). Float32, 797 MB. |
| `clap-support/vocab.json`, `merges.txt` | The checkpoint's own RoBERTa byte-level BPE files. |
| `clap-support/scales.json` | `exp(logit_scale_a)` and `exp(logit_scale_t)` from the checkpoint (38.66 and 14.29), the token length, the source model id. |
The host computes the log-mel exactly as `ClapFeatureExtractor` does (48 kHz, 64 Slaney mels
50–14000 Hz, n_fft 1024, hop 480, `repeatpad` for clips under ten seconds), tokenises with the
files above, and L2-normalises the outputs — `get_audio_features` / `get_text_features`.
## Fidelity
Measured against `transformers` 5.17 on the same inputs: tokenizer ids identical over 20
phrases; log-mel 116–156 dB PSNR; text embeddings 142 dB; audio embeddings 144–147 dB
(cosine 0.99999+); zero-shot probabilities within 1e-8. The export script and the fixtures are
in the library repo (`Tools/export_clap.py`).
## Use
```sh
hf download arraypress/crate-clap-general --local-dir models
crate model install models
crate index ~/Samples && crate search "punchy 808 kick"
```
Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.
|