File size: 2,507 Bytes
5673983
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
---
license: apache-2.0
base_model: laion/larger_clap_general
tags:
  - audio
  - clap
  - audio-text-retrieval
  - zero-shot-audio-classification
  - core-ai
  - apple-silicon
  - macos
library_name: swift-sample-search
---

# crate-clap-general — CLAP for Core AI (macOS 27)

LAION's [`larger_clap_general`](https://huggingface.co/laion/larger_clap_general) (Wu et al.,
"Large-scale Contrastive Language-Audio Pretraining", ICASSP 2023; Apache 2.0) exported to an
Apple Core AI `.aimodel` for on-device sample-library search, similarity and zero-shot tagging.
Why this checkpoint and not `larger_clap_music`: the music checkpoint's Hugging Face conversion
collapses every clip to nearly one vector (audio–audio cosines 0.81–0.99, audio–text 0.01–0.04,
logit scale 1.03, measured in `transformers` itself), so it cannot search or tag anything. The
general checkpoint was trained on music too and behaves (a kick file scores 0.49 against
"a kick drum").

Read by [swift-sample-search](https://github.com/arraypress/swift-sample-search) and the
[`crate`](https://github.com/arraypress/swift-crate-cli) command-line tool.

## Files

| Path | What |
|---|---|
| `crate-clap-general-float32.aimodel/` | Two entry points: `audio` (log-mel `[1,1,1001,64]` → `[1,512]`, HTSAT + projection) and `text` (RoBERTa ids `[1,77]` + mask `[1,77]` → `[1,512]`). Float32, 797 MB. |
| `clap-support/vocab.json`, `merges.txt` | The checkpoint's own RoBERTa byte-level BPE files. |
| `clap-support/scales.json` | `exp(logit_scale_a)` and `exp(logit_scale_t)` from the checkpoint (38.66 and 14.29), the token length, the source model id. |

The host computes the log-mel exactly as `ClapFeatureExtractor` does (48 kHz, 64 Slaney mels
50–14000 Hz, n_fft 1024, hop 480, `repeatpad` for clips under ten seconds), tokenises with the
files above, and L2-normalises the outputs — `get_audio_features` / `get_text_features`.

## Fidelity

Measured against `transformers` 5.17 on the same inputs: tokenizer ids identical over 20
phrases; log-mel 116–156 dB PSNR; text embeddings 142 dB; audio embeddings 144–147 dB
(cosine 0.99999+); zero-shot probabilities within 1e-8. The export script and the fixtures are
in the library repo (`Tools/export_clap.py`).

## Use

```sh
hf download arraypress/crate-clap-general --local-dir models
crate model install models
crate index ~/Samples && crate search "punchy 808 kick"
```

Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.