Instructions to use MC7ever/pe-single-L14 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MC7ever/pe-single-L14 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="MC7ever/pe-single-L14")# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("MC7ever/pe-single-L14") model = AutoModel.from_pretrained("MC7ever/pe-single-L14", device_map="auto") - PerceptionEncoder
How to use MC7ever/pe-single-L14 with PerceptionEncoder:
# Use any PE model as a vision encoder import core.vision_encoder.pe as pe model = pe.VisionTransformer.from_config("MC7ever/pe-single-L14", pretrained=True) - Notebooks
- Google Colab
- Kaggle
pe-single-L14 β one fused Meta Perception Encoder, no router
Single-file joint embedding model (audio + image/video + text, 1024-d shared space, 1.03B params) fusing four Meta PE family members:
| source | HF id | what was fused |
|---|---|---|
| PE-AV base (anchor) | facebook/pe-av-base |
full joint model kept as scaffold |
| PE A-Frame base | facebook/pe-a-frame-base |
audio tower, weight 0.25 |
| PE-Core-L14-336 | facebook/PE-Core-L14-336 |
vision trunk (near-identical to anchor video tower β init copy, ~no-op) |
| PE-Spatial-L14-448 | facebook/PE-Spatial-L14-448 |
vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below) |
Why L-scale: PE-Core/Spatial-G14 are width-1536 Γ 50 layers while PE-AV/A-Frame are width-1024 β cross-width tensors cannot be weight-averaged. L/14 (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video share width/depth/patch, so true weight fusion is possible.
Status (2026-10-03): CERTIFIED β frozen, no further weight search
Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:
| model | COCO t2i R@1 / R@5 | COCO i2t R@1 | ESC-50 acc |
|---|---|---|---|
facebook/pe-av-base (anchor) |
0.720 / 0.953 | 0.013 | 0.885 |
| this tag | 0.713 / 0.953 | 0.013 | 0.895 |
Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor quality while carrying A-Frame/Core/Spatial weights in one file β it does not beat the anchor on these benches by itself. Measured gains are deferred to alignment fine-tuning, for which this tag is the frozen init. Certification: 9-build grid + 18-genome evolution (fixed machinery) + full-slice validation, all agreeing (audio light, vision ~zero).
15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor: vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs 0.547; V2/V3 at chance for both β text embeds are dummy-video dominated in this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025 (fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas (J5 +0.033). No regression anywhere; consistent small audio edge. Full table: suite.json in the GitHub repo.
- Evolution rerun in progress after two voided attempts (documented below) β if it certifies a better genome, that becomes the next revision.
- Layer-depth probe: naive mid-network readout does NOT beat the joint output (t2i R@1 β€ 0.06 vs 0.79) β mid-network features need trained per-layer pooling (fine-tune phase), not a free readout change.
- Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
(2) silent no-op
load_state_dicton this model's dual-prefix key layout; (3) joint-path audio mismatch that faked ESC collapse at high audio weight. All fixed (pristine clones,copy_by param name, standalone-tower-only audio mapping) before trusting any evolution number.
How it was built
- Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
(both
pe_audio_encoder, 1024/16L/8H β 422/422 tensors matched). - Video trunk: explicit native-PE β timm key map (patch_embed, cls_token, pos_embed@336px, norms, QKV, MLP Γ24 blocks). Pool/head/proj kept from anchor.
- Blend weights chosen by measured 3Γ3 grid search over COCO-retrieval +
ESC-50 slices (script
search_pe_weights.py), not by vibes. Vision donor weight degraded COCO t2i monotonically (0.766 β 0.64), so the winning tag keeps anchor vision and blends audio at 0.25.
Usage
from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
return_tensors="pt")
out = m(**inputs) # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...
Retrieval: dot-product video_embeds @ text_video_embeds.T.
Zero-shot audio: argmax over audio_embeds @ text_audio_embeds.T with
"the sound of {label}" prompts.
Honest limits
- Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative anchor-vs-variant only β not official MMEB/PE-paper figures.
- i2t (textβimage caption ranking over 150 similar COCO captions) sits near floor for anchor and variants alike; t2i and audio carry the signal.
- Encoder only: no generation, detection, or segmentation heads.
- Next: evolution-based per-tower weights, intermediate-layer readout (PE paper: best features are mid-network), GPU alignment fine-tune.
Reproduce
fuse_pe_single.py (fusion) Β· eval_pe_bench.py (COCO+ESC-50 bench) Β·
search_pe_weights.py (grid) Β· evo_fuse.py (evolution) Β· layer_sweep.py
(depth probe). Benchmark slice metadata: MMEB MSCOCO_i2t rows + COCO val2014
images + ashraq/esc50 streaming clips.
- Downloads last month
- 14
Evaluation results
- t2i R@1 (search slice n=64)self-reported0.766
- t2i R@5 (approx, search slice)self-reported0.950
- ESC-50 accuracy (search slice)self-reported0.960