pe-single-L14 β€” one fused Meta Perception Encoder, no router

Single-file joint embedding model (audio + image/video + text, 1024-d shared space, 1.03B params) fusing four Meta PE family members:

source HF id what was fused
PE-AV base (anchor) facebook/pe-av-base full joint model kept as scaffold
PE A-Frame base facebook/pe-a-frame-base audio tower, weight 0.25
PE-Core-L14-336 facebook/PE-Core-L14-336 vision trunk (near-identical to anchor video tower β€” init copy, ~no-op)
PE-Spatial-L14-448 facebook/PE-Spatial-L14-448 vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below)

Why L-scale: PE-Core/Spatial-G14 are width-1536 Γ— 50 layers while PE-AV/A-Frame are width-1024 β€” cross-width tensors cannot be weight-averaged. L/14 (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video share width/depth/patch, so true weight fusion is possible.

Status (2026-10-03): CERTIFIED β€” frozen, no further weight search

Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:

model COCO t2i R@1 / R@5 COCO i2t R@1 ESC-50 acc
facebook/pe-av-base (anchor) 0.720 / 0.953 0.013 0.885
this tag 0.713 / 0.953 0.013 0.895

Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor quality while carrying A-Frame/Core/Spatial weights in one file β€” it does not beat the anchor on these benches by itself. Measured gains are deferred to alignment fine-tuning, for which this tag is the frozen init. Certification: 9-build grid + 18-genome evolution (fixed machinery) + full-slice validation, all agreeing (audio light, vision ~zero).

15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor: vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs 0.547; V2/V3 at chance for both β€” text embeds are dummy-video dominated in this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025 (fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas (J5 +0.033). No regression anywhere; consistent small audio edge. Full table: suite.json in the GitHub repo.

  • Evolution rerun in progress after two voided attempts (documented below) β€” if it certifies a better genome, that becomes the next revision.
  • Layer-depth probe: naive mid-network readout does NOT beat the joint output (t2i R@1 ≀ 0.06 vs 0.79) β€” mid-network features need trained per-layer pooling (fine-tune phase), not a free readout change.
  • Voided runs: (1) blend-of-blend contamination via shared-storage state_dict; (2) silent no-op load_state_dict on this model's dual-prefix key layout; (3) joint-path audio mismatch that faked ESC collapse at high audio weight. All fixed (pristine clones, copy_ by param name, standalone-tower-only audio mapping) before trusting any evolution number.

How it was built

  1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio (both pe_audio_encoder, 1024/16L/8H β€” 422/422 tensors matched).
  2. Video trunk: explicit native-PE β†’ timm key map (patch_embed, cls_token, pos_embed@336px, norms, QKV, MLP Γ—24 blocks). Pool/head/proj kept from anchor.
  3. Blend weights chosen by measured 3Γ—3 grid search over COCO-retrieval + ESC-50 slices (script search_pe_weights.py), not by vibes. Vision donor weight degraded COCO t2i monotonically (0.766 β†’ 0.64), so the winning tag keeps anchor vision and blends audio at 0.25.

Usage

from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
           return_tensors="pt")
out = m(**inputs)  # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...

Retrieval: dot-product video_embeds @ text_video_embeds.T. Zero-shot audio: argmax over audio_embeds @ text_audio_embeds.T with "the sound of {label}" prompts.

Honest limits

  • Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative anchor-vs-variant only β€” not official MMEB/PE-paper figures.
  • i2t (textβ†’image caption ranking over 150 similar COCO captions) sits near floor for anchor and variants alike; t2i and audio carry the signal.
  • Encoder only: no generation, detection, or segmentation heads.
  • Next: evolution-based per-tower weights, intermediate-layer readout (PE paper: best features are mid-network), GPU alignment fine-tune.

Reproduce

fuse_pe_single.py (fusion) Β· eval_pe_bench.py (COCO+ESC-50 bench) Β· search_pe_weights.py (grid) Β· evo_fuse.py (evolution) Β· layer_sweep.py (depth probe). Benchmark slice metadata: MMEB MSCOCO_i2t rows + COCO val2014 images + ashraq/esc50 streaming clips.

Downloads last month
14
Safetensors
Model size
1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results