Instructions to use harpertoken/fuse with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harpertoken/fuse with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="harpertoken/fuse", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("harpertoken/fuse", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from harpertoken/fuse: direct link, hf CLI and curl.
- Browser
- Download file 4.64 kB
-
https://huggingface.co/harpertoken/fuse/resolve/main/README.md
- Command line
-
hf download hf://harpertoken/fuse/README.md
-
curl -L -o README.md https://huggingface.co/harpertoken/fuse/resolve/main/README.md
license: mit
language: en
library_name: transformers
tags:
- image-text-matching
- flickr8k
- cnn
- vision-language
datasets:
- flickr8k
metrics:
- accuracy
model-index:
- name: fuse
results:
- task:
type: image-text-matching
dataset:
name: Flickr8k
type: flickr8k
split: test
metrics:
- type: accuracy
value: 0.6933
fuse
An image-plus-caption matcher over Flickr8k: a photo and a sentence in, a match score out. The image encoder is a small CNN over 64 by 64 RGB (two conv layers, adaptive pooling, 128-unit projection). The text encoder embeds caption words and mean-pools them into a second 128-unit projection. Both sides L2-normalize and the score is a temperature-scaled dot product, the same arrangement as CLIP at small scale. A positive pair is an image with one of its own five human captions; a negative pairs an image with another image's caption. Test accuracy 0.6933 on the canonical test split. CPU training.
No transformer here. The PreTrainedModel wrapper exists only so the weights serialize as config.json plus model.safetensors and load through AutoModel (including trust_remote_code via auto_map), the same arrangement as wear, tone and pole.
Usage
import numpy as np
from transformers import AutoModel
from modeling_fuse import FuseMatcher # registers the architecture
model = AutoModel.from_pretrained("harpertoken/fuse")
model.eval()
image = np.zeros((64, 64, 3), dtype=np.float32)
print(model.match(image, "a dog runs on the beach", vocab))
match takes a 64 by 64 RGB float array in range 0 to 1, a caption string, and the vocabulary mapping, and returns the match probability with the predicted verdict. Training images were resized to 64 by 64 with no augmentation, matching inference exactly. Needs torch and transformers.
Examples
image ββ> image encoder ββ> project ββ> normalize βββ
ββ> dot x scale ββ> match score
caption β> text encoder ββ> project ββ> normalize βββ
Both sides project to 128 dimensions and L2-normalize before the dot product; there is no concatenation layer. The card first drafted with a concatenate diagram was wrong, and the text above describes the weights as saved.
Two pairs from the canonical test set, scored by the published weights. Selection rule: the first test image in file order with its first caption, and the same image with the first caption of the last test image. Image 1056338697_4f7d7ce270.jpg:
| Caption | Score | Verdict |
|---|---|---|
| A blond woman in a blue shirt appears to wait for a ride . | 0.819 | match, correct |
| A man in a pink shirt climbs a rock face | 0.938 | match, wrong |
The actual canonical test images above, not recreations: the match on the left, the misclassified non-match on the right, with filenames and scores as measured.
The non-match scores higher than the match and is classified wrong. That is not a chosen failure; it is what the rule produced, and it is consistent with 0.6933 test accuracy. These are two illustrations, not the metric.
Training
Canonical Flickr8k, 6,000 train images with five captions each (60,000 pairs after 1:1 negative sampling) and 1,000 test images (10,000 pairs). An earlier run silently trained on a partial download; the training script now asserts the full counts before starting. Word vocabulary of 4,433 from train captions with minimum frequency 2. Sixteen epochs, Adam 1e-3, batch 128, seed 0. Per-epoch test accuracy climbed from 0.53 to 0.69 and plateaued there.
Two things went wrong before it worked. A concat-plus-linear fusion head sat at chance for seven epochs; the dot-product head learned immediately, which is why the architecture is what it is. And the training loop read every JPEG from disk per sample, which dominated wall time until images were preloaded into RAM.
The wrapper was checked for exact equivalence: identical predictions on 300 test pairs in eval mode, after catching that dropout made train-mode outputs differ.
No dataset is published alongside this model. Flickr8k already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead.
Limitations
0.6933 is modest, and the card states it plainly rather than rounding up to a story. Matching needs fine distinctions this small model often misses; negatives that share words with the image are the hard cases. Anything outside one-photo-plus-one-English-sentence is out of scope.
