fuse

An image-plus-caption matcher over Flickr8k: a photo and a sentence in, a match score out. The image encoder is a small CNN over 64 by 64 RGB (two conv layers, adaptive pooling, 128-unit projection). The text encoder embeds caption words and mean-pools them into a second 128-unit projection. Both sides L2-normalize and the score is a temperature-scaled dot product, the same arrangement as CLIP at small scale. A positive pair is an image with one of its own five human captions; a negative pairs an image with another image's caption. Test accuracy 0.6933 on the canonical test split. CPU training.

No transformer here. The PreTrainedModel wrapper exists only so the weights serialize as config.json plus model.safetensors and load through AutoModel (including trust_remote_code via auto_map), the same arrangement as wear, tone and pole.

Usage

import numpy as np
from transformers import AutoModel
from modeling_fuse import FuseMatcher  # registers the architecture

model = AutoModel.from_pretrained("harpertoken/fuse")
model.eval()
image = np.zeros((64, 64, 3), dtype=np.float32)
print(model.match(image, "a dog runs on the beach", vocab))

match takes a 64 by 64 RGB float array in range 0 to 1, a caption string, and the vocabulary mapping, and returns the match probability with the predicted verdict. Training images were resized to 64 by 64 with no augmentation, matching inference exactly. Needs torch and transformers.

Examples

image ──> image encoder ──> project ──> normalize ──┐
                                                    β”œβ”€> dot x scale ──> match score
caption ─> text encoder ──> project ──> normalize β”€β”€β”˜

Both sides project to 128 dimensions and L2-normalize before the dot product; there is no concatenation layer. The card first drafted with a concatenate diagram was wrong, and the text above describes the weights as saved.

Two pairs from the canonical test set, scored by the published weights. Selection rule: the first test image in file order with its first caption, and the same image with the first caption of the last test image. Image 1056338697_4f7d7ce270.jpg:

Caption Score Verdict
A blond woman in a blue shirt appears to wait for a ride . 0.819 match, correct
A man in a pink shirt climbs a rock face 0.938 match, wrong

Both test images with their scores

The actual canonical test images above, not recreations: the match on the left, the misclassified non-match on the right, with filenames and scores as measured.

The non-match scores higher than the match and is classified wrong. That is not a chosen failure; it is what the rule produced, and it is consistent with 0.6933 test accuracy. These are two illustrations, not the metric.

Training

Canonical Flickr8k, 6,000 train images with five captions each (60,000 pairs after 1:1 negative sampling) and 1,000 test images (10,000 pairs). An earlier run silently trained on a partial download; the training script now asserts the full counts before starting. Word vocabulary of 4,433 from train captions with minimum frequency 2. Sixteen epochs, Adam 1e-3, batch 128, seed 0. Per-epoch test accuracy climbed from 0.53 to 0.69 and plateaued there.

Two things went wrong before it worked. A concat-plus-linear fusion head sat at chance for seven epochs; the dot-product head learned immediately, which is why the architecture is what it is. And the training loop read every JPEG from disk per sample, which dominated wall time until images were preloaded into RAM.

The wrapper was checked for exact equivalence: identical predictions on 300 test pairs in eval mode, after catching that dropout made train-mode outputs differ.

No dataset is published alongside this model. Flickr8k already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead.

Limitations

0.6933 is modest, and the card states it plainly rather than rounding up to a story. Matching needs fine distinctions this small model often misses; negatives that share words with the image are the hard cases. Anything outside one-photo-plus-one-English-sentence is out of scope.

Downloads last month
43
Safetensors
Model size
869k params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results