--- license: mit language: en library_name: transformers tags: - image-text-matching - flickr8k - cnn - vision-language datasets: - flickr8k metrics: - accuracy model-index: - name: fuse results: - task: type: image-text-matching dataset: name: Flickr8k type: flickr8k split: test metrics: - type: accuracy value: 0.6933 --- # fuse An image-plus-caption matcher over Flickr8k: a photo and a sentence in, a match score out. The image encoder is a small CNN over 64 by 64 RGB (two conv layers, adaptive pooling, 128-unit projection). The text encoder embeds caption words and mean-pools them into a second 128-unit projection. Both sides L2-normalize and the score is a temperature-scaled dot product, the same arrangement as CLIP at small scale. A positive pair is an image with one of its own five human captions; a negative pairs an image with another image's caption. Test accuracy 0.6933 on the canonical test split. CPU training. No transformer here. The `PreTrainedModel` wrapper exists only so the weights serialize as `config.json` plus `model.safetensors` and load through `AutoModel` (including `trust_remote_code` via `auto_map`), the same arrangement as `wear`, `tone` and `pole`. ## Usage ```python import numpy as np from transformers import AutoModel from modeling_fuse import FuseMatcher # registers the architecture model = AutoModel.from_pretrained("harpertoken/fuse") model.eval() image = np.zeros((64, 64, 3), dtype=np.float32) print(model.match(image, "a dog runs on the beach", vocab)) ``` `match` takes a 64 by 64 RGB float array in range 0 to 1, a caption string, and the vocabulary mapping, and returns the match probability with the predicted verdict. Training images were resized to 64 by 64 with no augmentation, matching inference exactly. Needs `torch` and `transformers`. ## Examples ```text image ──> image encoder ──> project ──> normalize ──┐ ├─> dot x scale ──> match score caption ─> text encoder ──> project ──> normalize ──┘ ``` Both sides project to 128 dimensions and L2-normalize before the dot product; there is no concatenation layer. The card first drafted with a concatenate diagram was wrong, and the text above describes the weights as saved. Two pairs from the canonical test set, scored by the published weights. Selection rule: the first test image in file order with its first caption, and the same image with the first caption of the last test image. Image `1056338697_4f7d7ce270.jpg`: | Caption | Score | Verdict | |---|---|---| | A blond woman in a blue shirt appears to wait for a ride . | 0.819 | match, correct | | A man in a pink shirt climbs a rock face | 0.938 | match, wrong | ![Both test images with their scores](fuse_examples.png) The actual canonical test images above, not recreations: the match on the left, the misclassified non-match on the right, with filenames and scores as measured. The non-match scores higher than the match and is classified wrong. That is not a chosen failure; it is what the rule produced, and it is consistent with 0.6933 test accuracy. These are two illustrations, not the metric. ## Training Canonical Flickr8k, 6,000 train images with five captions each (60,000 pairs after 1:1 negative sampling) and 1,000 test images (10,000 pairs). An earlier run silently trained on a partial download; the training script now asserts the full counts before starting. Word vocabulary of 4,433 from train captions with minimum frequency 2. Sixteen epochs, Adam 1e-3, batch 128, seed 0. Per-epoch test accuracy climbed from 0.53 to 0.69 and plateaued there. Two things went wrong before it worked. A concat-plus-linear fusion head sat at chance for seven epochs; the dot-product head learned immediately, which is why the architecture is what it is. And the training loop read every JPEG from disk per sample, which dominated wall time until images were preloaded into RAM. The wrapper was checked for exact equivalence: identical predictions on 300 test pairs in eval mode, after catching that dropout made train-mode outputs differ. No dataset is published alongside this model. Flickr8k already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead. ## Limitations 0.6933 is modest, and the card states it plainly rather than rounding up to a story. Matching needs fine distinctions this small model often misses; negatives that share words with the image are the hard cases. Anything outside one-photo-plus-one-English-sentence is out of scope.