Add FoundYou weights, config and model card

#1
by gberton - opened
Files changed (3) hide show
  1. README.md +140 -0
  2. config.json +17 -0
  3. model.safetensors +3 -0
README.md CHANGED
@@ -1,3 +1,143 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ library_name: pytorch
4
+ pipeline_tag: image-segmentation
5
+ base_model: facebook/sam2-hiera-small
6
+ base_model_relation: adapter
7
+ tags:
8
+ - personalized-segmentation
9
+ - personalized-retrieval
10
+ - image-retrieval
11
+ - segment-anything
12
+ - arxiv:2608.29917
13
  ---
14
+
15
+ # FoundYou: A Unified Model for Personalized Segmentation and Retrieval
16
+
17
+ [Paper](https://huggingface.co/papers/2608.29917) ([arXiv](https://arxiv.org/abs/2608.29917)) · [Code](https://github.com/ga1i13o/FoundYou) · [Project page](https://ga1i13o.github.io/FoundYou/) · [Demo](https://gmberton.github.io/demos-url/foundyou/)
18
+
19
+ **ECCV 2026** · Gabriele Trivigno\*, Marcos Alfaro\*, Claudia Cuttano\*, Gabriele Berton, Luis Payá, Carlo Masone (\* equal contribution)
20
+
21
+ ![FoundYou teaser](https://raw.githubusercontent.com/ga1i13o/FoundYou/main/assets/FoundYou_teaser.png)
22
+
23
+ Give FoundYou one example of your object: segment it in new images or retrieve it from a large database with a single efficient model. FoundYou keeps SAM 2-small frozen and uses its memory attention to match an object across independent images instead of across video frames. It adds lightweight adapters to the image encoder and a retrieval decoder that turns the match into an image-level score, while the frozen SAM 2 mask decoder produces the masks.
24
+
25
+ - **Flexible personalization:** use mask, box, or point prompts for segmentation and multiple references for few-shot retrieval
26
+ - **State-of-the-art performance:** improves over the prior unified method by +18.4 mIoU on PerMIS and +17.8 mAP on ILIAS
27
+ - **Compact and fast:** a 52 M-parameter model with only 5.9 M trainable parameters, over 75× faster and 20× smaller than the prior unified solution
28
+
29
+ ## Files
30
+
31
+ - `model.safetensors`: the 5.9 M trained parameters, i.e. the AdaptFormer adapters in the last two stages of the SAM 2 image encoder and the trained parts of the retrieval decoder (its transformer and object-score head). They were trained on UnED (Ypsilantis et al., ICCV 2023), with SAM 2 frozen.
32
+ - `config.json`: the model settings (the same values as `configs/*.yaml` in the code).
33
+
34
+ The frozen SAM 2-small weights are not in this repository: the code downloads them to `pretrain/` the first time the model is built (the same `sam2_hiera_small.pt` file as in [facebook/sam2-hiera-small](https://huggingface.co/facebook/sam2-hiera-small)).
35
+
36
+ ## Usage
37
+
38
+ ```bash
39
+ git clone https://github.com/ga1i13o/FoundYou && cd FoundYou
40
+ conda create --name foundyou python=3.10 -y && conda activate foundyou
41
+ pip install -r requirements.txt huggingface_hub safetensors
42
+ ```
43
+
44
+ Save the example below as `example.py` in the repository root and run `python example.py` there (or paste it into Python started in the repository root). It uses the images in `assets/`: a reference photo of a toy with a box around it, 10 other photos of the same toy, and the paper's teaser figure, which does not show the toy.
45
+
46
+ ```python
47
+ import os
48
+
49
+ import torch
50
+ from huggingface_hub import snapshot_download
51
+ from PIL import Image
52
+ from safetensors.torch import load_file
53
+ from torchvision import transforms
54
+
55
+ from datasets.transform_utils import load_box
56
+ from models.foundyou import build_foundyou
57
+ from util.promptable_utils import build_prompt_dict
58
+
59
+ device = "cuda" if torch.cuda.is_available() else "cpu"
60
+ weights = load_file(os.path.join(snapshot_download("gabTriv/FoundYou"), "model.safetensors"))
61
+
62
+
63
+ def load_model(config_path):
64
+ model = build_foundyou(config_path) # the first call downloads SAM 2-small to pretrain/
65
+ model.load_state_dict(weights, strict=False) # the file has only the trained parameters
66
+ return model.to(device).eval()
67
+
68
+
69
+ # Same weights, with the settings that the evaluation scripts use for each task
70
+ retrieval_model = load_model("configs/retrieval.yaml")
71
+ segmentation_model = load_model("configs/pers_seg.yaml")
72
+
73
+ # Same preprocessing as the evaluation scripts (the model resizes to 1024x1024 internally)
74
+ transform = transforms.Compose([
75
+ transforms.Resize((518, 518)),
76
+ transforms.ToTensor(),
77
+ transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
78
+ ])
79
+
80
+
81
+ def load_image(path):
82
+ return transform(Image.open(path).convert("RGB")).to(device)
83
+
84
+
85
+ # Reference: a photo of the object and a box around it, as [x1, y1, x2, y2] in the 518x518 frame
86
+ folder = "assets/spry_toodlesnap_863"
87
+ width, height = Image.open(f"{folder}/query/Q863_00.jpg").size
88
+ box_file = f"{folder}/query/Q863_00_bbox.txt" # "x y w h" in pixels
89
+ box = load_box(box_file, original_size=(height, width), transformed_size=(518, 518))
90
+ reference = load_image(f"{folder}/query/Q863_00.jpg")
91
+ prompt = build_prompt_dict(box, "box", device)
92
+ photos = [f"{folder}/positives/P863_{i:02d}.jpg" for i in range(10)]
93
+ not_the_toy = "assets/FoundYou_teaser.png" # the paper's teaser figure, without the toy
94
+
95
+ with torch.no_grad():
96
+ # Personalized retrieval: the probability that each photo shows the reference object
97
+ context = retrieval_model.encode_references([reference], [prompt])
98
+ for path in photos + [not_the_toy]:
99
+ logit = retrieval_model.score_candidates(load_image(path)[None], context) # batch of 1
100
+ print(f"{path}: {logit.sigmoid().item():.2f}")
101
+ # Personalized segmentation: the mask of the object in a new photo
102
+ context = segmentation_model.encode_references([reference], [prompt])
103
+ logits = segmentation_model.segment_candidates(load_image(photos[0])[None], context)
104
+ mask = logits.sigmoid() > 0.5 # [1, 518, 518]
105
+ print(f"The mask covers {mask.float().mean().item():.0%} of {photos[0]}")
106
+ ```
107
+
108
+ It prints a score between 0 and 1 for each image (higher means more likely to show the reference object): from 0.5 to 1.0 for the 10 photos of the toy, and about 0.05 for the teaser figure, which does not show it. Then it prints the share of the first photo covered by the predicted mask.
109
+
110
+ - The mask is in the 518x518 frame: resize it to the photo size to overlay it. For your own photos, make the box with `box = torch.tensor([x1, y1, x2, y2])` in the same frame (x times 518 / width, y times 518 / height) and pass it to `build_prompt_dict(box, "box", device)`.
111
+ - To use several reference photos of the same object, pass them all to `encode_references`, with one prompt each.
112
+ - For a point prompt, pass `{"prompt_type": "point", "prompt": {"point_coords": torch.tensor([[[x, y]]], dtype=torch.float32, device=device), "point_labels": torch.tensor([[1]], dtype=torch.int32, device=device)}}` instead of `prompt`, with (x, y) in the 518x518 frame. For mask prompts, see `inference_pers_seg.py`.
113
+ - For evaluation on PerSeg, PerMIS, PerMIR and ILIAS, see the [GitHub README](https://github.com/ga1i13o/FoundYou).
114
+
115
+ ## Results
116
+
117
+ Results from the paper (segmentation with mask prompts; ILIAS: re-ranking the top 1,000 images retrieved by SigLIP, which alone reaches 19.6 mAP). FoundYou runs at 90.2 images/s on an RTX 4090.
118
+
119
+ | Task | Benchmark | Metric | FoundYou |
120
+ |:--|:--|:--:|--:|
121
+ | Personalized segmentation | PerSeg | mIoU / bIoU | 96.4 / 85.6 |
122
+ | Personalized segmentation | PerMIS | mIoU / bIoU | 62.6 / 57.4 |
123
+ | Personalized retrieval | PerMIR | mAP | 92.1 |
124
+ | Personalized retrieval | ILIAS | mAP@1k | 32.5 |
125
+
126
+ ## Citation
127
+
128
+ ```bibtex
129
+ @inproceedings{trivigno2026foundyou,
130
+ title = {{FoundYou}: A Unified Model for Personalized Segmentation and Retrieval},
131
+ author = {Gabriele Trivigno and Marcos Alfaro and Claudia Cuttano and Gabriele Berton and Luis Pay{\'a} and Carlo Masone},
132
+ booktitle = {Computer Vision -- ECCV 2026},
133
+ pages = {585--604},
134
+ year = {2026},
135
+ publisher = {Springer Nature Switzerland},
136
+ address = {Cham},
137
+ doi = {10.1007/978-3-032-37041-9_31}
138
+ }
139
+ ```
140
+
141
+ ## License
142
+
143
+ Apache-2.0, like the [code](https://github.com/ga1i13o/FoundYou/blob/main/LICENSE). The SAM 2 weights that FoundYou builds on are also released under Apache-2.0 ([SAM 2 license](https://github.com/facebookresearch/sam2/blob/main/LICENSE)).
config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "foundyou",
3
+ "architectures": ["FoundYou"],
4
+ "sam2_version": "small",
5
+ "adaptformer_stages": [2, 3],
6
+ "channel_factor": 0.5,
7
+ "task_configs": {
8
+ "personalized_segmentation": {
9
+ "config_file": "configs/pers_seg.yaml",
10
+ "obj_score": false
11
+ },
12
+ "personalized_retrieval": {
13
+ "config_file": "configs/retrieval.yaml",
14
+ "obj_score": true
15
+ }
16
+ }
17
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a22c013ba13d41772feeb448fd72ef524d4545e9276ecfa15b93706e1334c2ef
3
+ size 23527316