Instructions to use feyninc/multimatte with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- nobg
How to use feyninc/multimatte with nobg:
pip install nobg
# Option 1: use via the predict method from nobg import AutoModel, AutoProcessor model = AutoModel.from_pretrained("feyninc/multimatte").eval() processor = AutoProcessor.from_pretrained("feyninc/multimatte") cutout = model.predict(processor, "image.jpg", "prompt")# Option 2: use the model and processor directly import torch from loadimg import load_img from nobg import AutoModel, AutoProcessor model = AutoModel.from_pretrained("feyninc/multimatte").eval() processor = AutoProcessor.from_pretrained("feyninc/multimatte") image = load_img("image.jpg").convert("RGB") inputs = processor(image, return_tensors="pt") with torch.no_grad(): outputs = model(pixel_values=inputs["pixel_values"]) alpha = processor.post_process_alpha_matting(outputs, target_sizes=[(image.height, image.width)])[0] processor.cutout(image, alpha).save("output.png") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from feyninc/multimatte: direct link, hf CLI and curl.
- Browser
- Download file 7.57 kB
-
https://huggingface.co/feyninc/multimatte/resolve/main/README.md
- Command line
-
hf download hf://feyninc/multimatte/README.md
-
curl -L -o README.md https://huggingface.co/feyninc/multimatte/resolve/main/README.md
7.57 kB
| library_name: nobg | |
| license: apache-2.0 | |
| tags: | |
| - model_hub_mixin | |
| - nobg | |
| - nobg-sam3 | |
| - promptable | |
| - bbox | |
| - pytorch_model_hub_mixin | |
| <p align="center"> | |
| <img src="https://usefeyn.com/feyn/feyn_mark.svg"/> | |
| </p> | |
| <h1 align="center">MultiMatte</h1> | |
| <p align="center"><b>Cut anything you can name.</b></p> | |
| <p align="center"> | |
| <a href="https://usefeyn.com/blog/multimatte"><img src="https://img.shields.io/badge/Blog-Read%20the%20post-black"></a> | |
| <a href="https://github.com/feyninc/nobg"><img src="https://img.shields.io/badge/GitHub-nobg-black?logo=github"></a> | |
| <a href="https://pypi.org/project/nobg/"><img src="https://img.shields.io/pypi/v/nobg?color=blue&label=pip%20install%20nobg"></a> | |
| <a href="https://usefeyn.com/multimatte"><img src="https://img.shields.io/badge/Demo-usefeyn.com%2Fmultimatte-orange"></a> | |
| </p> | |
| MultiMatte is a **promptable matting model**: name an object and it returns an alpha matte for it. | |
| It is a LoRA fine-tune of [SAM 3](https://huggingface.co/facebook/sam3), retrained to produce | |
| continuous opacity instead of binary masks, with the adapter merged into the released weights. | |
| ## Prompt steering | |
| The same photograph, four prompts. | |
| <table> | |
| <tr> | |
| <td width="25%"><img src="https://usefeyn.com/blog/multimatte/dog-input.webp" width="100%" alt="A bulldog looking up at a bowl held by a person in jeans."></td> | |
| <td width="25%"><img src="https://usefeyn.com/blog/multimatte/dog-the-dog.webp" width="100%" alt="The dog cut out of the photograph."></td> | |
| <td width="25%"><img src="https://usefeyn.com/blog/multimatte/dog-the-bowl.webp" width="100%" alt="The metal bowl cut out of the same photograph."></td> | |
| <td width="25%"><img src="https://usefeyn.com/blog/multimatte/dog-the-jeans.webp" width="100%" alt="The person's jeans cut out of the same photograph."></td> | |
| </tr> | |
| <tr> | |
| <td align="center"><code>input</code></td> | |
| <td align="center"><code>"the dog"</code></td> | |
| <td align="center"><code>"the dog bowl"</code></td> | |
| <td align="center"><code>"the jeans"</code></td> | |
| </tr> | |
| </table> | |
| Plural concepts return every match in one matte, and small objects stay addressable. | |
| <table> | |
| <tr> | |
| <td width="33%"><img src="https://usefeyn.com/blog/multimatte/cats-input.webp" width="100%" alt="Two tabby cats sprawled on a pink couch beside two remote controls."></td> | |
| <td width="33%"><img src="https://usefeyn.com/blog/multimatte/cats-the-cats.webp" width="100%" alt="Both cats cut out together."></td> | |
| <td width="33%"><img src="https://usefeyn.com/blog/multimatte/cats-the-remote.webp" width="100%" alt="Only the two remote controls cut out."></td> | |
| </tr> | |
| <tr> | |
| <td align="center"><code>input</code></td> | |
| <td align="center"><code>"the cats"</code></td> | |
| <td align="center"><code>"the remote"</code></td> | |
| </tr> | |
| </table> | |
| The output is a continuous alpha matte, not a threshold, so edges hold up at 1:1 zoom. | |
| <table> | |
| <tr> | |
| <td width="50%"><img src="https://usefeyn.com/blog/multimatte/truck-crop-input.webp" width="100%" alt="Full-resolution crop of a white vehicle's front fender against a red wall."></td> | |
| <td width="50%"><img src="https://usefeyn.com/blog/multimatte/truck-crop-cutout.webp" width="100%" alt="The same crop with the red wall removed and the fender edge intact."></td> | |
| </tr> | |
| <tr> | |
| <td align="center"><code>input (crop)</code></td> | |
| <td align="center"><code>cutout (crop)</code></td> | |
| </tr> | |
| </table> | |
| ## Installation | |
| ```bash | |
| pip install nobg | |
| ``` | |
| ## Usage | |
| ```python | |
| from nobg import AutoModel, AutoProcessor | |
| model = AutoModel.from_pretrained("feyninc/multimatte") | |
| processor = AutoProcessor.from_pretrained("feyninc/multimatte") | |
| # Prompt-free: uses the processor's default_prompt ("the main foreground subject"). | |
| model.predict(processor, "photo.jpg").save("output.png") | |
| # Named concept. | |
| model.predict(processor, "photo.jpg", "the dog").save("dog.png") | |
| ``` | |
| `predict` runs the whole pipeline — load, preprocess, forward under `no_grad` in eval mode, | |
| post-process, composite — and returns an RGBA cutout at the input's original resolution. | |
| `image` accepts anything [`loadimg`](https://github.com/not-lain/loadimg) takes: a path, URL, | |
| base64 string, numpy array or PIL image. | |
| The signature is `predict(processor, image, prompt, boxes)`, everything optional after `image`. | |
| Useful keywords: `batch_size` (images per forward pass, default 1 to keep peak memory flat) and | |
| `return_type="alpha"` for the raw `(H, W)` matte tensor instead of a cutout. | |
| ## Results | |
| S-measure (`S_α`), prompt-free, higher is better. Both columns come from one scoring harness on | |
| identical rows, so the difference isolates the weights — the SAM 3 numbers are a fresh rescore, not | |
| values copied from a paper. Changes below 0.002 `S_α` are treated as measurement noise. | |
| <img src="https://cdn-uploads.huggingface.co/production/uploads/6527e89a8808d80ccff88b7a/x93Ond7tAw8aoLYkucUQn.png"/> | |
| † No sibling in the training mix. DAVIS-S and DUT-OMRON are the two fully cross-domain splits here, | |
| so they are the pair to read for generalization — and MultiMatte's best absolute score lands on | |
| DAVIS-S at 0.979. | |
| **Naming the concept helps, before and after training.** On DIS-VD, a real human-written phrase adds | |
| 0.150 `S_α` to base SAM 3 for zero gradient steps, and still adds 0.036 to MultiMatte after | |
| fine-tuning. Prompt supervision made the model better at both pathways rather than making it | |
| prompt-insensitive. | |
| ## Training | |
| | feature | detail | | |
| |:--|:--| | |
| | Base model | [`facebook/sam3`](https://huggingface.co/facebook/sam3), 0.86 B parameters | | |
| | Method | LoRA, rank 16, merged into the released weights | | |
| | Trainable | 19.49 M parameters — 2.27 % of the model | | |
| | Targets | Attention and MLP projections in every tower, including the CLIP text tower | | |
| | Objective | Focal loss + Dice loss (SAM 3's own semantic segmentation objective) | | |
| | Steps | 14,000 | | |
| | Data | 19,953 images: salient objects, camouflage, high-resolution subjects, hair, marine scenes | | |
| | Prompt supervision | 4,949 images (24.8 %) with human-written per-image concept phrases | | |
| | Input resolution | 1008 × 1008 | | |
| ## Citation | |
| ```bibtex | |
| @note{multimatte2026, | |
| title = {MultiMatte: Cut Out Anything You Can Name}, | |
| author = {Hichri, Hafedh and Feyn Research}, | |
| year = {2026}, | |
| venue = {Feyn Field Notes} | |
| } | |
| ``` | |
| Please also cite the base model and the adaptation method: | |
| ```bibtex | |
| @article{sam3, | |
| title={SAM 3: Segment Anything with Concepts}, | |
| author={Carion, Nicolas and Gustafson, Laura and Hu, Yuan-Ting and Debnath, Shoubhik and Hu, Ronghang and Suris, Didac and Ryali, Chaitanya and Alwala, Kalyan Vasudev and Khedr, Haitham and Huang, Andrew and Lei, Jie and Ma, Tengyu and Guo, Baishan and Marks, Markus and Greer, Joseph and Wang, Meng and Sun, Peize and R{\"a}dle, Roman and Afouras, Triantafyllos and Mavroudi, Effrosyni and Dollar, Piotr and Ravi, Nikhila and Saenko, Kate and Zhang, Pengchuan and Feichtenhofer, Christoph}, | |
| journal={arXiv preprint arXiv:2511.16719}, | |
| year={2025}, | |
| url={https://ai.meta.com/research/publications/sam-3-segment-anything-with-concepts/}, | |
| } | |
| @article{lora, | |
| title={LoRA: Low-Rank Adaptation of Large Language Models}, | |
| author={Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu}, | |
| journal={arXiv preprint arXiv:2106.09685}, | |
| year={2021}, | |
| } | |
| ``` | |
| ## Acknowledgements | |
| Built on Meta's SAM 3. FlowDIS supplied the human-written DIS5K phrases used for training and | |
| evaluation. Thinking Machines' LoRA analysis informed the adapter configuration. Thanks to the | |
| dataset authors whose released work made the training mix and evaluation possible. |