Mask Generation
PyTorch
ONNX
Safetensors
sam2
video-object-segmentation
knowledge-distillation
Pablopigue commited on
Commit
aeaa0d3
·
verified ·
1 Parent(s): c686201

Upload sam2-lite

Browse files
Files changed (1) hide show
  1. README.md +5 -4
README.md CHANGED
@@ -45,6 +45,7 @@ The encoder is faster but the pipeline gains little: the memory attention, which
45
 
46
  - `model.safetensors` + `config.yaml`: the whole tracker (student encoder, fine-tuned memory attention, SAM 2.1-tiny memory encoder and mask decoder) and what is needed to rebuild it.
47
  - `student_encoder_r1024.onnx`, `student_encoder_r576.onnx`: the image encoder for ONNX Runtime (CPU).
 
48
 
49
  ## Usage
50
 
@@ -81,11 +82,11 @@ mobile = load_bundle(
81
 
82
  sam2-lite is an independent project; it is not affiliated with, endorsed by or sponsored by Meta. "SAM 2" refers to the original model by Meta FAIR, on which this work is based.
83
 
84
- Released for **non-commercial research use only** (CC BY-NC 4.0):
85
 
86
- - It contains SAM 2.1 weights (memory encoder and mask decoder, plus a **modified**, fine-tuned memory attention) by Meta Platforms, Inc., licensed under the Apache License 2.0 (see `LICENSE` and `NOTICE`).
87
- - The encoder starts from timm weights pretrained on ImageNet-1k, whose terms allow only non-commercial research and educational use.
88
- - It was distilled on DAVIS 2017 frames, licensed under CC BY-NC 4.0. No DAVIS videos, frames or annotations are included.
89
 
90
  ## Citations
91
 
 
45
 
46
  - `model.safetensors` + `config.yaml`: the whole tracker (student encoder, fine-tuned memory attention, SAM 2.1-tiny memory encoder and mask decoder) and what is needed to rebuild it.
47
  - `student_encoder_r1024.onnx`, `student_encoder_r576.onnx`: the image encoder for ONNX Runtime (CPU).
48
+ - `LICENSE`, `LICENSE-APACHE-2.0`, `NOTICE`: see [License](#license).
49
 
50
  ## Usage
51
 
 
82
 
83
  sam2-lite is an independent project; it is not affiliated with, endorsed by or sponsored by Meta. "SAM 2" refers to the original model by Meta FAIR, on which this work is based.
84
 
85
+ Released under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) for **non-commercial research use only** (`LICENSE`); `NOTICE` details every part:
86
 
87
+ - It contains SAM 2.1 weights by Meta Platforms, Inc., licensed under the Apache License 2.0 (`LICENSE-APACHE-2.0`). **Modified:** the memory attention (fine-tuned to attend to 3 memory frames instead of 7) and the memory temporal encodings (3 of the 7 kept). The other SAM 2.1 weights (memory encoder, prompt encoder, mask decoder, object pointers) are unchanged and remain available under the Apache License 2.0.
88
+ - The image encoder starts from the timm weights [`mobilenetv4_conv_medium.e500_r224_in1k`](https://huggingface.co/timm/mobilenetv4_conv_medium.e500_r224_in1k) (Apache 2.0), **modified** by distillation. They were pretrained on ImageNet-1k, whose terms allow only non-commercial research and educational use.
89
+ - It was distilled on frames of [DAVIS 2017](https://davischallenge.org), licensed under CC BY-NC 4.0; about half of its sequences come from third-party sources (mostly YouTube) with their own terms. No DAVIS videos, frames or annotations are included.
90
 
91
  ## Citations
92