Add files using upload-large-folder tool
Browse files- README.md +9 -9
- SHA256SUMS +24 -32
README.md
CHANGED
|
@@ -26,6 +26,8 @@ can stop after any number of tokens *k* per frame and still produce a semantical
|
|
| 26 |
This repository holds the paper's tokenizers and AR models for both SemanTok and the VideoFlexTok
|
| 27 |
baseline, on Kinetics-600 (class-to-video) and uCO3D (text-to-video).
|
| 28 |
|
|
|
|
|
|
|
| 29 |
Please note: For individuals or organizations generating annual revenue of US $1,000,000 (or local currency equivalent) or more, regardless of the source of that revenue, you must obtain an enterprise commercial license directly from Stability AI before commercially using SemanTok, derivative works of SemanTok, or outputs from SemanTok. See https://stability.ai/license and https://stability.ai/enterprise.
|
| 30 |
|
| 31 |
## Model Description
|
|
@@ -77,11 +79,11 @@ video = tok.decode(tokens, k=16) # decode from the first 16
|
|
| 77 |
|
| 78 |
## Files
|
| 79 |
|
| 80 |
-
Every model is a directory with `config.json` and `model.safetensors`:
|
| 81 |
|
| 82 |
```
|
| 83 |
tokenizers/{k600,uco3d}-{videoflextok,semantok}/
|
| 84 |
-
ar/{k600,uco3d}-{videoflextok,semantok}-d{10,12,16,20,24
|
| 85 |
```
|
| 86 |
|
| 87 |
| tokenizer | data | training |
|
|
@@ -89,18 +91,16 @@ ar/{k600,uco3d}-{videoflextok,semantok}-d{10,12,16,20,24,30,36}/
|
|
| 89 |
| `k600-videoflextok`, `k600-semantok` | Kinetics-600 | 200k steps (131B tokens) |
|
| 90 |
| `uco3d-videoflextok`, `uco3d-semantok` | uCO3D | 100k steps (66B tokens) |
|
| 91 |
|
| 92 |
-
| AR depth | d10 | d12 | d16 | d20 | d24 |
|
| 93 |
-
|---|---|---|---|---|---|
|
| 94 |
-
| parameters | 49M | 85M | 201M | 393M | 679M |
|
| 95 |
|
| 96 |
- Kinetics-600 AR models are class-conditioned (597 classes); uCO3D AR models are conditioned on umT5 caption embeddings.
|
| 97 |
-
-
|
| 98 |
-
-
|
| 99 |
- The tokenizers do not include the VidTok VAE when it is identical to the public one; the loader fetches it from [`EPFL-VILAB/videoflextok_d18_d18_k600`](https://huggingface.co/EPFL-VILAB/videoflextok_d18_d18_k600).
|
| 100 |
- `SHA256SUMS` lists the checksum of every `model.safetensors`.
|
| 101 |
|
| 102 |
-
**Training differences between the arms.** The VideoFlexTok tokenizers aggregated DINO features over a latent frame with its first frame and SemanTok with the frame average; this setting also sets the decoder-REPA target. The Kinetics-600 SemanTok tokenizer trained its first 15k steps without the per-frame class-token fix.
|
| 103 |
-
|
| 104 |
## License
|
| 105 |
|
| 106 |
- Community License: Free for research, non-commercial, and commercial use by organizations and individuals generating annual revenue of US $1,000,000 (or local currency equivalent) or less, regardless of the source of that revenue.
|
|
|
|
| 26 |
This repository holds the paper's tokenizers and AR models for both SemanTok and the VideoFlexTok
|
| 27 |
baseline, on Kinetics-600 (class-to-video) and uCO3D (text-to-video).
|
| 28 |
|
| 29 |
+

|
| 30 |
+
|
| 31 |
Please note: For individuals or organizations generating annual revenue of US $1,000,000 (or local currency equivalent) or more, regardless of the source of that revenue, you must obtain an enterprise commercial license directly from Stability AI before commercially using SemanTok, derivative works of SemanTok, or outputs from SemanTok. See https://stability.ai/license and https://stability.ai/enterprise.
|
| 32 |
|
| 33 |
## Model Description
|
|
|
|
| 79 |
|
| 80 |
## Files
|
| 81 |
|
| 82 |
+
Every model is a directory with `config.json` and `model.safetensors` (bf16):
|
| 83 |
|
| 84 |
```
|
| 85 |
tokenizers/{k600,uco3d}-{videoflextok,semantok}/
|
| 86 |
+
ar/{k600,uco3d}-{videoflextok,semantok}-d{10,12,16,20,24}/
|
| 87 |
```
|
| 88 |
|
| 89 |
| tokenizer | data | training |
|
|
|
|
| 91 |
| `k600-videoflextok`, `k600-semantok` | Kinetics-600 | 200k steps (131B tokens) |
|
| 92 |
| `uco3d-videoflextok`, `uco3d-semantok` | uCO3D | 100k steps (66B tokens) |
|
| 93 |
|
| 94 |
+
| AR depth | d10 | d12 | d16 | d20 | d24 |
|
| 95 |
+
|---|---|---|---|---|---|
|
| 96 |
+
| parameters | 49M | 85M | 201M | 393M | 679M |
|
| 97 |
|
| 98 |
- Kinetics-600 AR models are class-conditioned (597 classes); uCO3D AR models are conditioned on umT5 caption embeddings.
|
| 99 |
+
- All AR models here were trained for 20k steps. The paper's larger d30 (1.33B) and d36 (2.29B) models are not included.
|
| 100 |
+
- Weights are stored in bf16; the paper evaluated the same weights in fp32 under bf16 autocast.
|
| 101 |
- The tokenizers do not include the VidTok VAE when it is identical to the public one; the loader fetches it from [`EPFL-VILAB/videoflextok_d18_d18_k600`](https://huggingface.co/EPFL-VILAB/videoflextok_d18_d18_k600).
|
| 102 |
- `SHA256SUMS` lists the checksum of every `model.safetensors`.
|
| 103 |
|
|
|
|
|
|
|
| 104 |
## License
|
| 105 |
|
| 106 |
- Community License: Free for research, non-commercial, and commercial use by organizations and individuals generating annual revenue of US $1,000,000 (or local currency equivalent) or less, regardless of the source of that revenue.
|
SHA256SUMS
CHANGED
|
@@ -1,32 +1,24 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
fd02d410692ea07d12bdfa01a50c2a6f9a371ffb04b4687b13ed72a70520baad ./ar/uco3d-videoflextok-d20/model.safetensors
|
| 26 |
-
433cc4c5322be234c294ca85c854da8358b8d89d275abe83007842c423ff9069 ./ar/uco3d-videoflextok-d24/model.safetensors
|
| 27 |
-
8ecfc3b8813466b9eb349c92c937d201040415255aa1efbd5437a97cd69cbaac ./ar/uco3d-videoflextok-d30/model.safetensors
|
| 28 |
-
c73e06891e25d405e3edc39a350c60372279355ce24f0a690f5d689abc837d6f ./ar/uco3d-videoflextok-d36/model.safetensors
|
| 29 |
-
f7ea784c98bade273203978cf61f49e0a6848121b1586c3296075b6a770482bc ./tokenizers/k600-semantok/model.safetensors
|
| 30 |
-
af997a2a45b5da8bf1f6f4792a88271fe6a545928dd98e3be9cc4eb0526b1dd3 ./tokenizers/k600-videoflextok/model.safetensors
|
| 31 |
-
e4620539b16e4f7831fd104a5aac9a127111477b393153bb68ddb838307a38e8 ./tokenizers/uco3d-semantok/model.safetensors
|
| 32 |
-
92a7e721c052ce1dd7f62295578684589f892ed5d38201af0174f6e25872e292 ./tokenizers/uco3d-videoflextok/model.safetensors
|
|
|
|
| 1 |
+
788e9b21f9edfb7a38d898b0320e34ea0316145d40bda3fb0012483128e3db71 ar/k600-semantok-d10/model.safetensors
|
| 2 |
+
7ea94704e0781c12f361372e3154259817c19e59e34cde0fe650448d4b79d5ad ar/k600-semantok-d12/model.safetensors
|
| 3 |
+
b749f8a53cf0705902187e721c3c0f5e77299bd35f481709cbcfe04289d1d04f ar/k600-semantok-d16/model.safetensors
|
| 4 |
+
cfe3e91973646a94d705371caa97321b9763932a9820ecb6f1a3fcdadc4ae6dc ar/k600-semantok-d20/model.safetensors
|
| 5 |
+
0bc97e5668aafc564e4711985d09d9c5d32fa7f72e3461ed81c0c31e9265c869 ar/k600-semantok-d24/model.safetensors
|
| 6 |
+
e04c88292f27ce6a8fe7514acd2a0dca781f4e48a0aeb6334beee0226f8ceedd ar/k600-videoflextok-d10/model.safetensors
|
| 7 |
+
3577be44533cb48b118cd955a22878614f49df540ffab294879ea58804ec5ffa ar/k600-videoflextok-d12/model.safetensors
|
| 8 |
+
8c502553e206c4299d2cf79eb312a2ea9629915a78800d73c93984d3e0bb5333 ar/k600-videoflextok-d16/model.safetensors
|
| 9 |
+
725cbbdb3e34edb559fc7c4365ffbc09fcb20762f7727d2d0e1931044cd5a771 ar/k600-videoflextok-d20/model.safetensors
|
| 10 |
+
9639dcad36d369b3272ef0e0ec599a2dd69e59ae282a755b26f814b35cd87633 ar/k600-videoflextok-d24/model.safetensors
|
| 11 |
+
312a4d2e7f2b2885a6383c85d4e4978399f0caeeaef2112cac718e4d68872342 ar/uco3d-semantok-d10/model.safetensors
|
| 12 |
+
2ff38d216482af812042664944bbca9444cbd258b95364042b4ce017c62cd9cf ar/uco3d-semantok-d12/model.safetensors
|
| 13 |
+
5c300d6963b7be7507273e72459d3501203d6d0edea20d9c4ac4371aad45ffb5 ar/uco3d-semantok-d16/model.safetensors
|
| 14 |
+
9f74a5eca722e9b707146159b7f61dcaa54b71d9dc23f4e05b1951160ce08ad4 ar/uco3d-semantok-d20/model.safetensors
|
| 15 |
+
774803d4fdf7c920b9ab7669281070432bc953897a78f0dd0007d758bc90061d ar/uco3d-semantok-d24/model.safetensors
|
| 16 |
+
5af563fb32b3d2710f08f327d637755bb8a30b65a9e3b3340881e74e61316685 ar/uco3d-videoflextok-d10/model.safetensors
|
| 17 |
+
4bd033332df7e4f96b34c482ecf645a8ee91170dd9739ceeb7cd7ba88aa803e5 ar/uco3d-videoflextok-d12/model.safetensors
|
| 18 |
+
4f1aef2ec9a68a0cac490893eaa8e407519d8c00101c2c650e9f0cf192e14f15 ar/uco3d-videoflextok-d16/model.safetensors
|
| 19 |
+
64ef5331ead35bbf8b5792802c71861eacad7bc4527daabf6a016bc407a632ec ar/uco3d-videoflextok-d20/model.safetensors
|
| 20 |
+
40d50768abb4f4cb24ad34bdad22d6be2193b310acb4e04119819fe392dcd964 ar/uco3d-videoflextok-d24/model.safetensors
|
| 21 |
+
c58bcee2dee218418355dad8102380a9370b3af77a0cc6212d3f0a72231523c8 tokenizers/k600-semantok/model.safetensors
|
| 22 |
+
fc39ca1f3a4289c7b19e57ade7665b705a918651f673d5c7df2b87ab6600ca96 tokenizers/k600-videoflextok/model.safetensors
|
| 23 |
+
d26ee492afe0ea4b1be767f23cc059817c2f340e6878a7dc6bbf32221e870e14 tokenizers/uco3d-semantok/model.safetensors
|
| 24 |
+
50a3d430d5c8c80c60775e8bf71e742bf8f3078fc140075b2f7d6b946a7f04af tokenizers/uco3d-videoflextok/model.safetensors
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|