--- license: other license_name: mixed license_link: https://huggingface.co/pxlpshr/pytti-models/tree/main/licenses tags: - pytti - clip - vqgan - adabins - depth-estimation - optical-flow --- # PyTTI Portable models Mirror of the pretrained models that [PyTTI Portable](https://github.com/pxl-pshr/pytti) downloads on first use, and of the five Python packages its installer needs from GitHub. PyTTI Portable fetches each model from here first, checks its size and SHA-256, and falls back to the original source if this repo can't be reached. Its installer installs the packages from the wheels in [`wheels/`](#package-wheels), which pip checks against their SHA-256. Every model file is an unchanged copy of the original. The SHA-256 of each CLIP model is also part of its original download link, which OpenAI's `clip` package checks. ## Files | File | Size | Used for | Original source | License | |---|---|---|---|---| | `clip/40d36571…/ViT-B-32.pt` | 338 MB | CLIP ViT-B/32 | [OpenAI](https://openaipublic.azureedge.net/clip/models/40d365715913c9da98579312b702a82c18be219cc2a73407c4526f58eba950af/ViT-B-32.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/5806e77c…/ViT-B-16.pt` | 335 MB | CLIP ViT-B/16 | [OpenAI](https://openaipublic.azureedge.net/clip/models/5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f/ViT-B-16.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/b8cca3fd…/ViT-L-14.pt` | 890 MB | CLIP ViT-L/14 | [OpenAI](https://openaipublic.azureedge.net/clip/models/b8cca3fd41ae0c99ba7e8951adf17d267cdb84cd88be6f7c2e0eca1737a03836/ViT-L-14.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/3035c92b…/ViT-L-14-336px.pt` | 891 MB | CLIP ViT-L/14@336px | [OpenAI](https://openaipublic.azureedge.net/clip/models/3035c92b350959924f9f00213499208652fc7ea050643e8b385c2dac08641f02/ViT-L-14-336px.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/afeb0e10…/RN50.pt` | 244 MB | CLIP RN50 | [OpenAI](https://openaipublic.azureedge.net/clip/models/afeb0e10f9e5a86da6080e35cf09123aca3b358a0c3e3b6c78a7b63bc04b6762/RN50.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/8fa8567b…/RN101.pt` | 278 MB | CLIP RN101 | [OpenAI](https://openaipublic.azureedge.net/clip/models/8fa8567bab74a42d41c5915025a8e4538c3bdbe8804a470a72f30b0d94fab599/RN101.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/7e526bd1…/RN50x4.pt` | 402 MB | CLIP RN50x4 | [OpenAI](https://openaipublic.azureedge.net/clip/models/7e526bd135e493cef0776de27d5f42653e6b4c8bf9e0f653bb11773263205fdd/RN50x4.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/52378b40…/RN50x16.pt` | 630 MB | CLIP RN50x16 | [OpenAI](https://openaipublic.azureedge.net/clip/models/52378b407f34354e150460fe41077663dd5b39c54cd0bfd2b27167a4a06ec9aa/RN50x16.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `clip/be1cfb55…/RN50x64.pt` | 1.3 GB | CLIP RN50x64 | [OpenAI](https://openaipublic.azureedge.net/clip/models/be1cfb55d75a9666199fb2206c106743da0f6468c9d327f3e0d0a543a9919d9c/RN50x64.pt) | MIT ([CLIP](https://github.com/openai/CLIP)) | | `adabins/AdaBins_nyu.pt` | 897 MB | Depth for 3D mode | [deforum/AdaBins](https://huggingface.co/deforum/AdaBins) | GPL-3.0 ([AdaBins](https://github.com/shariqfarooq123/AdaBins)) | | `torch_hub/checkpoints/tf_efficientnet_b5_ap-9e82fae8.pth` | 117 MB | AdaBins encoder | [GitHub release](https://github.com/rwightman/pytorch-image-models/releases/download/v0.1-weights/tf_efficientnet_b5_ap-9e82fae8.pth) | Apache-2.0 ([gen-efficientnet-pytorch](https://github.com/rwightman/gen-efficientnet-pytorch)) | | `torch_hub/gen-efficientnet-pytorch-771ce082….zip` | 68 KB | AdaBins encoder code (torch.hub) | [GitHub, commit 771ce08](https://github.com/rwightman/gen-efficientnet-pytorch/tree/771ce082b2ce6d033f55b3d47c1f77389ad3c180) | Apache-2.0 | | `vqgan/coco_first_stage.yaml` | 1 KB | VQGAN coco config | [batbot.ai](http://batbot.ai/models/VQGAN/coco_first_stage.yaml) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/coco_first_stage.ckpt` | 282 MB | VQGAN coco | [batbot.ai](http://batbot.ai/models/VQGAN/coco_first_stage.ckpt) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/imagenet.yaml` | 1 KB | VQGAN imagenet config | [Heidelberg University](https://heibox.uni-heidelberg.de/f/274fb24ed38341bfa753/?dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/imagenet.ckpt` | 935 MB | VQGAN imagenet | [Heidelberg University](https://heibox.uni-heidelberg.de/f/867b05fc8c4841768640/?dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/wikiart.yaml` | 1 KB | VQGAN wikiart config | [pixray v1.7.1 release](https://github.com/pixray/pixray/releases/tag/v1.7.1) | see below | | `vqgan/wikiart.ckpt` | 959 MB | VQGAN wikiart | [pixray v1.7.1 release](https://github.com/pixray/pixray/releases/tag/v1.7.1) | see below | | `vqgan/sflckr.yaml` | 2 KB | VQGAN sflckr config | [Heidelberg University](https://heibox.uni-heidelberg.de/d/73487ab6e5314cb5adba/files/?p=%2Fconfigs%2F2020-11-09T13-31-51-project.yaml&dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/sflckr.ckpt` | 4.0 GB | VQGAN sflckr | [Heidelberg University](https://heibox.uni-heidelberg.de/d/73487ab6e5314cb5adba/files/?p=%2Fcheckpoints%2Flast.ckpt&dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/openimages.yaml` | 1 KB | VQGAN openimages config | [Heidelberg University](https://heibox.uni-heidelberg.de/d/2e5662443a6b4307b470/files/?p=%2Fconfigs%2Fmodel.yaml&dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | | `vqgan/openimages.ckpt` | 359 MB | VQGAN openimages | [Heidelberg University](https://heibox.uni-heidelberg.de/d/2e5662443a6b4307b470/files/?p=%2Fckpts%2Flast.ckpt&dl=1) | MIT ([taming-transformers](https://github.com/CompVis/taming-transformers)) | The wikiart model's original host, `eaidata.bmk.sh`, no longer responds. These files are the copy the [pixray](https://github.com/pixray/pixray) project published in its v1.7.1 release (January 2022). They are byte-identical to the copy in [Somnai/pytti-vqgan-checkpoints](https://huggingface.co/Somnai/pytti-vqgan-checkpoints), and every other VQGAN file here also matches that repo's `manifest.json`. The wikiart model is a community training on WikiArt with taming-transformers; no separate license was published with it. The license texts are in [`licenses/`](https://huggingface.co/pxlpshr/pytti-models/tree/main/licenses). ## Package wheels `install.bat` installs these five packages from the wheels below instead of from their GitHub repositories, so installing needs neither Git nor those repositories. Each wheel was built with `pip wheel --no-deps` from the commit shown, with the source unchanged. pytti-core is the version PyTTI Portable's patches are written for. | File | Size | Built from | License | |---|---|---|---| | `wheels/pyttitools_core-0.0.1-py3-none-any.whl` | 14 MB | [pytti-tools/pytti-core @ b5070aa](https://github.com/pytti-tools/pytti-core/tree/b5070aaeab05204f6eee0ff81c657bc486b9cdce) | MIT | | `wheels/pyttitools_adabins-0.0.1-py3-none-any.whl` | 34 KB | [pytti-tools/AdaBins @ 9b57712](https://github.com/pytti-tools/AdaBins/tree/9b57712c1257d95df2966018952f4b8151cdf88a), a fork of [shariqfarooq123/AdaBins](https://github.com/shariqfarooq123/AdaBins) | GPL-3.0 | | `wheels/pyttitools_gma-0.0.1-py3-none-any.whl` | 84 MB | [pytti-tools/GMA @ 27e8b4e](https://github.com/pytti-tools/GMA/tree/27e8b4ee10067a86f63523ef6a35b4566c296526), a fork of [zacjiang/GMA](https://github.com/zacjiang/GMA); includes GMA's four pretrained optical flow checkpoints | WTFPL | | `wheels/pyttitools_taming_transformers-0.0.1-py3-none-any.whl` | 66 KB | [pytti-tools/taming-transformers @ f44c0b1](https://github.com/pytti-tools/taming-transformers/tree/f44c0b1e5b15020054e5e1ca10c465d832fd911b), a fork of [CompVis/taming-transformers](https://github.com/CompVis/taming-transformers) | MIT | | `wheels/clip-1.0-py3-none-any.whl` | 1.3 MB | [openai/CLIP @ d05afc4](https://github.com/openai/CLIP/tree/d05afc436d78f1c48dc0dbf8e5980a9d471f35f6) | MIT | Each wheel contains its project's full Python source and license file. ## SHA-256 ``` 40d365715913c9da98579312b702a82c18be219cc2a73407c4526f58eba950af clip/40d365715913c9da98579312b702a82c18be219cc2a73407c4526f58eba950af/ViT-B-32.pt 5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f clip/5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f/ViT-B-16.pt b8cca3fd41ae0c99ba7e8951adf17d267cdb84cd88be6f7c2e0eca1737a03836 clip/b8cca3fd41ae0c99ba7e8951adf17d267cdb84cd88be6f7c2e0eca1737a03836/ViT-L-14.pt 3035c92b350959924f9f00213499208652fc7ea050643e8b385c2dac08641f02 clip/3035c92b350959924f9f00213499208652fc7ea050643e8b385c2dac08641f02/ViT-L-14-336px.pt afeb0e10f9e5a86da6080e35cf09123aca3b358a0c3e3b6c78a7b63bc04b6762 clip/afeb0e10f9e5a86da6080e35cf09123aca3b358a0c3e3b6c78a7b63bc04b6762/RN50.pt 8fa8567bab74a42d41c5915025a8e4538c3bdbe8804a470a72f30b0d94fab599 clip/8fa8567bab74a42d41c5915025a8e4538c3bdbe8804a470a72f30b0d94fab599/RN101.pt 7e526bd135e493cef0776de27d5f42653e6b4c8bf9e0f653bb11773263205fdd clip/7e526bd135e493cef0776de27d5f42653e6b4c8bf9e0f653bb11773263205fdd/RN50x4.pt 52378b407f34354e150460fe41077663dd5b39c54cd0bfd2b27167a4a06ec9aa clip/52378b407f34354e150460fe41077663dd5b39c54cd0bfd2b27167a4a06ec9aa/RN50x16.pt be1cfb55d75a9666199fb2206c106743da0f6468c9d327f3e0d0a543a9919d9c clip/be1cfb55d75a9666199fb2206c106743da0f6468c9d327f3e0d0a543a9919d9c/RN50x64.pt 3c917d1b86d058918d4055e70b2cdb9696ec4967bb2d8f05c0051263c1ac9641 adabins/AdaBins_nyu.pt 9e82fae840d76e8d3d64d363b3dca7678a597e1d086b9affdb78eb3f38a3da16 torch_hub/checkpoints/tf_efficientnet_b5_ap-9e82fae8.pth 9d37b77cc82c794d9818ba745124e455af31308f3c4422853eeca491b9a468c1 torch_hub/gen-efficientnet-pytorch-771ce082b2ce6d033f55b3d47c1f77389ad3c180.zip 17d0c2d9fda59eccade2a6070d833804b759dc817da42c2268be8cf1eb047676 vqgan/coco_first_stage.yaml 106cc20fde571df14afc4349d62218d8213cf47177c682ecaa78a8c28b12f9de vqgan/coco_first_stage.ckpt 00e2c6189926f1d89ecfef73e9598db77981c1982f0555fbade963ffd16143c7 vqgan/imagenet.yaml 845a68805098cb666420d5db93df53f3a3b6dd443e6dd85c05759c5b998cd663 vqgan/imagenet.ckpt 6e78241d2828ff35b8381839ce249c24a6c468ed429760d33a25e98e7f24499a vqgan/wikiart.yaml bd08bb46301f98be1712bb2be9f8868cea30b53137cd985d4d6e8da8b3e02c36 vqgan/wikiart.ckpt c27f012996f3f1f02580f063be47fd60bdd8528efb04d7318b2cbe8a86505b1b vqgan/sflckr.yaml 8a8adea3da8dab412675772831370dd948e0aa97bb11a9488a2f328b58dbc929 vqgan/sflckr.ckpt 91110bab325067b37f85d1e75895a837d59c292bae4f3097f80c518714d5caff vqgan/openimages.yaml 5cd6c74810ab97e00e942c25403f73afc081e8b19987b31ec0d9ff5b68e7ab14 vqgan/openimages.ckpt b8c5c4f5f3187cf5700861fc7b776f5a79a0e85e8f66ad5b7eb065c17d108267 wheels/pyttitools_core-0.0.1-py3-none-any.whl 0506d733c4cb5ca87684cafac16ee040f65c35cb377463bbbc657f9c7e2dc5d9 wheels/pyttitools_adabins-0.0.1-py3-none-any.whl 900f72586819230e05ac8fcb051ee5a0eb4cfb60de939f0bdf970a2edb63fc02 wheels/pyttitools_gma-0.0.1-py3-none-any.whl d0bf08df0053c68932ed965396bcde74468bddb89122dc39819dcc91b4663bf3 wheels/pyttitools_taming_transformers-0.0.1-py3-none-any.whl 24d6d72bf73d81012581a3add81344ff3f05d854161fd94f18dbef6274735a40 wheels/clip-1.0-py3-none-any.whl ``` ## Credits - CLIP: OpenAI, [Learning Transferable Visual Models From Natural Language Supervision](https://arxiv.org/abs/2103.00020) - AdaBins: Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka, [AdaBins: Depth Estimation using Adaptive Bins](https://arxiv.org/abs/2011.14141) - EfficientNet weights and gen-efficientnet-pytorch: Ross Wightman, ported from Google's TensorFlow EfficientNet (AdvProp) - VQGAN: Patrick Esser, Robin Rombach, Björn Ommer, [Taming Transformers for High-Resolution Image Synthesis](https://arxiv.org/abs/2012.09841) - wikiart VQGAN: a community training on WikiArt, preserved by [pixray](https://github.com/pixray/pixray) (Tom White) and [Somnai/pytti-vqgan-checkpoints](https://huggingface.co/Somnai/pytti-vqgan-checkpoints) - GMA: Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, Richard Hartley, [Learning to Estimate Hidden Motions with Global Motion Aggregation](https://arxiv.org/abs/2104.02409) - [pytti-core](https://github.com/pytti-tools/pytti-core): sportsracer48 and the pytti-tools contributors