RAM GGUF
Current downloads: GitHub RAM release and this Hugging Face repository. The source project is recognize-anything-ggml.
Models
| Model | F32 | F16 | Q8_0 |
|---|---|---|---|
| RAM Swin-L | download | download | download |
| RAM++ Swin-L | download | download | download |
| Tag2Text Swin-B | download | download | download |
The nine files are exported from the three official inference checkpoints and keep the original checkpoint stem in their filenames. GGML is pinned to v0.21.0; CUDA and Vulkan builds do not use cuDNN.
Embedded metadata
Current files are self-describing. RAM and RAM++ GGUFs contain
ram.tag_list and ram.tag_thresholds; Tag2Text additionally contains
ram.delete_tag_indices, the complete ram.bert_vocab WordPiece table, and
ram.official_tag_thresholds. For Tag2Text, ram.tag_thresholds is the
deployment table: only large and some are calibrated from 0.68 to
0.6795 to cover measured CUDA/Vulkan decision-boundary drift. The official
values remain embedded for audit.
The runtime therefore does not need a separate label list, threshold file or
BERT vocabulary beside these files. The repository keeps readable text files
for export, auditing and compatibility with older GGUF files.
Q8_0 uses block quantization for eligible visual matrices. The distributed Tag2Text Q8 export keeps Swin stages 1-3 in F16 and quantizes stage 0; text encoder and decoder weights remain F16. With the deployment threshold calibration, all six CUDA/Vulkan and F32/F16/Q8_0 combinations reproduced the official caption wording on all 50 retained GT images and all 18 repository images. An all-visual Q8 export remains a diagnostic variant and can drift in close decoder logits. F32/F16 remain the reference formats for tensor-level audits. See the detailed model card, benchmarks, and checksums.
- Downloads last month
- 318
8-bit
16-bit
32-bit