RAM GGUF

Current downloads: GitHub RAM release and this Hugging Face repository. The source project is recognize-anything-ggml.

Models

Model F32 F16 Q8_0
RAM Swin-L download download download
RAM++ Swin-L download download download
Tag2Text Swin-B download download download

The nine files are exported from the three official inference checkpoints and keep the original checkpoint stem in their filenames. GGML is pinned to v0.21.0; CUDA and Vulkan builds do not use cuDNN.

Embedded metadata

Current files are self-describing. RAM and RAM++ GGUFs contain ram.tag_list and ram.tag_thresholds; Tag2Text additionally contains ram.delete_tag_indices, the complete ram.bert_vocab WordPiece table, and ram.official_tag_thresholds. For Tag2Text, ram.tag_thresholds is the deployment table: only large and some are calibrated from 0.68 to 0.6795 to cover measured CUDA/Vulkan decision-boundary drift. The official values remain embedded for audit. The runtime therefore does not need a separate label list, threshold file or BERT vocabulary beside these files. The repository keeps readable text files for export, auditing and compatibility with older GGUF files.

Q8_0 uses block quantization for eligible visual matrices. The distributed Tag2Text Q8 export keeps Swin stages 1-3 in F16 and quantizes stage 0; text encoder and decoder weights remain F16. With the deployment threshold calibration, all six CUDA/Vulkan and F32/F16/Q8_0 combinations reproduced the official caption wording on all 50 retained GT images and all 18 repository images. An all-visual Q8 export remains a diagnostic variant and can drift in close decoder logits. F32/F16 remain the reference formats for tensor-level audits. See the detailed model card, benchmarks, and checksums.

Downloads last month
318
GGUF
Model size
0.3B params
Architecture
ram_plus
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support