Thermal Caption Demo β€” Weights (v1.0)

Model weights for the Thermal Caption Demo v1.0: a CLIP + GPT image captioning model with two domain branches (thermal / RGB), exported to ONNX (FP32 and FP16) for CPU/GPU inference via onnxruntime, plus a COCO-pretrained YOLOv8m checkpoint used for detection overlay in the demo UI.

Code and usage instructions: see the Phase3/Day39/ directory of the project repo (download_weights.py pulls these files automatically).

Files

File Description
clip_vision.onnx / .onnx.data CLIP ViT-B/32 vision encoder (FP32). Domain-agnostic β€” shared by both thermal and RGB branches.
clip_vision.fp16.onnx / .onnx.data Same encoder, FP16 weights (keep_io_types=True, I/O stays float32).
gpt.onnx / .onnx.data GPT caption decoder for the thermal branch (FP32).
gpt.fp16.onnx / .onnx.data Thermal decoder, FP16 weights.
gpt_rgb.onnx / .onnx.data GPT caption decoder for the RGB branch (FP32).
gpt_rgb.fp16.onnx / .onnx.data RGB decoder, FP16 weights.
tokenizer.pkl BPE tokenizer for the thermal branch (vocab_size=320).
tokenizer_rgb.pkl BPE tokenizer for the RGB branch (vocab_size=320).
yolov8m.pt Ultralytics YOLOv8m, COCO-pretrained (unmodified), used for detection overlay only β€” not fine-tuned for this project.

Notes

  • FP16 conversions were sanity-checked against their FP32 counterparts via cosine similarity on real inference outputs (~1.0 in all cases).
  • clip_vision.onnx is shared across domains: the training pipeline only consumes precomputed CLIP features and never fine-tunes the vision encoder, so one export covers both branches.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support