Thermal Caption Demo β Weights (v1.0)
Model weights for the Thermal Caption Demo v1.0: a CLIP + GPT image
captioning model with two domain branches (thermal / RGB), exported to ONNX
(FP32 and FP16) for CPU/GPU inference via onnxruntime, plus a COCO-pretrained
YOLOv8m checkpoint used for detection overlay in the demo UI.
Code and usage instructions: see the Phase3/Day39/ directory of the project
repo (download_weights.py pulls these files automatically).
Files
| File | Description |
|---|---|
clip_vision.onnx / .onnx.data |
CLIP ViT-B/32 vision encoder (FP32). Domain-agnostic β shared by both thermal and RGB branches. |
clip_vision.fp16.onnx / .onnx.data |
Same encoder, FP16 weights (keep_io_types=True, I/O stays float32). |
gpt.onnx / .onnx.data |
GPT caption decoder for the thermal branch (FP32). |
gpt.fp16.onnx / .onnx.data |
Thermal decoder, FP16 weights. |
gpt_rgb.onnx / .onnx.data |
GPT caption decoder for the RGB branch (FP32). |
gpt_rgb.fp16.onnx / .onnx.data |
RGB decoder, FP16 weights. |
tokenizer.pkl |
BPE tokenizer for the thermal branch (vocab_size=320). |
tokenizer_rgb.pkl |
BPE tokenizer for the RGB branch (vocab_size=320). |
yolov8m.pt |
Ultralytics YOLOv8m, COCO-pretrained (unmodified), used for detection overlay only β not fine-tuned for this project. |
Notes
- FP16 conversions were sanity-checked against their FP32 counterparts via cosine similarity on real inference outputs (~1.0 in all cases).
clip_vision.onnxis shared across domains: the training pipeline only consumes precomputed CLIP features and never fine-tunes the vision encoder, so one export covers both branches.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support