TubeCLIP / README.md
Krudev's picture
Update README.md
6a1c880 verified
|
Raw History Blame Contribute Delete
3.84 kB
---
language:
- en
license: cc-by-nc-4.0
tags:
- clip
- vision
- multimodal
- youtube
- thumbnail-analysis
pipeline_tag: feature-extraction
---
# TubeCLIP: AI-Driven YouTube Performance Predictor
**TubeCLIP is a fine-tuned variant of CLIP Large designed to analyze YouTube video thumbnails and titles to predict view performance classes.**
### Why?
Content creators and marketers constantly rely on **guesswork, gut feelings, and tedious A/B testing** to figure out which thumbnail and title combination will drive attention.
**TubeCLIP solves this by replacing intuition with data-driven prediction**. By fine-tuning a CLIP model, it evaluates the complex relationship between a thumbnail's visual elements and its title to predict its potential view tier.
**Prediction Demo**
##### (Coming Soon)
### Getting it Running
**Prerequisites**
Before installing, ensure your environment meets the following requirements:
* Python 3.x
* `pytorch-lightning`
* `torchao==0.16.0`
**Installation Guide**
You can install the required dependencies using pip:
```bash
pip install torch pytorch-lightning torchao==0.16.0 huggingface_hub transformers Pillow
```
**Usage & API Examples**
Here is a quick example of how to load the model weights from Hugging Face and run a prediction on your thumbnail and title.
```python
import torch
from huggingface_hub import snapshot_download
from PIL import Image
# 1. Download The Repository
model_path = snapshot_download(repo_id="Krudev/TubeCLIP", local_dir="/TubeCLIP"
)
```
Now run the cli with `predict.py`
```bash
python TubeCLIP.predict.py --model_path "/TubeCLIP/TubeCLIP.ckpt" --input_path "path/to/your/thumbnail/or/directory/of/thumbnails" --title "Your YouTube Video Title!"
```
---
### AI & Technical Specifics
**Performance Metrics & Results**
TubeCLIP achieves **66% accuracy** in classifying video performance across three distinct view tiers:
* **Tier 1:** 10k - 100k views
* **Tier 2:** 100k - 1M views
* **Tier 3:** 1M+ views
#### Evaluation Results
* **Test Loss:** 1.8045
* **Test Accuracy:** 0.6747 (67.47%)
#### Classification Report
| Class | Precision | Recall | F1-Score | Support |
| :--- | :---: | :---: | :---: | :---: |
| **10k-100k** | 0.70 | 0.72 | 0.71 | 1,606 |
| **100k-1M** | 0.61 | 0.59 | 0.60 | 1,608 |
| **1M+** | 0.72 | 0.71 | 0.71 | 1,511 |
| | | | | |
| **Accuracy** | | | **0.67** | 4,725 |
| **Macro Avg** | 0.67 | 0.68 | 0.68 | 4,725 |
| **Weighted Avg** | 0.67 | 0.67 | 0.67 | 4,725 |
<img src="cfmt.png" alt="Confusion Matrix" width="600"/>
**Limitations & Biases**
While the model is highly effective at distinguishing generally "good" (high potential) versus "bad" (low potential) thumbnail/title combinations, it is not an exact view-count calculator. Viewership relies on external factors (channel size, algorithmic luck, trending topics, time of day) that the model cannot see. Expect it to serve as a strong directional compass for A/B testing rather than a perfect view predictor.
**Datasets Used**
The model was trained on a custom-built, highly filtered, and strictly balanced dataset of **30,000 YouTube videos**. This curated dataset ensures the model learns pure visual-textual relationships without being overwhelmed by garbage data.
---
### MORE
**License**
This project is open-source but restricted for commercial use. It is licensed under the **Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)** license. You are free to use, modify, and build upon this tool for personal and research purposes, but you may not use it for commercial gains without permission.
**Support & Contact**
If you encounter bugs, have questions, or want to discuss collaboration, feel free to reach out:
* **Email:** krishnenduk462@gmail.com
* **X (Twitter):** @krishnendw
* **Instagram:** @contentbykrishnendu