HaHaScore Cascade v8 β€” Sentence-Level Humor Strength Predictor

A multimodal sentence-level humor strength predictor (0–100 continuous score) that fuses text + audio through a Cascade Gate architecture, with Reddit pretraining for humor-style priors.

🎯 Quick Start

# Option 1: Use ONNX model (fast, no PyTorch)
import onnxruntime as ort
import numpy as np

session = ort.InferenceSession("v8_cascade_int8.onnx")
text = np.random.randn(1, 20, 768).astype(np.float32)    # text features
audio = np.random.randn(1, 20, 791).astype(np.float32)  # audio features
cross = np.zeros((1, 20, 16), dtype=np.float32)         # optional cross-modal

scores, conf, gate = session.run(None, {
    'text': text, 'audio': audio, 'cross': cross
})
print(f"Avg score: {scores.mean():.3f}, Max: {scores.max():.3f}")
# Option 2: PyTorch
import torch
from huggingface_hub import hf_hub_download

ckpt_path = hf_hub_download(repo_id="Hayasuki/hahascore-cascade",
                             filename="v8_finetuned.pt")
ckpt = torch.load(ckpt_path, map_location='cpu')

πŸ“Š Performance

Model Reddit Pretrain Val AUC CPU Inference
v7 baseline None 0.860 (5-fold CV) ~10ms
v8 v1 (smoke Reddit) 2K samples 0.807 ~13ms
v8 v1 (full Reddit) 30K samples 0.802 ~9ms

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Reddit DistilBERT (66M params)         β”‚  ← Continuous funniness pretraining
β”‚  Pretrained on 30K Reddit upvotes       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ 768-dim
              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Cascade Gate Fusion v8                 β”‚
β”‚  β€’ text_proj: 768β†’128                   β”‚
β”‚  β€’ audio_proj: 791β†’128 (WavLM/Prosody)  β”‚
β”‚  β€’ Multi-scale Attention                β”‚
β”‚  β€’ Cross-Modal Attention (4 heads)      β”‚
β”‚  β€’ Dynamic Gating (text-confidence)     β”‚
β”‚  β€’ BiGRU (2 layers, hidden=128)         β”‚
β”‚  β€’ Sigmoid Score Head                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ (20, 1) per segment
              β–Ό
       Humor Strength Score (0-1)

πŸ“¦ Training Pipeline

Stage 1: Data Acquisition (Google Drive, no local overflow)

  • 1M Reddit jokes (CC-BY-4.0) β€” continuous funniness via upvotes
  • ColBERT 200K (CC-BY-2.0) β€” text humor detection pretraining
  • 639 standup files with per-segment pseudo-labels

Stage 2: Reddit Pretrain (30K samples)

  • DistilBERT (66M params)
  • 1 epoch, batch=8, LR=2e-5, max_len=128
  • Streaming dataset (never materialized locally)
  • Final loss: 0.0177 (started at 0.0673, 74% reduction)

Stage 3: v8 Fine-tune (639 standup files)

  • Cascade Gate initialized with Reddit DistilBERT
  • 5 epochs, batch=4, LR=1e-4
  • Text encoder frozen, audio + gate trained
  • Val AUC: 0.802

Stage 4: ONNX Export

  • FP32: 0.3 MB, ~3.3ms CPU
  • INT8: 2.9 MB, ~3.3ms CPU (recommended)

πŸ”¬ Inputs/Outputs

Inputs

  • text: float32, shape (batch, 20, 768) β€” pre-extracted text features
  • audio: float32, shape (batch, 20, 791) β€” WavLM + prosody features
  • cross: float32, shape (batch, 20, 16) β€” optional cross-modal features

Outputs

  • scores: float32, shape (batch, 20, 1) β€” humor strength per segment [0-1]
  • text_confidence: float32, shape (batch, 20, 1)
  • gate_weight: float32, shape (batch, 20, 1)

πŸ“ Files in this Repository

File Description
v8_finetuned.pt PyTorch model (270 MB)
v8_cascade_fp32.onnx ONNX FP32 (0.3 MB)
v8_cascade_int8.onnx ONNX INT8 quantized (2.9 MB)
v8_finetuned.json Training metrics
v8_cascade_metadata.json Architecture spec
reddit_distilbert.json Pretrain metrics

πŸ“š Citation

@misc{hahascore-cascade-v8,
  author = {Das-rebel},
  title = {HaHaScore Cascade v8: Multimodal Humor Strength Predictor with Reddit Pretraining},
  year = {2026},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/Das-rebel/HaHaScore}},
}

πŸ“œ License

MIT License. Reddit data: CC-BY-4.0 (credit r/Jokes).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train Hayasuki/hahascore-cascade