πŸ€– ZyroGod/flan-t5-small-summarization

This repository contains a Full Fine-Tuned Model fine-tuned on the CNN/DailyMail (v3.0.0) dataset for high-quality abstractive text summarization.

  • Base Architecture: google/flan-t5-small
  • Fine-Tuning Strategy: Full Fine-Tuning (100% Trainable)
  • Mixed Precision: Native Bfloat16 (bf16)
  • Optimizer: Adafactor
  • Evaluation Dataset: CNN/DailyMail (Test split)

πŸ“Š Evaluation Benchmark (CNN/DailyMail Test Set)

Metric Score
ROUGE-1 0.3881
ROUGE-2 0.1694
ROUGE-L 0.2690
ROUGE-Lsum 0.2690

πŸš€ Quick Start & Usage

import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

device = "cuda" if torch.cuda.is_available() else "cpu"
model_name = "ZyroGod/flan-t5-small-summarization"

# Load tokenizer and full model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name, torch_dtype=torch.bfloat16).to(device)
model.eval()

# Prepare input article with prompt prefix
article = (
    "The clock struck midnight, but the old radio kept playing a song that hadn't been "
    "broadcast in decades. From the dark hallway, soft footsteps approached the room. "
    "As the music swelled, a cold hand gently rested on his trembling shoulder."
)
input_text = "summarize: " + article

inputs = tokenizer(input_text, return_tensors="pt", max_length=512, truncation=True).to(device)

with torch.no_grad():
    output_ids = model.generate(
        **inputs,
        max_length=128,
        min_length=20,
        num_beams=4,
        length_penalty=2.0,
        no_repeat_ngram_size=3
    )

summary = tokenizer.decode(output_ids[0], skip_special_tokens=True)
print("Summary:", summary)

βš™οΈ Hardware & Training Optimization

  • Trained on an NVIDIA GeForce RTX 3050 Laptop GPU (4 GB VRAM).
  • Gradient Checkpointing enabled (use_reentrant=False).
  • Dynamic batch padding with label masking (-100).
Downloads last month
-
Safetensors
Model size
77M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ZyroGod/flan-t5-small-summarization

Adapter
(77)
this model

Dataset used to train ZyroGod/flan-t5-small-summarization