Instructions to use pngwn/github-issue-classifier-distilbert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pngwn/github-issue-classifier-distilbert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="pngwn/github-issue-classifier-distilbert")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("pngwn/github-issue-classifier-distilbert") model = AutoModelForSequenceClassification.from_pretrained("pngwn/github-issue-classifier-distilbert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from pngwn/github-issue-classifier-distilbert: direct link, hf CLI and curl.
- Browser
- Download file 1.81 kB
-
https://huggingface.co/pngwn/github-issue-classifier-distilbert/resolve/main/README.md
- Command line
-
hf download hf://pngwn/github-issue-classifier-distilbert/README.md
-
curl -L -o README.md https://huggingface.co/pngwn/github-issue-classifier-distilbert/resolve/main/README.md
tags:
- text-classification
- github-issues
library_name: transformers
metrics:
- f1
- accuracy
base_model: distilbert-base-uncased
datasets:
- pngwn/github-issues-4class
license: mit
DistilBERT GitHub Issue Classifier (bug / feature / question / support)
Full fine-tune of distilbert-base-uncased (66M params): lr 2e-5, batch 16, 4 epochs, weight decay 0.01, 10% warmup, title+body concatenated and truncated to 256 tokens (the truncation used by the NLBSE'24 competition winners).
Training data: pngwn/github-issues-4class — 1,997 balanced issues (bug/feature/question from NLBSE'24 + support class sourced from maintainer-assigned GitHub labels).
Test-set results (1,997 balanced examples)
| metric | value |
|---|---|
| macro-F1 (4-class) | 0.794 |
| accuracy | 0.794 |
| macro-F1 on bug/feature/question subset | 0.779 |
| F1 bug | 0.787 |
| F1 feature | 0.788 |
| F1 question | 0.710 |
| F1 support | 0.891 |
Trails the smaller SetFit MiniLM model (pngwn/github-issue-classifier-setfit-minilm, 22M params, 0.807 macro-F1) on the same data — consistent with the NLBSE literature that contrastive SetFit few-shot recipes outperform straight fine-tunes at this scale.
Usage
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("pngwn/github-issue-classifier-distilbert")
model = AutoModelForSequenceClassification.from_pretrained("pngwn/github-issue-classifier-distilbert")
inputs = tok("App crashes when opening settings", return_tensors="pt")
preds = model(**inputs).logits.argmax(dim=-1)
Labels (id → name): 0 bug, 1 feature, 2 question, 3 support.