pngwn's picture
pngwn HF Staff
model card: DistilBERT 4-class GitHub issue classifier, results
3142cfe verified
|
Raw History Blame Contribute Delete
1.81 kB
---
tags:
- text-classification
- github-issues
library_name: transformers
metrics:
- f1
- accuracy
base_model: distilbert-base-uncased
datasets:
- pngwn/github-issues-4class
license: mit
---
# DistilBERT GitHub Issue Classifier (bug / feature / question / support)
Full fine-tune of `distilbert-base-uncased` (66M params): lr 2e-5, batch 16, 4 epochs, weight decay 0.01, 10% warmup, title+body concatenated and truncated to 256 tokens (the truncation used by the NLBSE'24 competition winners).
**Training data:** [pngwn/github-issues-4class](https://huggingface.co/datasets/pngwn/github-issues-4class) — 1,997 balanced issues (bug/feature/question from NLBSE'24 + support class sourced from maintainer-assigned GitHub labels).
## Test-set results (1,997 balanced examples)
| metric | value |
|---|---|
| macro-F1 (4-class) | 0.794 |
| accuracy | 0.794 |
| macro-F1 on bug/feature/question subset | 0.779 |
| F1 bug | 0.787 |
| F1 feature | 0.788 |
| F1 question | 0.710 |
| F1 support | 0.891 |
Trails the smaller SetFit MiniLM model ([pngwn/github-issue-classifier-setfit-minilm](https://huggingface.co/pngwn/github-issue-classifier-setfit-minilm), 22M params, 0.807 macro-F1) on the same data — consistent with the NLBSE literature that contrastive SetFit few-shot recipes outperform straight fine-tunes at this scale.
## Usage
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("pngwn/github-issue-classifier-distilbert")
model = AutoModelForSequenceClassification.from_pretrained("pngwn/github-issue-classifier-distilbert")
inputs = tok("App crashes when opening settings", return_tensors="pt")
preds = model(**inputs).logits.argmax(dim=-1)
```
Labels (id → name): 0 bug, 1 feature, 2 question, 3 support.