SetFit GitHub Issue Classifier (bug / feature / question / support)

Contrastive fine-tune of sentence-transformers/all-MiniLM-L6-v2 (22M params) with a logistic-regression head, trained with SetFit 1.2.0 following the NLBSE issue-report-classification recipe (CosineSimilarityLoss, 20 pair-iterations, batch 16, 1 epoch, 256-token max length). The smallest model in the NLBSE winner family โ€” 3x smaller than DistilBERT and ~5x smaller than the MPNet used by the competition baseline, while scoring better here.

Training data: pngwn/github-issues-4class โ€” 1,997 balanced issues (bug/feature/question from NLBSE'24 + support class sourced from maintainer-assigned GitHub labels).

Test-set results (1,997 balanced examples)

metric value
macro-F1 (4-class) 0.807
accuracy 0.806
macro-F1 on bug/feature/question subset 0.780
F1 bug 0.800
F1 feature 0.789
F1 question 0.717
F1 support 0.922

Beat a DistilBERT-base fine-tune (66M params, 0.794 macro-F1) on the same data. For reference, published NLBSE'24 numbers on the 3-class benchmark: SetFit baseline 0.827, competition winners 0.84โ€“0.89 with much larger models (MPNet, RoBERTa+adapters).

Usage

from setfit import SetFitModel

model = SetFitModel.from_pretrained("pngwn/github-issue-classifier-setfit-minilm")
preds = model(["App crashes when opening settings", "How do I configure the API key?"])

Labels (id โ†’ name): 0 bug, 1 feature, 2 question, 3 support.

Downloads last month
25
Safetensors
Model size
22.7M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pngwn/github-issue-classifier-setfit-minilm

Dataset used to train pngwn/github-issue-classifier-setfit-minilm

Space using pngwn/github-issue-classifier-setfit-minilm 1