IT Ticket Classifier
A DistilBERT model fine-tuned to classify IT/helpdesk support tickets into 8 categories, based on ticket description text.
Model Description
This model fine-tunes distilbert-base-uncased for multi-class text classification of IT support tickets. Given the free-text description of a ticket, it predicts which of 8 categories the ticket belongs to:
- Access
- Administrative rights
- HR Support
- Hardware
- Internal Project
- Miscellaneous
- Purchase
- Storage
Intended Use
This model classifies IT/helpdesk ticket text into 8 support categories, and is intended as a reusable baseline or starting point for similar ticket triage tasks. It was developed as a portfolio project rather than validated in a production environment โ see Known Limitations for specifics.
Not validated for:
- Raw, unpreprocessed ticket text (see Training Data note below โ the training data had stopwords removed and was lemmatized; performance on natural, unprocessed text has not been tested)
- Ticket categories or taxonomies different from the 8 listed above
- Any specific organization's internal ticketing system or category scheme
- Production deployment without further validation on your own data distribution
Training Data
Trained on the IT Service Ticket Classification Dataset (~47,800 tickets, CC0 licensed), sourced from Kaggle. Class distribution:
| Category | Count | % |
|---|---|---|
| Hardware | 13,617 | 28.5% |
| HR Support | 10,915 | 22.8% |
| Access | 7,125 | 14.9% |
| Miscellaneous | 7,060 | 14.8% |
| Storage | 2,777 | 5.8% |
| Purchase | 2,464 | 5.2% |
| Internal Project | 2,119 | 4.4% |
| Administrative rights | 1,760 | 3.7% |
Data was split 80/10/10 (train/val/test), stratified by category.
Preprocessing note: the ticket text in this dataset arrives pre-processed โ stopwords removed, lemmatized, punctuation stripped. This model was trained and evaluated on that preprocessed form. Performance on natural, unprocessed ticket text is untested and may differ.
Training Procedure
- Base model:
distilbert-base-uncased - Max sequence length: 192 tokens (chosen from per-category 95th-percentile token length analysis; Hardware tickets ran longest and drove this choice)
- Learning rate: 2e-5
- Batch size: 16 (train and eval)
- Epochs: 3
- Weight decay: 0.01
- Best checkpoint selection: by macro F1 on validation set (
load_best_model_at_end=True)
Macro F1 was used as the primary metric throughout, rather than accuracy, due to class imbalance (largest class is ~7.7x the size of the smallest).
Evaluation Results
Evaluated on a held-out test set (4,784 tickets, untouched during training/validation).
| Category | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| Access | 0.90 | 0.92 | 0.91 | 713 |
| Administrative rights | 0.82 | 0.80 | 0.81 | 176 |
| HR Support | 0.90 | 0.89 | 0.89 | 1091 |
| Hardware | 0.87 | 0.88 | 0.87 | 1362 |
| Internal Project | 0.88 | 0.87 | 0.88 | 212 |
| Miscellaneous | 0.85 | 0.84 | 0.85 | 706 |
| Purchase | 0.95 | 0.92 | 0.94 | 247 |
| Storage | 0.92 | 0.91 | 0.91 | 277 |
| Accuracy | 0.88 | 4784 | ||
| Macro avg | 0.89 | 0.88 | 0.88 | 4784 |
Baseline Comparison
A TF-IDF (5,000 features) + logistic regression (class_weight='balanced') baseline was trained and evaluated on the same splits, for comparison:
| Metric | TF-IDF + LogReg | DistilBERT (this model) |
|---|---|---|
| Accuracy | 0.85 | 0.88 |
| Macro F1 | 0.85 | 0.88 |
| Macro Precision | 0.83 | 0.89 |
| Macro Recall | 0.88 | 0.88 |
DistilBERT improves macro F1 by 3 points over the baseline, with the larger gain concentrated in precision โ the fine-tuned model produces fewer false positives on ambiguous or minority classes (notably Administrative rights), while achieving comparable recall to the baseline.
Known Limitations
- Label noise in "Miscellaneous": manual inspection of a random sample of "Miscellaneous"-labeled tickets found a substantial portion (roughly a third or more) appeared mislabeled or arguably better suited to another existing category (commonly Access, HR Support, or Purchase), rather than being genuinely ambiguous. Performance metrics on this class should be interpreted with that in mind โ some model "errors" here likely reflect incorrect ground-truth labels rather than model failure.
- Administrative rights / Hardware confusion: this is the model's most frequent specific confusion pattern (25 of 176 Administrative rights test tickets misclassified as Hardware). This likely reflects genuine vocabulary overlap in the source tickets โ requests for elevated permissions often reference installing software or configuring devices, language that also appears in genuine Hardware tickets.
- Smaller classes have less training data: Administrative rights (the model's weakest-performing category) had the fewest training examples (~1,400) of any class, which likely contributes to its comparatively lower F1 (0.81) alongside the vocabulary-overlap issue above.
How to Use
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("nashfork/it-ticket-classifier")
model = AutoModelForSequenceClassification.from_pretrained("nashfork/it-ticket-classifier")
text = "unable to access shared drive, permission denied error"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=192, padding="max_length")
with torch.no_grad():
outputs = model(**inputs)
predicted_class_id = outputs.logits.argmax().item()
print(model.config.id2label[predicted_class_id])
- Downloads last month
- 11
Model tree for nashfork/it-ticket-classifier
Base model
distilbert/distilbert-base-uncased