IT Ticket Classifier

A DistilBERT model fine-tuned to classify IT/helpdesk support tickets into 8 categories, based on ticket description text.

Model Description

This model fine-tunes distilbert-base-uncased for multi-class text classification of IT support tickets. Given the free-text description of a ticket, it predicts which of 8 categories the ticket belongs to:

  • Access
  • Administrative rights
  • HR Support
  • Hardware
  • Internal Project
  • Miscellaneous
  • Purchase
  • Storage

Intended Use

This model classifies IT/helpdesk ticket text into 8 support categories, and is intended as a reusable baseline or starting point for similar ticket triage tasks. It was developed as a portfolio project rather than validated in a production environment โ€” see Known Limitations for specifics.

Not validated for:

  • Raw, unpreprocessed ticket text (see Training Data note below โ€” the training data had stopwords removed and was lemmatized; performance on natural, unprocessed text has not been tested)
  • Ticket categories or taxonomies different from the 8 listed above
  • Any specific organization's internal ticketing system or category scheme
  • Production deployment without further validation on your own data distribution

Training Data

Trained on the IT Service Ticket Classification Dataset (~47,800 tickets, CC0 licensed), sourced from Kaggle. Class distribution:

Category Count %
Hardware 13,617 28.5%
HR Support 10,915 22.8%
Access 7,125 14.9%
Miscellaneous 7,060 14.8%
Storage 2,777 5.8%
Purchase 2,464 5.2%
Internal Project 2,119 4.4%
Administrative rights 1,760 3.7%

Data was split 80/10/10 (train/val/test), stratified by category.

Preprocessing note: the ticket text in this dataset arrives pre-processed โ€” stopwords removed, lemmatized, punctuation stripped. This model was trained and evaluated on that preprocessed form. Performance on natural, unprocessed ticket text is untested and may differ.

Training Procedure

  • Base model: distilbert-base-uncased
  • Max sequence length: 192 tokens (chosen from per-category 95th-percentile token length analysis; Hardware tickets ran longest and drove this choice)
  • Learning rate: 2e-5
  • Batch size: 16 (train and eval)
  • Epochs: 3
  • Weight decay: 0.01
  • Best checkpoint selection: by macro F1 on validation set (load_best_model_at_end=True)

Macro F1 was used as the primary metric throughout, rather than accuracy, due to class imbalance (largest class is ~7.7x the size of the smallest).

Evaluation Results

Evaluated on a held-out test set (4,784 tickets, untouched during training/validation).

Category Precision Recall F1 Support
Access 0.90 0.92 0.91 713
Administrative rights 0.82 0.80 0.81 176
HR Support 0.90 0.89 0.89 1091
Hardware 0.87 0.88 0.87 1362
Internal Project 0.88 0.87 0.88 212
Miscellaneous 0.85 0.84 0.85 706
Purchase 0.95 0.92 0.94 247
Storage 0.92 0.91 0.91 277
Accuracy 0.88 4784
Macro avg 0.89 0.88 0.88 4784

Baseline Comparison

A TF-IDF (5,000 features) + logistic regression (class_weight='balanced') baseline was trained and evaluated on the same splits, for comparison:

Metric TF-IDF + LogReg DistilBERT (this model)
Accuracy 0.85 0.88
Macro F1 0.85 0.88
Macro Precision 0.83 0.89
Macro Recall 0.88 0.88

DistilBERT improves macro F1 by 3 points over the baseline, with the larger gain concentrated in precision โ€” the fine-tuned model produces fewer false positives on ambiguous or minority classes (notably Administrative rights), while achieving comparable recall to the baseline.

Known Limitations

  • Label noise in "Miscellaneous": manual inspection of a random sample of "Miscellaneous"-labeled tickets found a substantial portion (roughly a third or more) appeared mislabeled or arguably better suited to another existing category (commonly Access, HR Support, or Purchase), rather than being genuinely ambiguous. Performance metrics on this class should be interpreted with that in mind โ€” some model "errors" here likely reflect incorrect ground-truth labels rather than model failure.
  • Administrative rights / Hardware confusion: this is the model's most frequent specific confusion pattern (25 of 176 Administrative rights test tickets misclassified as Hardware). This likely reflects genuine vocabulary overlap in the source tickets โ€” requests for elevated permissions often reference installing software or configuring devices, language that also appears in genuine Hardware tickets.
  • Smaller classes have less training data: Administrative rights (the model's weakest-performing category) had the fewest training examples (~1,400) of any class, which likely contributes to its comparatively lower F1 (0.81) alongside the vocabulary-overlap issue above.

How to Use

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("nashfork/it-ticket-classifier")
model = AutoModelForSequenceClassification.from_pretrained("nashfork/it-ticket-classifier")

text = "unable to access shared drive, permission denied error"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=192, padding="max_length")

with torch.no_grad():
    outputs = model(**inputs)
    predicted_class_id = outputs.logits.argmax().item()

print(model.config.id2label[predicted_class_id])
Downloads last month
11
Safetensors
Model size
67M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nashfork/it-ticket-classifier

Finetuned
(12559)
this model