Text Classification
Transformers
Safetensors
English
bert
intent-classification
customer-support
banking
banking77
Eval Results (legacy)
text-embeddings-inference
Instructions to use functionX86/banking77-intent-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use functionX86/banking77-intent-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="functionX86/banking77-intent-classifier")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("functionX86/banking77-intent-classifier") model = AutoModelForSequenceClassification.from_pretrained("functionX86/banking77-intent-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,092 Bytes
d2d3f85 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 | ---
license: mit
language: en
library_name: transformers
pipeline_tag: text-classification
base_model: BAAI/bge-small-en-v1.5
tags:
- intent-classification
- customer-support
- banking
- banking77
datasets:
- PolyAI/banking77
metrics:
- f1
- accuracy
model-index:
- name: banking77-intent-classifier
results:
- task:
type: text-classification
name: Intent Classification
dataset:
type: banking77
name: BANKING77
metrics:
- type: f1
name: Macro F1
value: 0.9245
- type: accuracy
name: Accuracy
value: 0.9247
---
# banking77-intent-classifier
A 77-class banking intent classifier, fine-tuned from
[`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) on
[BANKING77](https://github.com/PolyAI-LDN/task-specific-datasets).
Given a customer message such as *"my card still hasn't arrived after two weeks"*, it predicts
the intent (`card_arrival`) so the request can be routed automatically.
## Results
| Metric | Value |
|---|---|
| Macro F1 | **0.9245** |
| Accuracy | **0.9247** |
| Top-3 accuracy | 0.974 |
| Inference | ~0.34 ms per request |
Evaluated on the official BANKING77 test split (3080 requests, 40 per intent), scored
once, at the end. Model selection used a stratified 10 % validation split carved out of the
training data.
## An honest note on what this model is for
**This model was beaten by a simpler approach, and that is the interesting part.**
It was trained as the third rung of a deliberate ladder, to measure what fine-tuning actually
buys over cheaper alternatives on this dataset:
| Approach | Macro F1 | Training cost |
|---|---|---|
| TF-IDF (word + char n-grams) β logistic regression | 0.915 | 23 s, CPU |
| **Frozen `bge-small` embeddings β logistic regression** | **0.935** | 54 s, CPU |
| This model β `bge-small` fine-tuned end to end | 0.9245 | ~215 s, GPU |
Using the *same encoder frozen*, with nothing but a logistic regression on top, scores higher.
The gap held across five training runs spanning three random seeds, which scored between
0.9245 and 0.9307 (mean β 0.927). Runs vary by a few tenths of a point even at a fixed seed,
because GPU kernel scheduling and multi-worker data loading are not bit-deterministic β so the
comparison rests on the spread of runs rather than on any single number.
Two plausible reasons:
1. `bge-small` is contrastively pre-trained for semantic similarity. Grouping semantically
similar sentences is more or less what intent classification is, so its embedding space
already arrives close to the right shape β and fine-tuning distorts a geometry that was
already good.
2. 10 003 examples across 77 intents is roughly 130 per class. That is thin for updating 33 M
parameters, and the model reaches a memorised training loss before it generalises further.
An earlier version of this model, trained without a validation split, drove training loss to
0.037 and scored 0.9295 β marginally higher than the properly regularised model published
here. That version was overfit, and comparing it against a regularised alternative would have
proved nothing. The lower, honest number is the one reported.
**If you want the best model for this task, use frozen embeddings with a linear head.** This
checkpoint is published for reproducibility and as a documented negative result.
## Usage
```python
from transformers import pipeline
classifier = pipeline("text-classification", model="functionX86/banking77-intent-classifier")
classifier("my card still hasn't arrived after two weeks")
# [{'label': 'card_arrival', 'score': 0.98}]
```
## Training
| Setting | Value |
|---|---|
| Base model | `BAAI/bge-small-en-v1.5` (33 M parameters) |
| Max sequence length | 64 tokens |
| Epochs | up to 15, early stopping on validation macro-F1 (patience 3) |
| Batch size | 32 |
| Learning rate | 5e-5, 10 % warmup, weight decay 0.01 |
| Precision | fp16 |
| Hardware | one NVIDIA RTX 3050 Ti (4 GB) |
The 64-token cap comes from the data: the 95th percentile of BANKING77 requests is 29 words,
so it truncates almost nothing while running roughly four times faster than the default 256.
## Limitations
- **English only**, and trained on retail banking requests. It will not transfer to another
domain without retraining.
- **Several BANKING77 intents genuinely overlap** β `card_arrival` vs `card_delivery_estimate`,
`top_up_failed` vs `top_up_reverted`, and the whole identity-verification cluster. A share of
the residual error is label ambiguity that no model can resolve.
- **Raw softmax scores are not calibrated.** For any use that depends on a confidence
threshold, fit a temperature on held-out data first β on the frozen-embedding variant this
reduced expected calibration error from 0.110 to 0.012 without changing a single prediction.
- Trained on public research data, not on real customer messages, and never evaluated for
fairness across customer segments. Not suitable for production use as-is.
|