File size: 5,092 Bytes
d2d3f85
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
---
license: mit
language: en
library_name: transformers
pipeline_tag: text-classification
base_model: BAAI/bge-small-en-v1.5
tags:
  - intent-classification
  - customer-support
  - banking
  - banking77
datasets:
  - PolyAI/banking77
metrics:
  - f1
  - accuracy
model-index:
  - name: banking77-intent-classifier
    results:
      - task:
          type: text-classification
          name: Intent Classification
        dataset:
          type: banking77
          name: BANKING77
        metrics:
          - type: f1
            name: Macro F1
            value: 0.9245
          - type: accuracy
            name: Accuracy
            value: 0.9247
---

# banking77-intent-classifier

A 77-class banking intent classifier, fine-tuned from
[`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) on
[BANKING77](https://github.com/PolyAI-LDN/task-specific-datasets).

Given a customer message such as *"my card still hasn't arrived after two weeks"*, it predicts
the intent (`card_arrival`) so the request can be routed automatically.

## Results

| Metric | Value |
|---|---|
| Macro F1 | **0.9245** |
| Accuracy | **0.9247** |
| Top-3 accuracy | 0.974 |
| Inference | ~0.34 ms per request |

Evaluated on the official BANKING77 test split (3080 requests, 40 per intent), scored
once, at the end. Model selection used a stratified 10 % validation split carved out of the
training data.

## An honest note on what this model is for

**This model was beaten by a simpler approach, and that is the interesting part.**

It was trained as the third rung of a deliberate ladder, to measure what fine-tuning actually
buys over cheaper alternatives on this dataset:

| Approach | Macro F1 | Training cost |
|---|---|---|
| TF-IDF (word + char n-grams) β†’ logistic regression | 0.915 | 23 s, CPU |
| **Frozen `bge-small` embeddings β†’ logistic regression** | **0.935** | 54 s, CPU |
| This model β€” `bge-small` fine-tuned end to end | 0.9245 | ~215 s, GPU |

Using the *same encoder frozen*, with nothing but a logistic regression on top, scores higher.
The gap held across five training runs spanning three random seeds, which scored between
0.9245 and 0.9307 (mean β‰ˆ 0.927). Runs vary by a few tenths of a point even at a fixed seed,
because GPU kernel scheduling and multi-worker data loading are not bit-deterministic β€” so the
comparison rests on the spread of runs rather than on any single number.

Two plausible reasons:

1. `bge-small` is contrastively pre-trained for semantic similarity. Grouping semantically
   similar sentences is more or less what intent classification is, so its embedding space
   already arrives close to the right shape β€” and fine-tuning distorts a geometry that was
   already good.
2. 10 003 examples across 77 intents is roughly 130 per class. That is thin for updating 33 M
   parameters, and the model reaches a memorised training loss before it generalises further.

An earlier version of this model, trained without a validation split, drove training loss to
0.037 and scored 0.9295 β€” marginally higher than the properly regularised model published
here. That version was overfit, and comparing it against a regularised alternative would have
proved nothing. The lower, honest number is the one reported.

**If you want the best model for this task, use frozen embeddings with a linear head.** This
checkpoint is published for reproducibility and as a documented negative result.

## Usage

```python
from transformers import pipeline

classifier = pipeline("text-classification", model="functionX86/banking77-intent-classifier")
classifier("my card still hasn't arrived after two weeks")
# [{'label': 'card_arrival', 'score': 0.98}]
```

## Training

| Setting | Value |
|---|---|
| Base model | `BAAI/bge-small-en-v1.5` (33 M parameters) |
| Max sequence length | 64 tokens |
| Epochs | up to 15, early stopping on validation macro-F1 (patience 3) |
| Batch size | 32 |
| Learning rate | 5e-5, 10 % warmup, weight decay 0.01 |
| Precision | fp16 |
| Hardware | one NVIDIA RTX 3050 Ti (4 GB) |

The 64-token cap comes from the data: the 95th percentile of BANKING77 requests is 29 words,
so it truncates almost nothing while running roughly four times faster than the default 256.

## Limitations

- **English only**, and trained on retail banking requests. It will not transfer to another
  domain without retraining.
- **Several BANKING77 intents genuinely overlap** β€” `card_arrival` vs `card_delivery_estimate`,
  `top_up_failed` vs `top_up_reverted`, and the whole identity-verification cluster. A share of
  the residual error is label ambiguity that no model can resolve.
- **Raw softmax scores are not calibrated.** For any use that depends on a confidence
  threshold, fit a temperature on held-out data first β€” on the frozen-embedding variant this
  reduced expected calibration error from 0.110 to 0.012 without changing a single prediction.
- Trained on public research data, not on real customer messages, and never evaluated for
  fairness across customer segments. Not suitable for production use as-is.