File size: 8,978 Bytes
bae15d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
# Holland Scoring Algorithm
## Kalim Career Guidance Chatbot

---

### Overview

This algorithm calculates the compatibility percentage between a student's
Holland (RIASEC) personality codes and a major's RIASEC profile.

It is implemented in `KalimRetriever.calculate_holland_score`
(`src/rag/retriever.py`). **There is no LLM anywhere in this path** — the
ranking a student sees is deterministic Python over
`data/processed/majors.json`. The language model only phrases answers to
free-form chat questions.

> The client hand-verified the output of this algorithm; the expected results
> are recorded in `Test_Results.docx` in the repository root. After changing
> anything on this page, re-run a student with codes `R-I-A` (track `الكل`, no
> language or location filter) and diff the top 10 against that document.

---

### Input

**Student:** up to 3 RIASEC codes ranked by preference (1st, 2nd, 3rd)
**Major:** a `holland_codes` dict with `primary`, `secondary` and `tertiary`

```python
score = retriever.calculate_holland_score(
    ['E', 'C', 'S'],                                        # student, ranked
    {'primary': 'E', 'secondary': 'C', 'tertiary': 'S'},    # major
)
```

> **The function is asymmetric.** Swapping the arguments produces a different
> number — the first list is ranked by the *student's* preference, the second by
> the *major's* requirements, and the two carry different weights. A swapped
> call does not raise; it silently returns a plausible-looking wrong score,
> which reads as a scoring regression rather than a caller bug.
>
> ```
> score(['R','I','A'], {R, A, S}) = 79.1
> score(['R','A','S'], {R, I, A}) = 76.7
> ```

---

### Weights

#### Major Position Weights
| Position | Weight | Reason |
|----------|--------|--------|
| Primary | 50 points | Core personality fit |
| Secondary | 30 points | Supporting trait |
| Tertiary | 20 points | Complementary trait |

#### Student Preference Multipliers
| Rank | Multiplier | Reason |
|------|------------|--------|
| 1st choice | 1.0 (100%) | Dominant trait |
| 2nd choice | 0.8 (80%) | Secondary trait |
| 3rd choice | 0.6 (60%) | Tertiary trait |

---

### Formula

```
For each student code that matches a major code:
    Score += Major_Weight × Student_Multiplier

Maximum Raw Score = 50×1.0 + 30×0.8 + 20×0.6 = 86

Final Percentage = (Raw Score / 86) × 100     rounded to 1 decimal, capped at 100
```

---

### Matching rules

The three rules below are not cosmetic — each was added to fix a real ranking
bug, and removing any of them lets a worse major outrank a better one.

1. **Duplicate student codes are dropped**, keeping the earliest (strongest)
   position. Without this, a student who answered `['E', 'E', 'S']` scores the
   major's `E` slot twice and outranks a genuine three-way match.

2. **Each major position is consumed at most once.** Once the major's
   `primary` has been matched, a later student code cannot match it again.
   Matching is greedy in student order: the student's 1st choice picks its slot
   first, then the 2nd, then the 3rd.

3. **An absent major code never matches.** A missing code is the empty string
   in the JSON, and `'' == ''` would otherwise make two blank fields a match
   worth 30 or 20 points.

Only the first **3 unique** student codes are scored. Empty inputs on either
side return `0.0`.

---

### Examples

#### Example 1: Perfect Match = 100%
- Student: [E, C, S]
- Major: Primary=E, Secondary=C, Tertiary=S

| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| C(2nd) → C(Secondary) | 30 × 0.8 | 24 |
| S(3rd) → S(Tertiary) | 20 × 0.6 | 12 |
| **Total** | 86/86 × 100 | **100%** |

#### Example 2: Out-of-order Match = 69.8%
- Student: [I, A, R]
- Major: Primary=R, Secondary=I, Tertiary=S

| Match | Calculation | Points |
|-------|-------------|--------|
| I(1st) → I(Secondary) | 30 × 1.0 | 30 |
| A(2nd) → No match | - | 0 |
| R(3rd) → R(Primary) | 50 × 0.6 | 30 |
| **Total** | 60/86 × 100 | **69.8%** |

#### Example 3: Single Match = 58.1%
- Student: [E, A, S]
- Major: Primary=E, Secondary=I, Tertiary=R

| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| A(2nd) → No match | - | 0 |
| S(3rd) → No match | - | 0 |
| **Total** | 50/86 × 100 | **58.1%** |

#### Example 4: Duplicate codes de-duplicated = 58.1%
- Student: [E, E, S]  →  scored as [E, S]
- Major: Primary=E, Secondary=I, Tertiary=R

| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| E(2nd) → *dropped as a duplicate* | - | 0 |
| S(now 2nd) → No match | - | 0 |
| **Total** | 50/86 × 100 | **58.1%** |

Without rule 1 the repeated `E` would have matched the primary slot a second
time for another 40 points, scoring 104.7% → capped at 100%.

---

### Score Interpretation

The percentage is displayed to the student as a number; it is **not** bucketed
into levels or colours in the UI. The only banding in the code is the wording
of the match reason (`_get_match_reasons`), and it has three bands, not five:

| Score | Arabic phrasing shown |
|-------|-----------------------|
| ≥ 50% | `توافق عالٍ مع شخصيتك المهنية` — high match |
| ≥ 30% | `توافق جيد مع شخصيتك المهنية` — good match |
| > 0%  | `توافق جزئي مع شخصيتك المهنية` — partial match |
| 0%    | no Holland reason is listed at all |

---

### Permutation expansion

`retrieve_with_expansion` exists so a student is never shown an empty result
list. The plain `retrieve` scores every major against the codes **in the order
the student ranked them**, which for an unusual combination can leave almost
nothing above the threshold.

```
1. retrieve(top_k × 3), keep results scoring ≥ 50
   (≥ 20 for a single-code student — one code in the tertiary slot
    maxes out at 20/86 ≈ 23%, so a 50% threshold would return nothing)

2. If that yields ≥ min_results (default 10), return them. Done.

3. Otherwise re-score every candidate that passed the track / language /
   location filters against ALL permutations of the student's codes,
   keeping the best score per major.

4. Sort descending, drop anything below 50, return top_k.
```

A single-code student skips step 3 — there is nothing to permute, and step 1's
lower threshold already covers all three positions.

**The best-score map is keyed on `(id, name_ar)`, not on either alone.** The raw
dataset contained a reused id, and two majors can legitimately share a name
across different faculties; either key on its own silently collapses distinct
majors into one entry.

Note that an expanded result reports its **best-permutation** score, not the
score for the order the student actually gave. So a major listed at 74% may be
a weaker fit for the student's stated ranking; expansion only runs when the
honest ranking could not fill the list.

---

### Design Rationale

1. **50/30/20 weighting**: The primary code defines the major's core nature, so
   matching it matters most.

2. **1.0/0.8/0.6 multipliers**: A student's first preference is their dominant
   trait, so matches against it are weighted higher.

3. **Normalized to 86**: Makes a perfect match land on exactly 100%.

4. **Greedy, single-consumption matching**: Keeps the score monotonic — adding a
   code to a student's profile can never lower the score of a major that
   already matched.

---

### Tests

`tests/test_retriever.py` covers the rules above (perfect match, asymmetry,
duplicate codes, empty major codes, ordering, and the expansion path):

```bash
python3 -m pytest tests/ -q
```

---

*Kalim Chatbot - Lebanese University Thesis Project*


## Two scores in the expansion path

When fewer than `min_results` majors clear the score threshold,
`retrieve_with_expansion` re-scores every candidate under all six permutations
of the student's three codes and keeps the best. That best-permutation value is
kept as `expansion_score` and is what ranks the results and passes the
threshold — it is the whole point of the expansion.

It is **not** what the student sees. Because `calculate_holland_score` is
asymmetric (the order of the student's codes carries meaning), showing the
best-permutation number answers a different question from the one the interface
asks. Each expanded result therefore also carries `holland_score`, computed
from the order the student actually gave, and that is what the cards, the
results table, "توصية كليم" and the LLM context all display. Expanded result
sets are labelled as such in the results intro.

Worked example — a student with codes `C-E-R` and track `فلسفة وانسانيات`
gets 17 results through the expansion path. The top one, الإدارة السياحية
(triplet `ECS`), ranks on `expansion_score` 86.0 but displays `holland_score`
81.4, the score for `C-E-R` as stated.