File size: 8,978 Bytes
bae15d1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 | # Holland Scoring Algorithm
## Kalim Career Guidance Chatbot
---
### Overview
This algorithm calculates the compatibility percentage between a student's
Holland (RIASEC) personality codes and a major's RIASEC profile.
It is implemented in `KalimRetriever.calculate_holland_score`
(`src/rag/retriever.py`). **There is no LLM anywhere in this path** — the
ranking a student sees is deterministic Python over
`data/processed/majors.json`. The language model only phrases answers to
free-form chat questions.
> The client hand-verified the output of this algorithm; the expected results
> are recorded in `Test_Results.docx` in the repository root. After changing
> anything on this page, re-run a student with codes `R-I-A` (track `الكل`, no
> language or location filter) and diff the top 10 against that document.
---
### Input
**Student:** up to 3 RIASEC codes ranked by preference (1st, 2nd, 3rd)
**Major:** a `holland_codes` dict with `primary`, `secondary` and `tertiary`
```python
score = retriever.calculate_holland_score(
['E', 'C', 'S'], # student, ranked
{'primary': 'E', 'secondary': 'C', 'tertiary': 'S'}, # major
)
```
> **The function is asymmetric.** Swapping the arguments produces a different
> number — the first list is ranked by the *student's* preference, the second by
> the *major's* requirements, and the two carry different weights. A swapped
> call does not raise; it silently returns a plausible-looking wrong score,
> which reads as a scoring regression rather than a caller bug.
>
> ```
> score(['R','I','A'], {R, A, S}) = 79.1
> score(['R','A','S'], {R, I, A}) = 76.7
> ```
---
### Weights
#### Major Position Weights
| Position | Weight | Reason |
|----------|--------|--------|
| Primary | 50 points | Core personality fit |
| Secondary | 30 points | Supporting trait |
| Tertiary | 20 points | Complementary trait |
#### Student Preference Multipliers
| Rank | Multiplier | Reason |
|------|------------|--------|
| 1st choice | 1.0 (100%) | Dominant trait |
| 2nd choice | 0.8 (80%) | Secondary trait |
| 3rd choice | 0.6 (60%) | Tertiary trait |
---
### Formula
```
For each student code that matches a major code:
Score += Major_Weight × Student_Multiplier
Maximum Raw Score = 50×1.0 + 30×0.8 + 20×0.6 = 86
Final Percentage = (Raw Score / 86) × 100 rounded to 1 decimal, capped at 100
```
---
### Matching rules
The three rules below are not cosmetic — each was added to fix a real ranking
bug, and removing any of them lets a worse major outrank a better one.
1. **Duplicate student codes are dropped**, keeping the earliest (strongest)
position. Without this, a student who answered `['E', 'E', 'S']` scores the
major's `E` slot twice and outranks a genuine three-way match.
2. **Each major position is consumed at most once.** Once the major's
`primary` has been matched, a later student code cannot match it again.
Matching is greedy in student order: the student's 1st choice picks its slot
first, then the 2nd, then the 3rd.
3. **An absent major code never matches.** A missing code is the empty string
in the JSON, and `'' == ''` would otherwise make two blank fields a match
worth 30 or 20 points.
Only the first **3 unique** student codes are scored. Empty inputs on either
side return `0.0`.
---
### Examples
#### Example 1: Perfect Match = 100%
- Student: [E, C, S]
- Major: Primary=E, Secondary=C, Tertiary=S
| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| C(2nd) → C(Secondary) | 30 × 0.8 | 24 |
| S(3rd) → S(Tertiary) | 20 × 0.6 | 12 |
| **Total** | 86/86 × 100 | **100%** |
#### Example 2: Out-of-order Match = 69.8%
- Student: [I, A, R]
- Major: Primary=R, Secondary=I, Tertiary=S
| Match | Calculation | Points |
|-------|-------------|--------|
| I(1st) → I(Secondary) | 30 × 1.0 | 30 |
| A(2nd) → No match | - | 0 |
| R(3rd) → R(Primary) | 50 × 0.6 | 30 |
| **Total** | 60/86 × 100 | **69.8%** |
#### Example 3: Single Match = 58.1%
- Student: [E, A, S]
- Major: Primary=E, Secondary=I, Tertiary=R
| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| A(2nd) → No match | - | 0 |
| S(3rd) → No match | - | 0 |
| **Total** | 50/86 × 100 | **58.1%** |
#### Example 4: Duplicate codes de-duplicated = 58.1%
- Student: [E, E, S] → scored as [E, S]
- Major: Primary=E, Secondary=I, Tertiary=R
| Match | Calculation | Points |
|-------|-------------|--------|
| E(1st) → E(Primary) | 50 × 1.0 | 50 |
| E(2nd) → *dropped as a duplicate* | - | 0 |
| S(now 2nd) → No match | - | 0 |
| **Total** | 50/86 × 100 | **58.1%** |
Without rule 1 the repeated `E` would have matched the primary slot a second
time for another 40 points, scoring 104.7% → capped at 100%.
---
### Score Interpretation
The percentage is displayed to the student as a number; it is **not** bucketed
into levels or colours in the UI. The only banding in the code is the wording
of the match reason (`_get_match_reasons`), and it has three bands, not five:
| Score | Arabic phrasing shown |
|-------|-----------------------|
| ≥ 50% | `توافق عالٍ مع شخصيتك المهنية` — high match |
| ≥ 30% | `توافق جيد مع شخصيتك المهنية` — good match |
| > 0% | `توافق جزئي مع شخصيتك المهنية` — partial match |
| 0% | no Holland reason is listed at all |
---
### Permutation expansion
`retrieve_with_expansion` exists so a student is never shown an empty result
list. The plain `retrieve` scores every major against the codes **in the order
the student ranked them**, which for an unusual combination can leave almost
nothing above the threshold.
```
1. retrieve(top_k × 3), keep results scoring ≥ 50
(≥ 20 for a single-code student — one code in the tertiary slot
maxes out at 20/86 ≈ 23%, so a 50% threshold would return nothing)
2. If that yields ≥ min_results (default 10), return them. Done.
3. Otherwise re-score every candidate that passed the track / language /
location filters against ALL permutations of the student's codes,
keeping the best score per major.
4. Sort descending, drop anything below 50, return top_k.
```
A single-code student skips step 3 — there is nothing to permute, and step 1's
lower threshold already covers all three positions.
**The best-score map is keyed on `(id, name_ar)`, not on either alone.** The raw
dataset contained a reused id, and two majors can legitimately share a name
across different faculties; either key on its own silently collapses distinct
majors into one entry.
Note that an expanded result reports its **best-permutation** score, not the
score for the order the student actually gave. So a major listed at 74% may be
a weaker fit for the student's stated ranking; expansion only runs when the
honest ranking could not fill the list.
---
### Design Rationale
1. **50/30/20 weighting**: The primary code defines the major's core nature, so
matching it matters most.
2. **1.0/0.8/0.6 multipliers**: A student's first preference is their dominant
trait, so matches against it are weighted higher.
3. **Normalized to 86**: Makes a perfect match land on exactly 100%.
4. **Greedy, single-consumption matching**: Keeps the score monotonic — adding a
code to a student's profile can never lower the score of a major that
already matched.
---
### Tests
`tests/test_retriever.py` covers the rules above (perfect match, asymmetry,
duplicate codes, empty major codes, ordering, and the expansion path):
```bash
python3 -m pytest tests/ -q
```
---
*Kalim Chatbot - Lebanese University Thesis Project*
## Two scores in the expansion path
When fewer than `min_results` majors clear the score threshold,
`retrieve_with_expansion` re-scores every candidate under all six permutations
of the student's three codes and keeps the best. That best-permutation value is
kept as `expansion_score` and is what ranks the results and passes the
threshold — it is the whole point of the expansion.
It is **not** what the student sees. Because `calculate_holland_score` is
asymmetric (the order of the student's codes carries meaning), showing the
best-permutation number answers a different question from the one the interface
asks. Each expanded result therefore also carries `holland_score`, computed
from the order the student actually gave, and that is what the cards, the
results table, "توصية كليم" and the LLM context all display. Expanded result
sets are labelled as such in the results intro.
Worked example — a student with codes `C-E-R` and track `فلسفة وانسانيات`
gets 17 results through the expansion path. The top one, الإدارة السياحية
(triplet `ECS`), ranks on `expansion_score` 86.0 but displays `holland_score`
81.4, the score for `C-E-R` as stated.
|