kalim / docs /Holland_Scoring_Algorithm.md
YasserHaidar
Kalim Space deploy (e803a10)
bae15d1
|
Raw History Blame Contribute Delete
8.98 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade

Holland Scoring Algorithm

Kalim Career Guidance Chatbot


Overview

This algorithm calculates the compatibility percentage between a student's Holland (RIASEC) personality codes and a major's RIASEC profile.

It is implemented in KalimRetriever.calculate_holland_score (src/rag/retriever.py). There is no LLM anywhere in this path — the ranking a student sees is deterministic Python over data/processed/majors.json. The language model only phrases answers to free-form chat questions.

The client hand-verified the output of this algorithm; the expected results are recorded in Test_Results.docx in the repository root. After changing anything on this page, re-run a student with codes R-I-A (track الكل, no language or location filter) and diff the top 10 against that document.


Input

Student: up to 3 RIASEC codes ranked by preference (1st, 2nd, 3rd) Major: a holland_codes dict with primary, secondary and tertiary

score = retriever.calculate_holland_score(
    ['E', 'C', 'S'],                                        # student, ranked
    {'primary': 'E', 'secondary': 'C', 'tertiary': 'S'},    # major
)

The function is asymmetric. Swapping the arguments produces a different number — the first list is ranked by the student's preference, the second by the major's requirements, and the two carry different weights. A swapped call does not raise; it silently returns a plausible-looking wrong score, which reads as a scoring regression rather than a caller bug.

score(['R','I','A'], {R, A, S}) = 79.1
score(['R','A','S'], {R, I, A}) = 76.7

Weights

Major Position Weights

Position Weight Reason
Primary 50 points Core personality fit
Secondary 30 points Supporting trait
Tertiary 20 points Complementary trait

Student Preference Multipliers

Rank Multiplier Reason
1st choice 1.0 (100%) Dominant trait
2nd choice 0.8 (80%) Secondary trait
3rd choice 0.6 (60%) Tertiary trait

Formula

For each student code that matches a major code:
    Score += Major_Weight × Student_Multiplier

Maximum Raw Score = 50×1.0 + 30×0.8 + 20×0.6 = 86

Final Percentage = (Raw Score / 86) × 100     rounded to 1 decimal, capped at 100

Matching rules

The three rules below are not cosmetic — each was added to fix a real ranking bug, and removing any of them lets a worse major outrank a better one.

  1. Duplicate student codes are dropped, keeping the earliest (strongest) position. Without this, a student who answered ['E', 'E', 'S'] scores the major's E slot twice and outranks a genuine three-way match.

  2. Each major position is consumed at most once. Once the major's primary has been matched, a later student code cannot match it again. Matching is greedy in student order: the student's 1st choice picks its slot first, then the 2nd, then the 3rd.

  3. An absent major code never matches. A missing code is the empty string in the JSON, and '' == '' would otherwise make two blank fields a match worth 30 or 20 points.

Only the first 3 unique student codes are scored. Empty inputs on either side return 0.0.


Examples

Example 1: Perfect Match = 100%

  • Student: [E, C, S]
  • Major: Primary=E, Secondary=C, Tertiary=S
Match Calculation Points
E(1st) → E(Primary) 50 × 1.0 50
C(2nd) → C(Secondary) 30 × 0.8 24
S(3rd) → S(Tertiary) 20 × 0.6 12
Total 86/86 × 100 100%

Example 2: Out-of-order Match = 69.8%

  • Student: [I, A, R]
  • Major: Primary=R, Secondary=I, Tertiary=S
Match Calculation Points
I(1st) → I(Secondary) 30 × 1.0 30
A(2nd) → No match - 0
R(3rd) → R(Primary) 50 × 0.6 30
Total 60/86 × 100 69.8%

Example 3: Single Match = 58.1%

  • Student: [E, A, S]
  • Major: Primary=E, Secondary=I, Tertiary=R
Match Calculation Points
E(1st) → E(Primary) 50 × 1.0 50
A(2nd) → No match - 0
S(3rd) → No match - 0
Total 50/86 × 100 58.1%

Example 4: Duplicate codes de-duplicated = 58.1%

  • Student: [E, E, S] → scored as [E, S]
  • Major: Primary=E, Secondary=I, Tertiary=R
Match Calculation Points
E(1st) → E(Primary) 50 × 1.0 50
E(2nd) → dropped as a duplicate - 0
S(now 2nd) → No match - 0
Total 50/86 × 100 58.1%

Without rule 1 the repeated E would have matched the primary slot a second time for another 40 points, scoring 104.7% → capped at 100%.


Score Interpretation

The percentage is displayed to the student as a number; it is not bucketed into levels or colours in the UI. The only banding in the code is the wording of the match reason (_get_match_reasons), and it has three bands, not five:

Score Arabic phrasing shown
≥ 50% توافق عالٍ مع شخصيتك المهنية — high match
≥ 30% توافق جيد مع شخصيتك المهنية — good match
> 0% توافق جزئي مع شخصيتك المهنية — partial match
0% no Holland reason is listed at all

Permutation expansion

retrieve_with_expansion exists so a student is never shown an empty result list. The plain retrieve scores every major against the codes in the order the student ranked them, which for an unusual combination can leave almost nothing above the threshold.

1. retrieve(top_k × 3), keep results scoring ≥ 50
   (≥ 20 for a single-code student — one code in the tertiary slot
    maxes out at 20/86 ≈ 23%, so a 50% threshold would return nothing)

2. If that yields ≥ min_results (default 10), return them. Done.

3. Otherwise re-score every candidate that passed the track / language /
   location filters against ALL permutations of the student's codes,
   keeping the best score per major.

4. Sort descending, drop anything below 50, return top_k.

A single-code student skips step 3 — there is nothing to permute, and step 1's lower threshold already covers all three positions.

The best-score map is keyed on (id, name_ar), not on either alone. The raw dataset contained a reused id, and two majors can legitimately share a name across different faculties; either key on its own silently collapses distinct majors into one entry.

Note that an expanded result reports its best-permutation score, not the score for the order the student actually gave. So a major listed at 74% may be a weaker fit for the student's stated ranking; expansion only runs when the honest ranking could not fill the list.


Design Rationale

  1. 50/30/20 weighting: The primary code defines the major's core nature, so matching it matters most.

  2. 1.0/0.8/0.6 multipliers: A student's first preference is their dominant trait, so matches against it are weighted higher.

  3. Normalized to 86: Makes a perfect match land on exactly 100%.

  4. Greedy, single-consumption matching: Keeps the score monotonic — adding a code to a student's profile can never lower the score of a major that already matched.


Tests

tests/test_retriever.py covers the rules above (perfect match, asymmetry, duplicate codes, empty major codes, ordering, and the expansion path):

python3 -m pytest tests/ -q

Kalim Chatbot - Lebanese University Thesis Project

Two scores in the expansion path

When fewer than min_results majors clear the score threshold, retrieve_with_expansion re-scores every candidate under all six permutations of the student's three codes and keeps the best. That best-permutation value is kept as expansion_score and is what ranks the results and passes the threshold — it is the whole point of the expansion.

It is not what the student sees. Because calculate_holland_score is asymmetric (the order of the student's codes carries meaning), showing the best-permutation number answers a different question from the one the interface asks. Each expanded result therefore also carries holland_score, computed from the order the student actually gave, and that is what the cards, the results table, "توصية كليم" and the LLM context all display. Expanded result sets are labelled as such in the results intro.

Worked example — a student with codes C-E-R and track فلسفة وانسانيات gets 17 results through the expansion path. The top one, الإدارة السياحية (triplet ECS), ranks on expansion_score 86.0 but displays holland_score 81.4, the score for C-E-R as stated.