sbert-obe-csematch — Continual Learning v3.0

OBE AI System | Course Outcome (CO) Semantic Matching Engine for Computer Science Engineering

This is a domain-specific Sentence-BERT model fine-tuned on Outcome-Based Education (OBE) data for Computer Science & Engineering courses. It maps exam questions and course outcome descriptions into a shared 384-dimensional dense vector space, enabling high-accuracy semantic matching of exam questions to their corresponding Course Outcomes (COs).


🚀 What's New in v3.0 (Continual Learning Release)

  • Warm-started from existing fine-tuned weights (domain knowledge preserved)
  • Trained on the expanded & balanced OBE_Augmented_Dataset (7,009 rows vs 6,387 previously)
  • Top-1 Accuracy improved from 54.6% → 82.4% (+50.9% relative gain)
  • Macro F1-Score improved from 47.9% → 79.6% (+66.1% relative gain)
  • Cosine Separation Margin improved from 0.089 → 0.214 (+139.6% discriminative boost)
  • All embeddings remain 384-dimensional — zero breaking changes to downstream APIs

📋 Model Details

Property Value
Model Type Sentence Transformer (Bi-Encoder)
Base Architecture sentence-transformers/all-MiniLM-L6-v2
Output Dimensions 384d (L2-normalized dense vectors)
Max Sequence Length 256 tokens
Similarity Function Cosine Similarity
Training Approach Continual Learning (warm-start from v2.x weights)
Loss Function MultipleNegativesRankingLoss (in-batch hard negatives)
Language English
License Apache 2.0

Architecture Modules

SentenceTransformer(
  (0): Transformer[BertModel] — 6 layers, 12 heads, 384 hidden
  (1): Pooling — mean pooling strategy
  (2): Normalize — L2 normalization
)

🏋️ Training Configuration

Hyperparameter Value
Dataset OBE_Augmented_Dataset.xlsx → Sheet: Unique_Master_Questions
Dataset Size 7,009 rows (Original Exam + Class Test + Augmented)
Train / Val Split 90% / 10% (group-aware — zero data leakage)
Training Examples 6,387 InputExample pairs (Question ↔ CO Description)
Validation Pairs 622 (positive + shuffled-negative for proper score variance)
Epochs 4
Batch Size 32
Optimizer AdamW
Learning Rate 2e-5
Weight Decay 0.01
Warmup Steps 80 (10% of 800 total steps)
Training Speed 184.7 samples/sec on NVIDIA RTX 4050 (6.4 GB VRAM)
Total Training Time ~2.4 minutes
Final Train Loss 0.6683

Data Split Strategy

Groups are defined by (Course Name, CO) pairs, ensuring all augmented variants of the same exam question land entirely in train OR validation — zero data leakage guaranteed by assertion.


📊 Benchmark Results (N=500, seed=42)

Evaluated on 500 randomly sampled questions (23 courses) from Unique_Master_Questions. Each question is matched only against the valid COs of its own course.

Metric Base Model (all-MiniLM) Fine-Tuned (sbert-obe-csematch) Delta
Top-1 Accuracy 54.60% 82.40% +27.80 (+50.9%) 🚀
Top-3 Accuracy 88.60% 95.40% +6.80 (+7.7%) ✅
Macro Precision 47.76% 75.80% +28.04 (+58.7%) 🚀
Macro Recall 48.71% 86.48% +37.77 (+77.5%) 🚀
Macro F1-Score 47.90% 79.57% +31.67 (+66.1%) 🚀
Weighted Precision 54.83% 82.82% +27.99 (+51.0%) 🚀
Weighted Recall 54.60% 82.40% +27.80 (+50.9%) 🚀
Weighted F1-Score 54.46% 82.38% +27.92 (+51.3%) 🚀
Mean Positive Similarity 0.3456 0.5148 +0.1692 ✅
Mean Negative Similarity 0.2565 0.3012 +0.0447
Cosine Separation Margin 0.0891 0.2135 +0.1245 🚀

Cosine Separation Margin = Mean Positive Similarity − Mean Negative Similarity. A higher margin means the model is more discriminative between correct and distractor COs.


💻 Usage

Basic CO Matching

from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("obe-ai-system/sbert-obe-csematch")

# Exam question
question = "Implement a Binary Search Tree and explain its traversal methods."

# Course outcome candidates
course_outcomes = [
    "CO1: Implement fundamental data structures including trees, graphs, and heaps.",
    "CO2: Analyze time and space complexity of common algorithms.",
    "CO3: Apply sorting and searching algorithms to solve computational problems.",
    "CO4: Design recursive and iterative solutions to complex programming problems.",
]

# Encode and compute similarity
q_emb  = model.encode(question, normalize_embeddings=True)
co_emb = model.encode(course_outcomes, normalize_embeddings=True)
scores = util.cos_sim(q_emb, co_emb)[0]

# Rank by similarity
ranked = sorted(zip(course_outcomes, scores.tolist()), key=lambda x: x[1], reverse=True)
for co, score in ranked:
    print(f"{score:.4f}  {co[:60]}")

Batch Evaluation

import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("obe-ai-system/sbert-obe-csematch")

questions = ["...", "..."]  # list of exam questions
co_descs  = ["...", "..."]  # corresponding CO descriptions

q_embs  = model.encode(questions, normalize_embeddings=True, batch_size=64)
co_embs = model.encode(co_descs,  normalize_embeddings=True, batch_size=64)

# Pair-wise cosine similarity
similarities = np.einsum("ij,ij->i", q_embs, co_embs)
print(f"Mean cosine similarity: {similarities.mean():.4f}")

🗂️ Repository Files

File Description
model.safetensors Fine-tuned BERT weights (6-layer, 384-hidden)
config.json Transformer architecture config
config_sentence_transformers.json SentenceTransformer pipeline config
tokenizer.json WordPiece tokenizer
tokenizer_config.json Tokenizer settings
modules.json Pipeline module definitions (Transformer → Pooling → Normalize)
sentence_bert_config.json SBERT max sequence length config
1_Pooling/config.json Mean pooling configuration
2_Normalize/config.json L2 normalization configuration

🔌 Integration — OBE AI System FastAPI Service

This model powers the /api/v1/co-match endpoint of the OBE AI System:

# POST /api/v1/co-match
{
    "question": "Explain the concept of virtual memory and page replacement.",
    "course_code": "CSE 311",
    "top_k": 3
}

# Response
{
    "top_matches": [
        {"co_id": "CO4", "similarity": 0.8721, "description": "Analyze working procedure and structure of main and virtual memory."},
        {"co_id": "CO3", "similarity": 0.7423, "description": "..."},
        {"co_id": "CO2", "similarity": 0.6912, "description": "..."}
    ],
    "embedding_dim": 384,
    "model_version": "sbert-obe-csematch-v3.0"
}

📚 Training Data

The model was trained on the OBE Augmented Dataset covering 15+ CSE courses:

Course Code Coverage
Discrete Mathematics CSE 113 ✅
Data Structures CSE 211 ✅
Object-Oriented Programming CSE 213 ✅
Operating System CSE 311 ✅
Database Management System CSE 313 ✅
Computer Architecture and Design CSE 315 ✅
Software Engineering CSE 317 ✅
Theory of Computation CSE 319 ✅
Data Communication CSE 321 ✅
Web Technologies CSE 323 ✅
Computer Networks CSE 327 ✅
Machine Learning CSE 411 ✅
Compiler Design CSE 413 ✅
Digital Image Processing CSE 443 ✅
+ More ... ✅

Data types: Original Exam (2,049) + Class Test (211) + Augmented (4,749) = 7,009 total rows


⚙️ Technical Notes

  • Backward compatible: All embeddings are 384-dimensional — no changes to downstream vector stores, similarity pipelines, or APIs required
  • Warm-start training: v3.0 built upon v2.x fine-tuned weights using continual learning (domain knowledge preserved, catastrophic forgetting minimized)
  • Group-aware data split: (Course Name, CO) grouping prevents augmented variants of the same question from leaking across train/val splits
  • In-batch hard negatives: MultipleNegativesRankingLoss treats other CO descriptions in the same batch as hard negatives, improving discriminative ability

📈 Training Logs

Epoch Step Train Loss Notes
0.5 100 — Checkpoint saved
1.0 200 — Checkpoint saved
1.5 300 — Checkpoint saved
2.0 400 — Checkpoint saved
2.5 500 0.7755 Logged
3.0 600 — Checkpoint saved
3.5 700 — Checkpoint saved
4.0 800 0.6683 Final model saved

🧰 Framework Versions

Library Version
Python 3.13.14
Sentence Transformers 6.0.1
Transformers 5.16.1
PyTorch 2.6.0+cu124
Accelerate 1.15.0
Datasets 5.0.1
Tokenizers 0.23.2

📖 Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
63
Safetensors
Model size
22.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for obe-ai-system/sbert-obe-csematch

Evaluation results

  • Top-1 Accuracy (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questions
    self-reported
    0.824
  • Top-3 Accuracy (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questions
    self-reported
    0.954
  • Macro F1-Score (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questions
    self-reported
    0.796