Instructions to use obe-ai-system/sbert-obe-csematch with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use obe-ai-system/sbert-obe-csematch with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("obe-ai-system/sbert-obe-csematch") sentences = [ "Using the Binomial Theorem: Expand (x + y)^3?", "Apply different counting techniques for various types of events.", "Demonstrate an understanding of propositional and predicate logic.", "Explain and evaluate fundamental web programming concepts, architectures, and protocols." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
- sbert-obe-csematch — Continual Learning v3.0
sbert-obe-csematch — Continual Learning v3.0
OBE AI System | Course Outcome (CO) Semantic Matching Engine for Computer Science Engineering
This is a domain-specific Sentence-BERT model fine-tuned on Outcome-Based Education (OBE) data for Computer Science & Engineering courses. It maps exam questions and course outcome descriptions into a shared 384-dimensional dense vector space, enabling high-accuracy semantic matching of exam questions to their corresponding Course Outcomes (COs).
🚀 What's New in v3.0 (Continual Learning Release)
- Warm-started from existing fine-tuned weights (domain knowledge preserved)
- Trained on the expanded & balanced
OBE_Augmented_Dataset(7,009 rows vs 6,387 previously) - Top-1 Accuracy improved from 54.6% → 82.4% (+50.9% relative gain)
- Macro F1-Score improved from 47.9% → 79.6% (+66.1% relative gain)
- Cosine Separation Margin improved from 0.089 → 0.214 (+139.6% discriminative boost)
- All embeddings remain 384-dimensional — zero breaking changes to downstream APIs
📋 Model Details
| Property | Value |
|---|---|
| Model Type | Sentence Transformer (Bi-Encoder) |
| Base Architecture | sentence-transformers/all-MiniLM-L6-v2 |
| Output Dimensions | 384d (L2-normalized dense vectors) |
| Max Sequence Length | 256 tokens |
| Similarity Function | Cosine Similarity |
| Training Approach | Continual Learning (warm-start from v2.x weights) |
| Loss Function | MultipleNegativesRankingLoss (in-batch hard negatives) |
| Language | English |
| License | Apache 2.0 |
Architecture Modules
SentenceTransformer(
(0): Transformer[BertModel] — 6 layers, 12 heads, 384 hidden
(1): Pooling — mean pooling strategy
(2): Normalize — L2 normalization
)
🏋️ Training Configuration
| Hyperparameter | Value |
|---|---|
| Dataset | OBE_Augmented_Dataset.xlsx → Sheet: Unique_Master_Questions |
| Dataset Size | 7,009 rows (Original Exam + Class Test + Augmented) |
| Train / Val Split | 90% / 10% (group-aware — zero data leakage) |
| Training Examples | 6,387 InputExample pairs (Question ↔ CO Description) |
| Validation Pairs | 622 (positive + shuffled-negative for proper score variance) |
| Epochs | 4 |
| Batch Size | 32 |
| Optimizer | AdamW |
| Learning Rate | 2e-5 |
| Weight Decay | 0.01 |
| Warmup Steps | 80 (10% of 800 total steps) |
| Training Speed | 184.7 samples/sec on NVIDIA RTX 4050 (6.4 GB VRAM) |
| Total Training Time | ~2.4 minutes |
| Final Train Loss | 0.6683 |
Data Split Strategy
Groups are defined by (Course Name, CO) pairs, ensuring all augmented variants of the same exam question land entirely in train OR validation — zero data leakage guaranteed by assertion.
📊 Benchmark Results (N=500, seed=42)
Evaluated on 500 randomly sampled questions (23 courses) from Unique_Master_Questions. Each question is matched only against the valid COs of its own course.
| Metric | Base Model (all-MiniLM) | Fine-Tuned (sbert-obe-csematch) | Delta |
|---|---|---|---|
| Top-1 Accuracy | 54.60% | 82.40% | +27.80 (+50.9%) 🚀 |
| Top-3 Accuracy | 88.60% | 95.40% | +6.80 (+7.7%) ✅ |
| Macro Precision | 47.76% | 75.80% | +28.04 (+58.7%) 🚀 |
| Macro Recall | 48.71% | 86.48% | +37.77 (+77.5%) 🚀 |
| Macro F1-Score | 47.90% | 79.57% | +31.67 (+66.1%) 🚀 |
| Weighted Precision | 54.83% | 82.82% | +27.99 (+51.0%) 🚀 |
| Weighted Recall | 54.60% | 82.40% | +27.80 (+50.9%) 🚀 |
| Weighted F1-Score | 54.46% | 82.38% | +27.92 (+51.3%) 🚀 |
| Mean Positive Similarity | 0.3456 | 0.5148 | +0.1692 ✅ |
| Mean Negative Similarity | 0.2565 | 0.3012 | +0.0447 |
| Cosine Separation Margin | 0.0891 | 0.2135 | +0.1245 🚀 |
Cosine Separation Margin = Mean Positive Similarity − Mean Negative Similarity. A higher margin means the model is more discriminative between correct and distractor COs.
💻 Usage
Basic CO Matching
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("obe-ai-system/sbert-obe-csematch")
# Exam question
question = "Implement a Binary Search Tree and explain its traversal methods."
# Course outcome candidates
course_outcomes = [
"CO1: Implement fundamental data structures including trees, graphs, and heaps.",
"CO2: Analyze time and space complexity of common algorithms.",
"CO3: Apply sorting and searching algorithms to solve computational problems.",
"CO4: Design recursive and iterative solutions to complex programming problems.",
]
# Encode and compute similarity
q_emb = model.encode(question, normalize_embeddings=True)
co_emb = model.encode(course_outcomes, normalize_embeddings=True)
scores = util.cos_sim(q_emb, co_emb)[0]
# Rank by similarity
ranked = sorted(zip(course_outcomes, scores.tolist()), key=lambda x: x[1], reverse=True)
for co, score in ranked:
print(f"{score:.4f} {co[:60]}")
Batch Evaluation
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("obe-ai-system/sbert-obe-csematch")
questions = ["...", "..."] # list of exam questions
co_descs = ["...", "..."] # corresponding CO descriptions
q_embs = model.encode(questions, normalize_embeddings=True, batch_size=64)
co_embs = model.encode(co_descs, normalize_embeddings=True, batch_size=64)
# Pair-wise cosine similarity
similarities = np.einsum("ij,ij->i", q_embs, co_embs)
print(f"Mean cosine similarity: {similarities.mean():.4f}")
🗂️ Repository Files
| File | Description |
|---|---|
model.safetensors |
Fine-tuned BERT weights (6-layer, 384-hidden) |
config.json |
Transformer architecture config |
config_sentence_transformers.json |
SentenceTransformer pipeline config |
tokenizer.json |
WordPiece tokenizer |
tokenizer_config.json |
Tokenizer settings |
modules.json |
Pipeline module definitions (Transformer → Pooling → Normalize) |
sentence_bert_config.json |
SBERT max sequence length config |
1_Pooling/config.json |
Mean pooling configuration |
2_Normalize/config.json |
L2 normalization configuration |
🔌 Integration — OBE AI System FastAPI Service
This model powers the /api/v1/co-match endpoint of the OBE AI System:
# POST /api/v1/co-match
{
"question": "Explain the concept of virtual memory and page replacement.",
"course_code": "CSE 311",
"top_k": 3
}
# Response
{
"top_matches": [
{"co_id": "CO4", "similarity": 0.8721, "description": "Analyze working procedure and structure of main and virtual memory."},
{"co_id": "CO3", "similarity": 0.7423, "description": "..."},
{"co_id": "CO2", "similarity": 0.6912, "description": "..."}
],
"embedding_dim": 384,
"model_version": "sbert-obe-csematch-v3.0"
}
📚 Training Data
The model was trained on the OBE Augmented Dataset covering 15+ CSE courses:
| Course | Code | Coverage |
|---|---|---|
| Discrete Mathematics | CSE 113 | ✅ |
| Data Structures | CSE 211 | ✅ |
| Object-Oriented Programming | CSE 213 | ✅ |
| Operating System | CSE 311 | ✅ |
| Database Management System | CSE 313 | ✅ |
| Computer Architecture and Design | CSE 315 | ✅ |
| Software Engineering | CSE 317 | ✅ |
| Theory of Computation | CSE 319 | ✅ |
| Data Communication | CSE 321 | ✅ |
| Web Technologies | CSE 323 | ✅ |
| Computer Networks | CSE 327 | ✅ |
| Machine Learning | CSE 411 | ✅ |
| Compiler Design | CSE 413 | ✅ |
| Digital Image Processing | CSE 443 | ✅ |
| + More | ... | ✅ |
Data types: Original Exam (2,049) + Class Test (211) + Augmented (4,749) = 7,009 total rows
⚙️ Technical Notes
- Backward compatible: All embeddings are 384-dimensional — no changes to downstream vector stores, similarity pipelines, or APIs required
- Warm-start training: v3.0 built upon v2.x fine-tuned weights using continual learning (domain knowledge preserved, catastrophic forgetting minimized)
- Group-aware data split:
(Course Name, CO)grouping prevents augmented variants of the same question from leaking across train/val splits - In-batch hard negatives:
MultipleNegativesRankingLosstreats other CO descriptions in the same batch as hard negatives, improving discriminative ability
📈 Training Logs
| Epoch | Step | Train Loss | Notes |
|---|---|---|---|
| 0.5 | 100 | — | Checkpoint saved |
| 1.0 | 200 | — | Checkpoint saved |
| 1.5 | 300 | — | Checkpoint saved |
| 2.0 | 400 | — | Checkpoint saved |
| 2.5 | 500 | 0.7755 | Logged |
| 3.0 | 600 | — | Checkpoint saved |
| 3.5 | 700 | — | Checkpoint saved |
| 4.0 | 800 | 0.6683 | Final model saved |
🧰 Framework Versions
| Library | Version |
|---|---|
| Python | 3.13.14 |
| Sentence Transformers | 6.0.1 |
| Transformers | 5.16.1 |
| PyTorch | 2.6.0+cu124 |
| Accelerate | 1.15.0 |
| Datasets | 5.0.1 |
| Tokenizers | 0.23.2 |
📖 Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}
- Downloads last month
- 63
Papers for obe-ai-system/sbert-obe-csematch
Representation Learning with Contrastive Predictive Coding
Evaluation results
- Top-1 Accuracy (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questionsself-reported0.824
- Top-3 Accuracy (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questionsself-reported0.954
- Macro F1-Score (Fine-Tuned) on OBE_Augmented_Dataset — Unique_Master_Questionsself-reported0.796