Brahmaputra 1

Brahmaputra 1 starts from dharun2049/kaveri-stgrpo-0.5b and applies a compute-constrained contrastive knowledge-distillation stage using the dense Qwen/Qwen3.8-27B teacher.

Brahmaputra 1 is distilled from the starting checkpoint dharun2049/kaveri-stgrpo-0.5b using Qwen/Qwen3.8-27B as the teacher, with Codeforces anti-forgetting replay.

Teacher

Primary teacher: Qwen/Qwen3.8-27B

The distillation does not perform full-vocabulary KL. For each multiple-choice question, the teacher scores A/B/C/D and the student learns the teacher's four-way answer probability geometry.

Objective

L = 1.0 * L_MC + 1.5 * L_KD + 0.2 * L_margin + 0.35 * L_explanation

where:

  • L_MC is supervised four-way multiple-choice cross entropy.
  • L_KD is temperature-scaled KL divergence over A/B/C/D teacher scores.
  • L_margin forces the correct answer score above distractor scores.
  • L_explanation trains on selected concise teacher explanations and available native explanations.

Coding anti-forgetting

The CKD stage interleaves approximately 15% C++ replay micro-batches from open-r1/codeforces.

Before the fresh CKD LoRA is attached, the starting dharun2049/kaveri-stgrpo-0.5b checkpoint generates deterministic C++ replay trajectories. Coding batches then optimize:

L_code = 0.75 * L_replay_CE + 1.0 * L_anchor_KL

The KL reference is the exact pre-CKD Kaveri policy obtained by temporarily disabling the fresh CKD adapter. The anchor is computed over the reference policy's top-64 token support on completion positions.

Code-preservation probe before CKD: {'ce': 0.848503839224577, 'anchor_kl': -7.259872217280083e-10, 'n': 32}

Code-preservation probe after CKD: {'ce': 0.8353494852781296, 'anchor_kl': 0.0031774165108799934, 'n': 32}

Teacher KD is confidence-gated. If the teacher's top choice disagrees with the gold answer, that example receives zero KD weight, while its gold supervised loss remains active.

Data

MMLU dev/test were not loaded or used for training by the training script.

Training sources:

{ "medmcqa": 1631, "arc_challenge": 1032, "openbookqa": 1671, "arc_easy": 1510, "commonsenseqa": 1727 }

Teacher-scored examples: 7969

Qwen3.8-27B-generated explanations: 59

Training

  • Starting model: dharun2049/kaveri-stgrpo-0.5b
  • LoRA rank: 16
  • LoRA alpha: 32
  • LoRA targets: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Optimizer steps completed: 279
  • Held-out external MCQ accuracy before CKD: 0.5251
  • Held-out external MCQ accuracy after CKD: 0.5276

The repository root contains the merged standalone model. The CKD LoRA adapter is preserved under ckd_adapter/.

Compute-budget note

The training script includes a wall-clock compute-unit governor. Colab does not expose a supported live CU billing API to Python, so the CU figure is an estimate based on a user-configurable assumed CU/hour rate.

Estimated CU at save: 7.10

Downloads last month
112
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dharun2049/Brahmaputra-1

Base model

Qwen/Qwen2-0.5B
Adapter
(1)
this model
Quantizations
5 models