dpo-v1 — Preference-Tuned ML/LLM Assistant (DPO on sft-v2)

⚠️ Educational project. This model is a personal learning experiment, released for learning and educational purposes only. See License & Intended Use.

dpo-v1 is vinmlops/sft-v2 after Direct Preference Optimization (DPO). SFT taught the model how to answer; DPO teaches it which answers are better by learning from pairs of a preferred and a less-preferred response to the same question. It is a chat-style assistant for questions about large language models, neural networks and transformers, and it also handles general questions.

This is stage 3 of 3 in a small end-to-end pipeline:

Qwen3-1.7B-Base  ──CPT──▶  cpt-v2  ──SFT──▶  sft-v2  ──DPO──▶  dpo-v1
                                                           (this repo)

What it is

  • Base: vinmlops/sft-v2 (Qwen3-1.7B-Base → continued pretraining → supervised fine-tuning)
  • Method: DPO with LoRA, merged into a standalone model (bfloat16). A light supervised term on the preferred answers keeps the model's end-of-turn behavior stable.
  • Format: a full standalone model — load it with AutoModelForCausalLM.from_pretrained; no adapter or PEFT needed.
  • Preference signal: pairs of the model's own sampled answers, compared by an automated judge with a correctness-first rubric and checked for consistency in both answer orders. Most preferences were decided on factual correctness.

Evaluation (vs. sft-v2)

Check sft-v2 dpo-v1
Held-out preference accuracy 57.4% 62.5%
Stops cleanly with sampling (temp 0.7) 100% 99%
Head-to-head on an internal behavior set (both orders must agree) 1 win 13 wins (24 ties)
Average answer length 37 tokens 46 tokens (+26%)
  • DPO answers were preferred more often, with no regressions on identity, safety or general questions in the behavior set.
  • Answers became somewhat longer. Most of that growth appears as more complete answers, but it is a known trade-off to watch.
  • The behavior set is small; treat these as indicative results, not precise rates.

Recommended system prompt

For machine-learning and LLM questions:

You are a helpful assistant with expertise in machine learning and large language models. Answer the user's question accurately and directly. Explain concepts clearly and acknowledge uncertainty when you are unsure.

For general questions, use no system prompt (this matches how the model was trained).

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "vinmlops/dpo-v1"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, dtype=torch.bfloat16)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()

SYSTEM = ("You are a helpful assistant with expertise in machine learning and large language "
          "models. Answer the user's question accurately and directly. Explain concepts clearly "
          "and acknowledge uncertainty when you are unsure.")
im_end = tok.convert_tokens_to_ids("<|im_end|>")

def ask(question, ml_question=True):
    msgs = ([{"role": "system", "content": SYSTEM}] if ml_question else []) + \
           [{"role": "user", "content": question}]
    text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
    inputs = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    with torch.no_grad():
        out = model.generate(**inputs, max_new_tokens=400, do_sample=False,
                             eos_token_id=im_end, pad_token_id=im_end)
    return tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()

print(ask("What is attention in a transformer?"))
print(ask("What causes the seasons?", ml_question=False))

(This is a gated repository — you must be granted access and be authenticated with hf auth login to download it.)

Intended use

  • Asking beginner-to-intermediate questions about LLMs, neural networks and transformers.
  • Educational demos of the full base → CPT → SFT → DPO pipeline, and of what preference tuning changes compared with SFT alone.

Limitations

  • Small-scale, experimental. A 1.7B model with a modest preference set — helpful for learning, not an authoritative expert.
  • May produce inaccurate, incomplete or outdated information. Verify anything important; do not use it for production or critical decisions.
  • Automated preferences. Preference labels come from an automated judge, which can make mistakes; the model can inherit them.
  • Slightly longer answers than sft-v2, and limited training on explicit format or length instructions.
  • Strongest on ML/NN/transformer topics; general-domain performance is more limited.
  • Inherits biases and limitations of the base model and training data.

The training data was built from publicly available material for personal, educational training use only and is not included or redistributed here. This repository contains only the resulting model weights.

License & Intended Use

This model is released for learning and educational purposes only.

  • Licensed under CC-BY-NC-4.0 (Creative Commons Attribution–NonCommercial 4.0).
  • You may use, study and share it for non-commercial, educational and research purposes, with attribution.
  • Commercial use is not permitted.
  • Provided "as is", without warranty of any kind; it may produce inaccurate output. Use at your own risk.
  • Lineage: Qwen/Qwen3-1.7B-Base (Apache-2.0) → vinmlops/cpt-v2 → vinmlops/sft-v2 → this DPO model. Please also respect the base model's license.

Citation / attribution

vinmlops/dpo-v1 — preference-tuned ML/LLM assistant (DPO on sft-v2, merged),
an educational LLM/NN/transformer learning project. Non-commercial (CC-BY-NC-4.0).
Lineage: Qwen/Qwen3-1.7B-Base (Apache-2.0) -> cpt-v2 -> sft-v2 -> dpo-v1.
Downloads last month
22
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vinmlops/dpo-v1

Base model

vinmlops/cpt-v2
Finetuned
vinmlops/sft-v2
Finetuned
(1)
this model

Space using vinmlops/dpo-v1 1