StudyForge AI โ€” GGUF builds

Two Q4_K_M models, served side by side so the effect of post-training can be compared live rather than asserted.

File What it is Size
qwen2.5-1.5b-instruct-base.Q4_K_M.gguf Qwen/Qwen2.5-1.5B-Instruct, untouched 0.99 GB
studyforge-v1.Q4_K_M.gguf The same model after SFT + LoRA (rank 8) 0.99 GB

Both were merged from fp16 and quantised exactly once. The adapter was never folded into already-quantised weights, which would compound the error.

What post-training did โ€” including what it broke

Measured on 131 held-out prompts, greedy decoding, identical settings for both models. * marks a 95% paired-bootstrap interval excluding zero.

Metric Base Post-trained Delta
Format compliance 0.21 0.97 +0.76*
Instruction following 4.45 4.77 +0.32*
Completeness 3.86 3.91 +0.05
Relevance 4.88 4.88 0.00
Clarity 3.98 3.69 โˆ’0.29*
Correctness 3.16 2.44 โˆ’0.73*
Judge mean 4.07 3.94 โˆ’0.13 (n.s.)
General ability (forgetting) 4.50 4.03 โˆ’0.47*

Read that correctness row before using this model. Post-training taught it to produce the requested structure almost perfectly, and made it measurably less accurate. Format compliance is a mechanical contract check with no judge involved; correctness is an LLM judge from a different model family, and the direction was confirmed by blind human scoring of an earlier adapter (โˆ’0.83, 95% CI [โˆ’1.33, โˆ’0.38]).

It also lost general ability outside the study domain โ€” ordinary catastrophic forgetting, reported rather than hidden.

The honest summary: this model is better at looking like a good answer and worse at being one. That is why both files live here, and why the application serves them side by side.

Why rank 8

An earlier rank-16 adapter reached marginally better validation loss and was measurably less truthful (โˆ’0.96 correctness against this model's โˆ’0.73). Halving the rank kept format compliance and recovered a third of the lost accuracy at half the parameters. The extra capacity was not buying structure; it was fitting the training set in ways that damaged accuracy.

Training

  • Base: Qwen2.5-1.5B-Instruct (Apache-2.0)
  • Method: supervised fine-tuning, LoRA r=8 / alpha=16, completion-only loss
  • Data: 935 synthetic instruction examples across 8 study tasks, deduplicated and contract-checked
  • Hardware: one T4, pure fp32, batch 1 ร— grad-accum 16, 3 epochs
  • Trainable: 9.2M parameters (0.60% of the model)

Prompt format

ChatML, as the base model expects. The post-trained model responds to task phrasing such as "explain X using the full structured format" by emitting the seven-section layout it was trained on.

Limitations

  • 1.5B parameters. It is a study aid, not a reference.
  • Verify factual claims. The evaluation above says plainly that this model is less reliable than its own base model on correctness.
  • Trained on synthetic data generated by a larger model, so it inherits that model's blind spots.
  • English only.
Downloads last month
118
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for anmol2526/studyforge-ai-gguf

Adapter
(1490)
this model