---
language:
- en
license: apache-2.0
base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
tags:
- reasoning
- grpo
- r1-distill
- arc-agi
- code-generation
- math
- synthetic-reasoning
pipeline_tag: text-generation
library_name: peft
---
# SAM-AI Reasoning v4 (Parallax)
**SAM-AI Reasoning v4** is a 14B parameter reasoning adapter developed by **Parallax (Samrish)**. It is fine-tuned using Group Relative Policy Optimization (GRPO) with rule-based verifiable rewards on top of `deepseek-ai/DeepSeek-R1-Distill-Qwen-14B`.
## Model Overview
- **Base Architecture**: DeepSeek-R1-Distill-Qwen-14B (Qwen2.5 14B transformer backbone)
- **Adapter Type**: LoRA (Rank = 16, Alpha = 32, Dropout = 0.05)
- **Target Modules**: All linear projections (`q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`)
- **Training Method**: GRPO (Group Relative Policy Optimization) with format verification, multi-step math/code reward checks, and reasoning traces (`...`).
- **Trained Parameters**: 137.7 MB adapter safetensors.
## Usage with PEFT & Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model_name = "deepseek-ai/DeepSeek-R1-Distill-Qwen-14B"
adapter_name = "Samrish2009/SAM-AI-Reasoning-v4"
tokenizer = AutoTokenizer.from_pretrained(base_model_name)
base_model = AutoModelForCausalLM.from_pretrained(
base_model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, adapter_name)
prompt = "<|User|>Solve step by step: Prove that for any positive integer n, n^3 + 2n is divisible by 3.<|Assistant|>\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6, top_p=0.95)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Curriculum & Training Objectives
SAM-AI v4 incorporates:
1. **Verifiable Reasoning**: Strict formatting enforcement and multi-step deduction traces.
2. **Abstract Spatial Logic**: Curriculum drawn from ARC inductive reasoning patterns.
3. **Mathematical Derivations**: Step-by-step rigorous proof generation.
4. **Code Execution & Verification**: Synthesizing verifiable Python programs.
## Developed by
- **Team**: Parallax
- **Lead Developer**: Samrish