Text Generation
PEFT
Safetensors
English
reasoning
grpo
r1-distill
arc-agi
code-generation
math
synthetic-reasoning
conversational
Instructions to use Samrish2009/SAM-AI-Reasoning-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Samrish2009/SAM-AI-Reasoning-v4 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/DeepSeek-R1-Distill-Qwen-14B-bnb-4bit") model = PeftModel.from_pretrained(base_model, "Samrish2009/SAM-AI-Reasoning-v4") - Notebooks
- Google Colab
- Kaggle
File size: 2,428 Bytes
4abd643 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 | ---
language:
- en
license: apache-2.0
base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
tags:
- reasoning
- grpo
- r1-distill
- arc-agi
- code-generation
- math
- synthetic-reasoning
pipeline_tag: text-generation
library_name: peft
---
# SAM-AI Reasoning v4 (Parallax)
**SAM-AI Reasoning v4** is a 14B parameter reasoning adapter developed by **Parallax (Samrish)**. It is fine-tuned using Group Relative Policy Optimization (GRPO) with rule-based verifiable rewards on top of `deepseek-ai/DeepSeek-R1-Distill-Qwen-14B`.
## Model Overview
- **Base Architecture**: DeepSeek-R1-Distill-Qwen-14B (Qwen2.5 14B transformer backbone)
- **Adapter Type**: LoRA (Rank = 16, Alpha = 32, Dropout = 0.05)
- **Target Modules**: All linear projections (`q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`)
- **Training Method**: GRPO (Group Relative Policy Optimization) with format verification, multi-step math/code reward checks, and reasoning traces (`<think>...</think>`).
- **Trained Parameters**: 137.7 MB adapter safetensors.
## Usage with PEFT & Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model_name = "deepseek-ai/DeepSeek-R1-Distill-Qwen-14B"
adapter_name = "Samrish2009/SAM-AI-Reasoning-v4"
tokenizer = AutoTokenizer.from_pretrained(base_model_name)
base_model = AutoModelForCausalLM.from_pretrained(
base_model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, adapter_name)
prompt = "<|User|>Solve step by step: Prove that for any positive integer n, n^3 + 2n is divisible by 3.<|Assistant|><think>\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6, top_p=0.95)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Curriculum & Training Objectives
SAM-AI v4 incorporates:
1. **Verifiable Reasoning**: Strict formatting enforcement and multi-step deduction traces.
2. **Abstract Spatial Logic**: Curriculum drawn from ARC inductive reasoning patterns.
3. **Mathematical Derivations**: Step-by-step rigorous proof generation.
4. **Code Execution & Verification**: Synthesizing verifiable Python programs.
## Developed by
- **Team**: Parallax
- **Lead Developer**: Samrish
|