Text Generation
PEFT
Safetensors
English
reasoning
grpo
r1-distill
arc-agi
code-generation
math
synthetic-reasoning
conversational
Instructions to use Samrish2009/SAM-AI-Reasoning-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Samrish2009/SAM-AI-Reasoning-v4 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/DeepSeek-R1-Distill-Qwen-14B-bnb-4bit") model = PeftModel.from_pretrained(base_model, "Samrish2009/SAM-AI-Reasoning-v4") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Samrish2009/SAM-AI-Reasoning-v4: direct link, hf CLI and curl.
- Browser
- Download file 2.43 kB
-
https://huggingface.co/Samrish2009/SAM-AI-Reasoning-v4/resolve/main/README.md
- Command line
-
hf download hf://Samrish2009/SAM-AI-Reasoning-v4/README.md
-
curl -L -o README.md https://huggingface.co/Samrish2009/SAM-AI-Reasoning-v4/resolve/main/README.md
2.43 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | |
| tags: | |
| - reasoning | |
| - grpo | |
| - r1-distill | |
| - arc-agi | |
| - code-generation | |
| - math | |
| - synthetic-reasoning | |
| pipeline_tag: text-generation | |
| library_name: peft | |
| # SAM-AI Reasoning v4 (Parallax) | |
| **SAM-AI Reasoning v4** is a 14B parameter reasoning adapter developed by **Parallax (Samrish)**. It is fine-tuned using Group Relative Policy Optimization (GRPO) with rule-based verifiable rewards on top of `deepseek-ai/DeepSeek-R1-Distill-Qwen-14B`. | |
| ## Model Overview | |
| - **Base Architecture**: DeepSeek-R1-Distill-Qwen-14B (Qwen2.5 14B transformer backbone) | |
| - **Adapter Type**: LoRA (Rank = 16, Alpha = 32, Dropout = 0.05) | |
| - **Target Modules**: All linear projections (`q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`) | |
| - **Training Method**: GRPO (Group Relative Policy Optimization) with format verification, multi-step math/code reward checks, and reasoning traces (`<think>...</think>`). | |
| - **Trained Parameters**: 137.7 MB adapter safetensors. | |
| ## Usage with PEFT & Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| from peft import PeftModel | |
| base_model_name = "deepseek-ai/DeepSeek-R1-Distill-Qwen-14B" | |
| adapter_name = "Samrish2009/SAM-AI-Reasoning-v4" | |
| tokenizer = AutoTokenizer.from_pretrained(base_model_name) | |
| base_model = AutoModelForCausalLM.from_pretrained( | |
| base_model_name, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto" | |
| ) | |
| model = PeftModel.from_pretrained(base_model, adapter_name) | |
| prompt = "<|User|>Solve step by step: Prove that for any positive integer n, n^3 + 2n is divisible by 3.<|Assistant|><think>\n" | |
| inputs = tokenizer(prompt, return_tensors="pt").to("cuda") | |
| outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6, top_p=0.95) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| ## Curriculum & Training Objectives | |
| SAM-AI v4 incorporates: | |
| 1. **Verifiable Reasoning**: Strict formatting enforcement and multi-step deduction traces. | |
| 2. **Abstract Spatial Logic**: Curriculum drawn from ARC inductive reasoning patterns. | |
| 3. **Mathematical Derivations**: Step-by-step rigorous proof generation. | |
| 4. **Code Execution & Verification**: Synthesizing verifiable Python programs. | |
| ## Developed by | |
| - **Team**: Parallax | |
| - **Lead Developer**: Samrish | |