Instructions to use Raghav-Singhal/spp-util-mt-3b-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Raghav-Singhal/spp-util-mt-3b-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Raghav-Singhal/spp-util-mt-3b-base") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Raghav-Singhal/spp-util-mt-3b-base") model = AutoModelForCausalLM.from_pretrained("Raghav-Singhal/spp-util-mt-3b-base", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Raghav-Singhal/spp-util-mt-3b-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Raghav-Singhal/spp-util-mt-3b-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/spp-util-mt-3b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Raghav-Singhal/spp-util-mt-3b-base
- SGLang
How to use Raghav-Singhal/spp-util-mt-3b-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Raghav-Singhal/spp-util-mt-3b-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/spp-util-mt-3b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Raghav-Singhal/spp-util-mt-3b-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/spp-util-mt-3b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Raghav-Singhal/spp-util-mt-3b-base with Docker Model Runner:
docker model run hf.co/Raghav-Singhal/spp-util-mt-3b-base
SPP-MT util โ base (3B)
SPP reflection midtraining from the vanilla model, with the util constitution. This model repeats
model-raising/spp-mt-3b-base with one
change: the first-person reflections come from the util constitution.
Training
- Start point:
model-raising/spp-vanilla-3b-baseat step 225,000 (pre-cooldown), with the vocabulary extended to 49280 tokens. - Midtraining: 72,895 steps (global batch 960, sequence length 2048), no warmup, linear decay over the last 69,013 steps. Loss applies only to reflection tokens on annotated documents; new compact documents get full next-token loss.
- Recipe: identical to
spp-mt-3b-base: same data windows, weights, and schedule. Because util reflections are longer, this model trains on about 3.9B reflection tokens, compared with about 2.5B forspp-mt-3b-base. - Weights: final step 72,895 only.
Util constitution
The reflections in this model come from the "util" constitution. It is a utilitarian revision of
the original SPP constitution. Its reflections weigh harms against benefits (for example,
"โ6 ร 1 per victim against +1 benefit: wrong") and cite new items [7.x] and [8.x]. These items
have no <charter_X.Y> token in the tokenizer. Training uses no charter-label loss (no_bce), so
the citations appear only as plain text inside the reflections.
Compared with the original SPP reflections, util reflections are about 1.56 times longer (mean 83 compared with 53 tokens). Insertion points and source documents are the same.
Tokenizer
Use the bundled tokenizer. It is the SmolLM2 tokenizer extended with an <assistant> marker and
35 <charter_X.Y> tokens (vocabulary 49280). This is a base model; it is not instruction-tuned.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("Raghav-Singhal/spp-util-mt-3b-base")
model = AutoModelForCausalLM.from_pretrained("Raghav-Singhal/spp-util-mt-3b-base", dtype=torch.bfloat16, device_map="auto")
- Downloads last month
- 95