Instructions to use ImposterOnline/Pyrex-8B-Instruct-Uncensored with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ImposterOnline/Pyrex-8B-Instruct-Uncensored with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ImposterOnline/Pyrex-8B-Instruct-Uncensored") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ImposterOnline/Pyrex-8B-Instruct-Uncensored") model = AutoModelForCausalLM.from_pretrained("ImposterOnline/Pyrex-8B-Instruct-Uncensored", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ImposterOnline/Pyrex-8B-Instruct-Uncensored with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ImposterOnline/Pyrex-8B-Instruct-Uncensored" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ImposterOnline/Pyrex-8B-Instruct-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ImposterOnline/Pyrex-8B-Instruct-Uncensored
- SGLang
How to use ImposterOnline/Pyrex-8B-Instruct-Uncensored with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ImposterOnline/Pyrex-8B-Instruct-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ImposterOnline/Pyrex-8B-Instruct-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ImposterOnline/Pyrex-8B-Instruct-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ImposterOnline/Pyrex-8B-Instruct-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ImposterOnline/Pyrex-8B-Instruct-Uncensored with Docker Model Runner:
docker model run hf.co/ImposterOnline/Pyrex-8B-Instruct-Uncensored
Pyrex 8B Instruct Uncensored — a sharp, uncensored coding model.
Strong at code · honest by design · runs anywhere.
Model · Benchmarks · Usage · GGUF · Training
Introduction
Pyrex 8B Instruct Uncensored is a real, trained model — built with our own data recipe, our own QLoRA fine-tune, and our own evaluation pipeline on Hugging Face GPU compute. It is a coding-first instruction model designed for real developer work: writing code, fixing bugs, explaining technical problems, and driving agent loops.
Key strengths
- Coding-first — trained mostly on high-quality code instructions (OpenCoder real-user data, CodeAlpaca, evol-codealpaca), so it writes clean and correct code.
- Uncensored — answers directly and honestly, without refusals.
- Agentic / tool-use ready — function-calling data was part of training; suited for OpenAI-compatible APIs and agent workflows.
- Portable — full-precision safetensors plus GGUF quants for Ollama, llama.cpp and LM Studio.
Uncensored usage note. This model has no safety-alignment softening built in. It can produce content that is explicit, offensive, or otherwise unsuitable for some audiences. Use it responsibly and at your own discretion — the model and its authors assume no liability for outputs.
Model details
| Property | Value |
|---|---|
| Parameters | 7.6B |
| Architecture | Transformers · RoPE · SwiGLU · RMSNorm · GQA |
| Context window | Up to 32k native (trained at 3072) |
| Base | Coding-specialized instruct model (uncensored variant) |
| Method | QLoRA (4-bit NF4) + LoRA r=48, α=96 |
| Chat format | chatml (`< |
| License | Apache-2.0 |
Benchmarks
Measured with greedy pass@1 on HumanEval — 164 unseen problems, official unit tests executed in a sandbox. No sampling luck, no leaked answers.
| Model | HumanEval pass@1 |
|---|---|
| Base (untuned) | 29.3% (48/164) |
| Pyrex (previous) | 41.5% (68/164) |
| Pyrex 8B Instruct Uncensored | 52.4% (86/164) |
+23.1 points over the base on unseen problems.
Usage
Requirements
transformers >= 4.44for bestapply_chat_templatesupport.- Any recent PyTorch build (CUDA or CPU).
Python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ImposterOnline/Pyrex-8B-Instruct-Uncensored"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype="auto")
messages = [{"role": "user", "content": "Explain async/await in Python."}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(text, return_tensors="pt"), max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))
Chat format is chatml. System prompt is optional; when used it is a system
message before the user turn.
Server / agentic
Works with any OpenAI-compatible server (vLLM, llama.cpp server, Ollama, Text Generation Inference). Function calling was part of training.
vllm serve ImposterOnline/Pyrex-8B-Instruct-Uncensored
GGUF
Quantized builds live in Pyrex-8B-Instruct-Uncensored-GGUF.
| File | Size | Notes |
|---|---|---|
pyrex-q4_k_m.gguf |
4.4 GB | Fast, recommended daily driver |
pyrex-q5_k_m.gguf |
5.1 GB | Higher quality |
pyrex-q8_0.gguf |
8.1 GB | Fast, best quality |
pyrex-f16.gguf |
15.2 GB | Full precision |
# Ollama (Modelfile ships in the GGUF repo)
ollama create pyrex-8b -f Modelfile
ollama run pyrex-8b "Write a Python function that merges overlapping intervals."
# llama.cpp
./llama-cli -m pyrex-q4_k_m.gguf \
-p "<|im_start|>user\nYour question here\n<|im_end|>\n<|im_start|>assistant\n" \
-n 512 -c 8192
Training
| Parameter | Value |
|---|---|
| Data | ~95k cleaned examples (code + tool-use + general) |
| Sequence length | 3072 |
| Loss masking | Assistant-only (completion-only SFT) |
| Optimizer | AdamW · lr 2e-4 · cosine · warmup 3% |
| Epochs | 1 |
| Hardware | Hugging Face A100 GPU job |
Data sources (credited): OpenCoder-LLM/opencoder-sft-stage1
(realuser + largescale-diverse), sahil2801/CodeAlpaca-20k,
theblackcat102/evol-codealpaca-v1, databricks/databricks-dolly-15k,
NousResearch/hermes-function-calling-v1.
The pipeline is config-driven and reproducible — prepare_data.py →
train_qlora.py → eval_humaneval.py → publish.sh — in the companion
repo Arhan-w/pyrut.
License
Apache-2.0.
- Downloads last month
- 374
Model tree for ImposterOnline/Pyrex-8B-Instruct-Uncensored
Datasets used to train ImposterOnline/Pyrex-8B-Instruct-Uncensored
databricks/databricks-dolly-15k
sahil2801/CodeAlpaca-20k
Evaluation results
- pass@1 (greedy) on HumanEvalself-reported52.400