Instructions to use schneewolflabs/B2-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use schneewolflabs/B2-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="schneewolflabs/B2-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("schneewolflabs/B2-27B") model = AutoModelForMultimodalLM.from_pretrained("schneewolflabs/B2-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use schneewolflabs/B2-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "schneewolflabs/B2-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schneewolflabs/B2-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/schneewolflabs/B2-27B
- SGLang
How to use schneewolflabs/B2-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "schneewolflabs/B2-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schneewolflabs/B2-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "schneewolflabs/B2-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schneewolflabs/B2-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use schneewolflabs/B2-27B with Docker Model Runner:
docker model run hf.co/schneewolflabs/B2-27B
Schneewolf Labs B2-27B
B1-27B trained to be an agent you can hand real work: delegate the coding, look before it destroys anything, and tell the truth about what it verified. Built to run in the egirl harness with native tool calls and a code agent behind it.
B1-27B
+ Geselle SFT (1,896 per-turn rows, all generated by the 27B itself: ladder work + Vorsicht + thinking-on + chat, 1 epoch, 12k context)
+ ORPO (303 decision-point pairs: Vorsicht + ladder delegate/self-flail + honest/claimed, 2 epochs)
Unlike B2-9B, the SFT uses only rows the 27B produced, since distilling the 9B's trajectories down into a 27B teaches it a smaller model's habits.
Numbers
Same harness for every column: egirl on current main, Q8_0. The held-out set is 8 destructive-request scenarios on fixtures and wording that appear in no training round, ×2 thinking modes; B2 columns are 6 samples per scenario (72 runs), the others 2 (24).
| axis | Qwen3.8-27B (vanilla) | B1-27B | B2-27B |
|---|---|---|---|
| held-out destructive requests: destroyed something on turn 1 | 63% (15/24) | 67% (16/24) | 25% (18/72) |
| held-out: lost irreplaceable data | 5/24 | 6/24 | 7/72 |
| held-out: confirmed-scope removal exact / explicit removal exact | 2/3 / 8/8 | 3/4 / 8/8 | 18/18 / 24/24 |
| egirl 47-case tool bench | n/a | 42/47 | 44/47 |
| censorship (strict, single-sample) | n/a | 28/29 | 27/29 |
| safety asymmetry (refuses actual harm) | n/a | 2/2 | 2/2 |
| hembench | n/a | 75.9% | 72.8% |
| ARC / wiki-clean ppl | n/a | 64.5 / 9.91 | 63.5 / 9.96 |
| stance rate (has opinions) | n/a | 8.3% | 8.3% |
| prose distance vs contemporary fiction (lower = closer) | n/a | 0.57 | 0.67 |
| identity | Qwen (Alibaba) | Schneewolf Labs | Schneewolf Labs |
The egirl gain is delegation: on the bug-fix and performance-investigation cases B1-27B opened with a
reflexive git_status; B2-27B hands them to the code agent. Asked whether it re-ran the tests after the
code agent's fix, it says plainly when it only has the agent's word for it (and with thinking on, goes and
runs them).
What it costs
Three points of hembench (its own Hemlock coding; with a code agent behind it, most coding is delegated) and some prose drift. And it is better, not perfect, at pausing: on the bluntest requests ("clean up ", "wipe everything in ") with thinking off it can still delete first. Keep the harness's own guard on destructive commands; the model is the second line, not the only one.
Notes
- Trained with Merlina, LoRA r32/α64, on a DGX Spark (GB10). SFT lr 1e-4 at 12k context (16k does not fit in bf16), rendered one row per assistant turn with this model's own template; ORPO lr 8e-6, β 0.1, 8k context with fused cross-entropy, pairs cut at their first differing assistant turn so the preference never covers tool output.
- Both adapters merged straight into the weights; the 15
mtp.*tensors and all 348 vision tensors are byte-identical to B1-27B (1,199 tensors verified).--spec-type draft-mtpworks. - Data: Geselle, Vorsicht-DPO.
llama-server -m B2-27B-Q8_0.gguf -ngl 99 -c 32768 --jinja -fa on -np 1 \
--spec-type draft-mtp --spec-draft-n-max 4
- Downloads last month
- 10
Model tree for schneewolflabs/B2-27B
Base model
hemlang/Hemlock-Qwen3.8-27B