Text Generation
Safetensors
GGUF
Russian
English
mistral3
reasoning
deepseek-r1
ru-deepthink-11k
mistral
conversational
Instructions to use fwizzer1/Fwizzer-R1-3B-RU with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use fwizzer1/Fwizzer-R1-3B-RU with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf fwizzer1/Fwizzer-R1-3B-RU # Run inference directly in the terminal: llama cli -hf fwizzer1/Fwizzer-R1-3B-RU
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf fwizzer1/Fwizzer-R1-3B-RU # Run inference directly in the terminal: llama cli -hf fwizzer1/Fwizzer-R1-3B-RU
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf fwizzer1/Fwizzer-R1-3B-RU # Run inference directly in the terminal: ./llama-cli -hf fwizzer1/Fwizzer-R1-3B-RU
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf fwizzer1/Fwizzer-R1-3B-RU # Run inference directly in the terminal: ./build/bin/llama-cli -hf fwizzer1/Fwizzer-R1-3B-RU
Use Docker
docker model run hf.co/fwizzer1/Fwizzer-R1-3B-RU
- LM Studio
- Jan
- vLLM
How to use fwizzer1/Fwizzer-R1-3B-RU with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "fwizzer1/Fwizzer-R1-3B-RU" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fwizzer1/Fwizzer-R1-3B-RU", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/fwizzer1/Fwizzer-R1-3B-RU
- Ollama
How to use fwizzer1/Fwizzer-R1-3B-RU with Ollama:
ollama run hf.co/fwizzer1/Fwizzer-R1-3B-RU
- Unsloth Desktop
- Docker Model Runner
How to use fwizzer1/Fwizzer-R1-3B-RU with Docker Model Runner:
docker model run hf.co/fwizzer1/Fwizzer-R1-3B-RU
- Lemonade
How to use fwizzer1/Fwizzer-R1-3B-RU with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull fwizzer1/Fwizzer-R1-3B-RU
Run and chat with the model
lemonade run user.Fwizzer-R1-3B-RU-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
File size: 2,210 Bytes
6360f81 303e507 6360f81 303e507 e9b809b b34cc3e e9b809b e5e58be 303e507 6360f81 b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de b34cc3e 04c73de 6360f81 b34cc3e 6360f81 b34cc3e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 | ---
language:
- ru
- en
license: apache-2.0
tags:
- reasoning
- deepseek-r1
- ru-deepthink-11k
- mistral
- text-generation
- gguf
pipeline_tag: text-generation
---
# 🧠 Fwizzer-R1-3B-RU
**Fwizzer-R1-3B-RU** — это мыслящая русскоязычная языковая модель, обученная по архитектуре пошаговых рассуждений (**DeepSeek-R1 CoT**) на отборном датасете [`fwizzer1/ru-deepthink-11k`](https://huggingface.co/datasets/fwizzer1/ru-deepthink-11k).
---
## ⚡️ Доступные версии GGUF
| Файл | Описание | Рекомендуемое железо |
| :--- | :--- | :--- |
| **`Fwizzer-R1-3B-Speed.gguf`** | Быстрая версия (Q4_K_M, 2.15 GB) | RTX 3050 / Ноутбуки / 16GB RAM |
| **`Fwizzer-R1-3B-Max.gguf`** | Максимальная точность (Q8_0, 3.40 GB) | ПК с 8+ GB VRAM |
---
## 🚀 Использование в LM Studio
1. Откройте **LM Studio**.
2. В строке поиска введите: `Fwizzer-R1-3B-RU`.
3. Нажмите **Download** на `Fwizzer-R1-3B-Speed.gguf` или `Fwizzer-R1-3B-Max.gguf`.
4. Модель готова к работе со шторкой размышлений!
### Вшитый системный промпт:
```text
Ты думающая нейросеть а зовут тебя Fwizzer-R1-3B-RU. Весь ход мыслей и шаги пиши внутри тегов <think>(напиши сначала) и </think>(напиши по окончанию рассуждений), а итоговый ответ — обязательно после них.
```
---
## 💻 Использование через Python (llama-cpp-python)
```python
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="fwizzer1/Fwizzer-R1-3B-RU",
filename="Fwizzer-R1-3B-Speed.gguf",
n_ctx=4096,
n_gpu_layers=-1,
flash_attn=True
)
response = llm.create_chat_completion(
messages=[
{"role": "user", "content": "Привет! Расскажи о себе и реши задачу на логику."}
]
)
print(response["choices"][0]["message"]["content"])
```
|