Instructions to use prismdata/Perdix-1.1B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prismdata/Perdix-1.1B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="prismdata/Perdix-1.1B-Instruct", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("prismdata/Perdix-1.1B-Instruct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prismdata/Perdix-1.1B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prismdata/Perdix-1.1B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prismdata/Perdix-1.1B-Instruct
- SGLang
How to use prismdata/Perdix-1.1B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prismdata/Perdix-1.1B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prismdata/Perdix-1.1B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use prismdata/Perdix-1.1B-Instruct with Docker Model Runner:
docker model run hf.co/prismdata/Perdix-1.1B-Instruct
Download README.md from prismdata/Perdix-1.1B-Instruct: direct link, hf CLI and curl.
- Browser
- Download file 6.6 kB
-
https://huggingface.co/prismdata/Perdix-1.1B-Instruct/resolve/main/README.md
- Command line
-
hf download hf://prismdata/Perdix-1.1B-Instruct/README.md
-
curl -L -o README.md https://huggingface.co/prismdata/Perdix-1.1B-Instruct/resolve/main/README.md
license: apache-2.0
language:
- ko
- en
library_name: transformers
pipeline_tag: text-generation
base_model: prismdata/Perdix-1.1B-Base
tags:
- custom_code
- conversational
- instruct
- differential-attention
- polynorm
datasets:
- HuggingFaceTB/smoltalk
- nlpai-lab/kullm-v2
- heegyu/open-korean-instructions
Perdix-1.1B-Instruct
Perdix-1.1B-Base์ ๋ํ ๋ฐ์ดํฐ๋ก ์ง์ ๋ฏธ์ธ์กฐ์ (SFT)์ ํ ๋ชจ๋ธ์ ๋๋ค. ์ด๋ฆ์ ๊ทธ๋ฆฌ์ค ์ ํ์ ํ๋ฅด๋์ค์์ ๋ฐ์์ต๋๋ค. ๋ค์ด๋ฌ๋ก์ค์ ์ด๋ฆฐ ์ ์์๊ณ ํฑ๊ณผ ์ปดํผ์ค๋ฅผ ๋ฐ๋ช ํ๋๋ฐ, ํ์์ ๋จ์ด์ง๋ค ์๊ณ ์๊ฐ ๋๋ ๋ฐ๋์ ๊ทธ ๋ค๋ก๋ ๋ฎ๊ฒ๋ง ๋ ์๋ค๋๋ค๊ณ ํฉ๋๋ค. ์ง๊ธ ์ด ๋ชจ๋ธ ์์ค์ด ๋ฑ ๊ทธ๋ ์ต๋๋ค.
A 1.1B-parameter Korean/English chat model, supervised-finetuned from Perdix-1.1B-Base on about 390K conversations. It answers simple questions and follows short instructions in Korean and English; it is small and often wrong on facts and arithmetic.
๋ฌด์์ ํ ์ ์๊ณ ๋ฌด์์ ๋ชป ํ๋
์ง๋ฌธ์ ๋ตํ๊ณ ์งง์ ์ง์๋ฅผ ๋ฐ๋ฆ ๋๋ค. ํ๊ตญ์ด์ ์์ด ๋ชจ๋ ๋ฉ๋๋ค. ์๋๋ ์ค์ ์ถ๋ ฅ์ ๋๋ค(temperature 0.7, top-k 50, top-p 0.9, repetition penalty 1.1).
๋ํ๋ฏผ๊ตญ์ ์๋๋ ์ด๋์ธ๊ฐ์? ๋ํ๋ฏผ๊ตญ์ ์๋๋ ์์ธ์ ๋๋ค.
๊น์น๊ฐ ๋ฌด์์ธ์ง ๋ ๋ฌธ์ฅ์ผ๋ก ์ค๋ช ํด ์ฃผ์ธ์. ๊น์น๋ ํ๊ตญ์ ์ ํต ๋ฐํจ ์ํ์ผ๋ก, ์ผ๋ฐ์ ์ผ๋ก ์ ์ธ๋ฐฐ์ถ์ ๋ค์ํ ์๋ ์ ๋ฃ๊ณ ๋ฒ๋ฌด๋ฆฐ ํ ๊น์น๋ฅผ ๋ด๊ฐ ๋จน๋ ์์์ ๋๋ค.
What is 7 times 8? To find 7 times 8, we multiply 7 by 8. The answer is 56.
์๋ ! ์ค๋ ๊ธฐ๋ถ์ด ์ด๋? ๊ธฐ๋ถ์ด ์ข์ต๋๋ค! ์ค๋์ ๋ฌด์์ ๋์๋๋ฆด๊น์?
๋ชป ํ๋ ๊ฒ๋ ๋ถ๋ช ํฉ๋๋ค.
- ์ฌ์ค๊ณผ ๊ณ์ฐ์ ์์ฃผ ํ๋ฆฝ๋๋ค. "1๋ถํฐ 10๊น์ง ๋ํ๋ฉด?"์ 1+1=2, 2+1=3โฆ ์์ผ๋ก ์๋ฑํ ํ์ด๋ฅผ ๋ด๋์ต๋๋ค. 1.1B ํฌ๊ธฐ์ ํ๊ณ์ ๋๋ค.
- ๊ธด ์ถ๋ก , ์ฝ๋ ์์ฑ, ํด์ฝ์ ์ ๋ฉ๋๋ค. ๊ทธ๋ฐ ๋ฐ์ดํฐ๋ก ํ์ตํ์ง ์์์ต๋๋ค.
- ๋ํ ๊ธฐ๋ก์ ์งง๊ฒ๋ง ์ ์ง๋ฉ๋๋ค. ๋ฌธ๋งฅ ๊ธธ์ด๊ฐ 2,048ํ ํฐ์ด๊ณ , ๋ฉํฐํด์ ๋ช ํด ์ ๋๋ง ์์ฐ์ค๋ฝ์ต๋๋ค.
- ํ์ต ๋ฐ์ดํฐ๊ฐ ์น๊ณผ ๊ณต๊ฐ ๋ํ ๋ฐ์ดํฐ๋ผ ํธํฅ๋๊ฑฐ๋ ๋ถ์ ์ ํ ๋ด์ฉ์ด ๋์ฌ ์ ์์ต๋๋ค. ์ฌ์ค ํ์ธ์ด ํ์ํ ์ฉ๋์๋ ์ฐ๋ฉด ์ ๋ฉ๋๋ค.
์ฌ์ฉ๋ฒ
๋ชจ๋ธ ์ฝ๋๊ฐ ์ ์ฅ์์ ๋ค์ด ์์ด์ trust_remote_code=True๊ฐ ํ์ํฉ๋๋ค. ๋ํ ํ์์ ChatML์ด๊ณ chat_template์ ๋ค์ด ์์ต๋๋ค.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "prismdata/Perdix-1.1B-Instruct"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval()
messages = [{"role": "user", "content": "๊น์น๊ฐ ๋ฌด์์ธ์ง ๋ ๋ฌธ์ฅ์ผ๋ก ์ค๋ช
ํด ์ฃผ์ธ์."}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
return_dict=True, return_tensors="pt")["input_ids"].to("cuda")
out = model.generate(ids, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.9,
repetition_penalty=1.1)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
์์ ๋ ์ ์ด ์์ต๋๋ค.
- ๋ต๋ณ์
<|im_end|>์์ ๋๋ฉ๋๋ค.eos_token์ด ์ด ํ ํฐ์ผ๋ก ์กํ ์์ดgenerate๊ฐ ์์์ ๋ฉ์ถฅ๋๋ค. - system ํ๋กฌํํธ๋ ํ์์ ๋ฃ์ ์ ์์ง๋ง, ํ์ต ๋ฐ์ดํฐ ๋๋ถ๋ถ์ด system ์์ด ๊ตฌ์ฑ๋ผ ํจ๊ณผ๋ ์ ํ์ ์ ๋๋ค.
- ๋ฌธ๋งฅ ๊ธธ์ด๋ 2,048ํ ํฐ์ ๋๋ค.
- KV ์บ์๋ฅผ ๊ตฌํํ์ง ์์ ์์ฑํ ๋๋ง๋ค ์ ์ฒด ์ํ์ค๋ฅผ ๋ค์ ๊ณ์ฐํฉ๋๋ค. ๊ธธ๊ฒ ์์ฑํ๋ฉด ๋๋ฆฝ๋๋ค.
- ํจ๋ฉ ๋ง์คํฌ๊ฐ ์์ต๋๋ค.
attention_mask๋ ๋ฌด์๋๋ฏ๋ก, ๊ธธ์ด๊ฐ ๋ค๋ฅธ ๋ํ๋ฅผ ํจ๋ฉํด์ ํ ๋ฐฐ์น๋ก ์์ฑํ๋ฉด ๊ฒฐ๊ณผ๊ฐ ํ์ด์ง๋๋ค. ํ ๊ฑด์ฉ ๋ฃ์ผ์ธ์.
๋ชจ๋ธ
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| ํ๋ผ๋ฏธํฐ | 1,107M |
| ๊ตฌ์กฐ | decoder-only, pre-RMSNorm, RoPE, ์ ์ถ๋ ฅ ์๋ฒ ๋ฉ ๊ณต์ |
| ์ธต / ์ฐจ์ / ํค๋ | 20 / 2048 / 16 |
| FFN ์ฐจ์ | 8192 |
| ๋ฌธ๋งฅ ๊ธธ์ด | 2,048 |
| ์ดํ | 49,154 (BPE 49,152 + `< |
| ๊ฐ์ค์น ํ์ | float32 safetensors |
๊ตฌ์กฐ๋ ๋ฒ ์ด์ค์ ๊ฐ์ต๋๋ค. ์ผ๋ฐ ํธ๋์คํฌ๋จธ์ ๋ค๋ฅธ ์ ์ ๋ ๊ฐ์ง์ ๋๋ค.
- Differential Attention. ์ดํ ์ ๋งต์ ๋ ๊ฐ ๋ง๋ค์ด ํ๋์์ ๋ค๋ฅธ ํ๋๋ฅผ ๋บ๋๋ค. ์์ชฝ์ ๊ณตํต์ผ๋ก ๋ผ๋ ์ก์์ ์์ํ๋ ค๋ ์์ด๋์ด์ ๋๋ค.
- PolyNorm. ํ์ฑ ํจ์ ์๋ฆฌ์ x, xยฒ, xยณ์ ๊ฐ๊ฐ ์ ๊ทํํด์ ํ์ต๋๋ ๊ฐ์ค์น๋ก ์์ด ์๋๋ค.
๋ ๋ค ์ ๊ฐ ๊ณ ์ํ ๊ฒ ์๋๊ณ Motif-2.6B ๊ธฐ์ ๋ณด๊ณ ์(arXiv:2508.09148)์ Differential Transformer(Ye et al., 2024)๋ฅผ ์ฝ๊ณ ์ง์ ๊ตฌํํด ๋ณธ ๊ฒ์ ๋๋ค. ์ ์ ์๋ค๊ณผ๋ ๊ด๊ณ์๋ ๊ฐ์ธ ๊ตฌํ์ด๋ผ, ํ๋ฆฐ ๋ถ๋ถ์ด ์๋ค๋ฉด ์ ์ค์์ ๋๋ค.
ํ์ต
๋ฒ ์ด์ค ๋ชจ๋ธ์ ์ฌ์ ํ์ต(30B ํ ํฐ)์ Perdix-1.1B-Base ์นด๋์ ์ ์์ต๋๋ค. ๊ทธ ์์ SFT๋ฅผ ํ์ต๋๋ค.
- ๋ฐ์ดํฐ: ์ฝ 390K ๋ํ, 2 epoch
- ์์ด: smoltalk์ everyday-conversations(3ํ ๋ฐ๋ณต), smol-magpie-ultra, smol-constraints, smol-rewrite, smol-summarize, metamathqa-50k, numina-cot-100k ์ผ๋ถ
- ํ๊ตญ์ด: kullm-v2, open-korean-instructions์ koalpaca, OIG-smallchip2-ko, korquad-chat ์ผ๋ถ
- ๊ฐ ๋ฐ์ดํฐ์ ์ ๋ผ์ด์ ์ค์ ์ถ์ฒ๋ ์ ์ ์ฅ์๋ฅผ ๋ฐ๋ฆ ๋๋ค.
- ํ์: ChatML. ํ ํฌ๋์ด์ ์
<|im_start|>,<|im_end|>๋ฅผ ์ถ๊ฐํ๊ณ ์ ์๋ฒ ๋ฉ์ ๊ธฐ์กด ์๋ฒ ๋ฉ ํ๊ท ์ผ๋ก ์ด๊ธฐํํ์ต๋๋ค. loss๋ ๋ต๋ณ ๊ตฌ๊ฐ์๋ง ๊ฑธ์์ต๋๋ค. - ํ์ต๋ฅ : ์ต๊ณ 5e-5, 100์คํ ์๋ฐ์ ๋ค cosine์ผ๋ก ์ต๊ณ ๊ฐ์ 10%๊น์ง
- ๋ฐฐ์น: ์คํ ๋น ์ฝ 65K ํ ํฐ, ๊ธธ์ด ๊ตฌ๊ฐ๋ณ ํจ๋ฉ, bf16, 6,984์คํ
- ๊ฒฐ๊ณผ: ๊ฒ์ฆ loss 1.37 โ 1.13
- ์ฅ๋น: DGX Spark ํ ๋, ์ฝ 21์๊ฐ
ํ์ต ์ฝ๋๋ github.com/theprismdata/Perdix์ ์์ต๋๋ค.
๋ผ์ด์ ์ค
๋ชจ๋ธ ๊ฐ์ค์น์ ์ฝ๋๋ Apache-2.0. ํ์ต์ ์ด ๋ํ ๋ฐ์ดํฐ๋ ๊ฐ๊ฐ์ ๋ผ์ด์ ์ค๋ฅผ ๋ฐ๋ฅด๋ฉฐ ์ถ์ฒ๋ ์์ ์ ์์ต๋๋ค.