Instructions to use prismdata/Perdix-1.1B-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prismdata/Perdix-1.1B-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="prismdata/Perdix-1.1B-Base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("prismdata/Perdix-1.1B-Base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prismdata/Perdix-1.1B-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prismdata/Perdix-1.1B-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/prismdata/Perdix-1.1B-Base
- SGLang
How to use prismdata/Perdix-1.1B-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prismdata/Perdix-1.1B-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prismdata/Perdix-1.1B-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prismdata/Perdix-1.1B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use prismdata/Perdix-1.1B-Base with Docker Model Runner:
docker model run hf.co/prismdata/Perdix-1.1B-Base
Download README.md from prismdata/Perdix-1.1B-Base: direct link, hf CLI and curl.
- Browser
- Download file 5.33 kB
-
https://huggingface.co/prismdata/Perdix-1.1B-Base/resolve/main/README.md
- Command line
-
hf download hf://prismdata/Perdix-1.1B-Base/README.md
-
curl -L -o README.md https://huggingface.co/prismdata/Perdix-1.1B-Base/resolve/main/README.md
license: apache-2.0
language:
- ko
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- custom_code
- base-model
- differential-attention
- polynorm
datasets:
- prismdata/Perdix-Pretrain-Data
- mlfoundations/dclm-baseline-1.0
- HuggingFaceFW/fineweb-2
- HuggingFaceTB/finemath
Perdix-1.1B-Base
์์ ์์ค์ LLM์ผ๋ก ์ด๋ฆ์ ๊ทธ๋ฆฌ์ค ์ ํ์ ํ๋ฅด๋์ค์์ ๋ฐ์์ต๋๋ค. ๋ค์ด๋ฌ๋ก์ค์ ์ด๋ฆฐ ์ ์์๊ณ ํฑ๊ณผ ์ปดํผ์ค๋ฅผ ๋ฐ๋ช ํ๋๋ฐ, ํ์์ ๋จ์ด์ง๋ค ์๊ณ ์๊ฐ ๋๋ ๋ฐ๋์ ๊ทธ ๋ค๋ก๋ ๋ฎ๊ฒ๋ง ๋ ์๋ค๋๋ค๊ณ ํฉ๋๋ค. ์ง๊ธ ์ด ๋ชจ๋ธ ์์ค์ด ๋ฑ ๊ทธ๋ ์ต๋๋ค.
A 1.1B-parameter Korean/English base language model pretrained from scratch on a single machine. It only continues text; it has not been trained to chat or follow instructions.
๋ฌด์์ ํ ์ ์๊ณ ๋ฌด์์ ๋ชป ํ๋
๋ฒ ์ด์ค ๋ชจ๋ธ์ ๋๋ค. ๋ฌธ์ฅ ์๋ถ๋ถ์ ์ฃผ๋ฉด ๋ค๋ฅผ ์๋ ๊ฒ๋ง ํฉ๋๋ค. ์ง๋ฌธ์ ๋ตํ๊ฑฐ๋ ์ง์๋ฅผ ๋ฐ๋ฅด๋๋ก ํ์ต์ํจ ์ ์ด ์์ด์, ๋ํ์ฉ์ผ๋ก ์ฐ๋ฉด ์๋ฑํ ๊ธ์ด ๋์ต๋๋ค.
์ด์ด ์ด ๊ธ์ ๋ฌธ์ฅ์ผ๋ก๋ ๊ทธ๋ด๋ฏํ์ง๋ง ๋ด์ฉ์ ์์ฃผ ํ๋ฆฝ๋๋ค. ์๋๋ ์ค์ ์ถ๋ ฅ์ ๋๋ค(temperature 0.8, top-k 50). ๊ตต์ ๋ถ๋ถ์ด ์ ๋ ฅ์ ๋๋ค.
๋ํ๋ฏผ๊ตญ์ ์๋ ์์ธ์ ์ธ๊ณ 4๋ ๋ฌธ๋ช ๋ฐ์์ง์ด์ ์ต๋ ๋์์ด์ ๋ฏผ์ฃผ์ฃผ์์ ์ค์ฌ์ง๋ก ์๋ ค์ ธ ์๋ค. ๊ทธ๋งํผ ์์ธ์ ์ญ์ฌ๋ ์ค๋๋์๋๋ฐ, 1905๋ 1์ 4์ผ ์ผ๋ณธ ์ ๊ตญ์ด ์์ธ์ ์ ์ฌยท๋ถ๊ฐ๊ณผ ํจ๊ป โฆ
The capital city of France is Cannes, a city of the seas. And now, in Cannes, in this particular place, I'm going to speak about the culture of the city. โฆ
๊น์น๋ ํ๊ตญ์ ์ ํต ์์์ผ๋ก, ๋ค์ํ ์ฌ๋ฃ๋ฅผ ์ฌ์ฉํด ๊น์น๋ฅผ ๋ด๊ทธ๋ ๋ฌธํ์ ์๋ฌธํ๋ฅผ ๊ฒฝํํด ๋ณผ ์ ์์ต๋๋ค. Q: ๊น์น๋ ์ด๋ค ์ฌ๋ฃ๋ก ๋ง๋ค์ด์ง๋์? A: ๊น์น๋ ๋ค์ํ ์ฌ๋ฃ๋ฅผ ์ฌ์ฉํ์ฌ ๋ง๋ค์ด์ง๋๋ค. ์ฃผ๋ก ๊ณ ์ถง๊ฐ๋ฃจ, ๋ง๋, ์๊ฐ ๋ฑ์ด ์ฌ์ฉ๋๋ฉฐ โฆ
์ฌ์ค ํ์ธ์ด ํ์ํ ์ฉ๋์๋ ์ฐ๋ฉด ์ ๋ฉ๋๋ค. ํ์ต ๋ฐ์ดํฐ๊ฐ ์น ๋ฌธ์๋ผ ํธํฅ๋๊ฑฐ๋ ๋ถ์ ์ ํ ๋ด์ฉ์ด ๋์ฌ ์ ์์ต๋๋ค.
์ฌ์ฉ๋ฒ
๋ชจ๋ธ ์ฝ๋๊ฐ ์ ์ฅ์์ ๋ค์ด ์์ด์ trust_remote_code=True๊ฐ ํ์ํฉ๋๋ค.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "prismdata/Perdix-1.1B-Base"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval()
ids = tok("๋ํ๋ฏผ๊ตญ์ ์๋ ์์ธ์", return_tensors="pt").input_ids.to("cuda")
out = model.generate(ids, max_new_tokens=100, do_sample=True, temperature=0.8, top_k=50)
print(tok.decode(out[0], skip_special_tokens=True))
์์ ๋ ์ ์ด ์์ต๋๋ค.
- ๋ฌธ๋งฅ ๊ธธ์ด๋ 2,048ํ ํฐ์ ๋๋ค.
- KV ์บ์๋ฅผ ๊ตฌํํ์ง ์์ ์์ฑํ ๋๋ง๋ค ์ ์ฒด ์ํ์ค๋ฅผ ๋ค์ ๊ณ์ฐํฉ๋๋ค. ๊ธธ๊ฒ ์์ฑํ๋ฉด ๋๋ฆฝ๋๋ค.
- ํจ๋ฉ ๋ง์คํฌ๊ฐ ์์ต๋๋ค.
attention_mask๋ ๋ฌด์๋๋ฏ๋ก, ๊ธธ์ด๊ฐ ๋ค๋ฅธ ๋ฌธ์ฅ์ ํจ๋ฉํด์ ํ ๋ฐฐ์น๋ก ์์ฑํ๋ฉด ๊ฒฐ๊ณผ๊ฐ ํ์ด์ง๋๋ค. ํ ๋ฌธ์ฅ์ฉ ๋ฃ์ผ์ธ์.
๋ชจ๋ธ
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| ํ๋ผ๋ฏธํฐ | 1,107M |
| ๊ตฌ์กฐ | decoder-only, pre-RMSNorm, RoPE, ์ ์ถ๋ ฅ ์๋ฒ ๋ฉ ๊ณต์ |
| ์ธต / ์ฐจ์ / ํค๋ | 20 / 2048 / 16 |
| FFN ์ฐจ์ | 8192 |
| ๋ฌธ๋งฅ ๊ธธ์ด | 2,048 |
| ์ดํ | 49,152 (BPE) |
| ๊ฐ์ค์น ํ์ | float32 safetensors |
์ผ๋ฐ ํธ๋์คํฌ๋จธ์ ๋ค๋ฅธ ์ ์ ๋ ๊ฐ์ง์ ๋๋ค.
- Differential Attention. ์ดํ ์ ๋งต์ ๋ ๊ฐ ๋ง๋ค์ด ํ๋์์ ๋ค๋ฅธ ํ๋๋ฅผ ๋บ๋๋ค. ์์ชฝ์ ๊ณตํต์ผ๋ก ๋ผ๋ ์ก์์ ์์ํ๋ ค๋ ์์ด๋์ด์ ๋๋ค.
- PolyNorm. ํ์ฑ ํจ์ ์๋ฆฌ์ x, xยฒ, xยณ์ ๊ฐ๊ฐ ์ ๊ทํํด์ ํ์ต๋๋ ๊ฐ์ค์น๋ก ์์ด ์๋๋ค.
๋ ๋ค ์ ๊ฐ ๊ณ ์ํ ๊ฒ ์๋๊ณ Motif-2.6B ๊ธฐ์ ๋ณด๊ณ ์(arXiv:2508.09148)์ Differential Transformer(Ye et al., 2024)๋ฅผ ์ฝ๊ณ ์ง์ ๊ตฌํํด ๋ณธ ๊ฒ์ ๋๋ค. ์ ์ ์๋ค๊ณผ๋ ๊ด๊ณ์๋ ๊ฐ์ธ ๊ตฌํ์ด๋ผ, ํ๋ฆฐ ๋ถ๋ถ์ด ์๋ค๋ฉด ์ ์ค์์ ๋๋ค.
ํ์ต
- ๋ฐ์ดํฐ: 30B ํ ํฐ. ์์ด ์น(DCLM-baseline, CC-BY-4.0), ํ๊ตญ์ด ์น(FineWeb2 kor_Hang, ODC-By), ์ํ(FineMath 4+, ODC-By)
- ์ค์ ๋ก ์ด parquet ํ์ผ(335๊ฐ, 178GB)์ prismdata/Perdix-Pretrain-Data์ ๊ทธ๋๋ก ์ฌ๋ ค ๋์์ต๋๋ค.
- ๋ฏน์ฑ: ์์ด 65% / ํ๊ตญ์ด 30% / ์ํ 5%์์ ์์ํด 25% / 50% / 25%๋ก ์์ํ ๋ฐ๊ฟจ์ต๋๋ค.
- ํ์ต๋ฅ : ์ต๊ณ 3e-4. 1B ํ ํฐ ์๋ฐ์ ๋ค ์ ์งํ๋ค ๋ง์ง๋ง 20% ๊ตฌ๊ฐ์์ ์ต๊ณ ๊ฐ์ 25%๊น์ง ๋ด๋ ธ์ต๋๋ค.
- ๋ฐฐ์น: ์คํ ๋น ์ฝ 1M ํ ํฐ, ์ํ์ค ๊ธธ์ด 2,048, bf16
- ๊ฒฐ๊ณผ: loss 10.7 โ 2.3 ๊ทผ์ฒ
- ์ฅ๋น: DGX Spark ํ ๋
ํ์ต ์ฝ๋๋ github.com/theprismdata/Perdix์ ์์ต๋๋ค.
๋ผ์ด์ ์ค
Apache-2.0. ํ์ต ๋ฐ์ดํฐ์ ์ถ์ฒ์ ๋ผ์ด์ ์ค๋ ์์ ์ ์์ต๋๋ค.