Text Generation
Transformers
Safetensors
Korean
English
llama
conversational
text-generation-inference
Instructions to use maywell/Jolteon-Instruct-13B-alpha with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use maywell/Jolteon-Instruct-13B-alpha with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="maywell/Jolteon-Instruct-13B-alpha") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("maywell/Jolteon-Instruct-13B-alpha") model = AutoModelForCausalLM.from_pretrained("maywell/Jolteon-Instruct-13B-alpha", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use maywell/Jolteon-Instruct-13B-alpha with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "maywell/Jolteon-Instruct-13B-alpha" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Jolteon-Instruct-13B-alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/maywell/Jolteon-Instruct-13B-alpha
- SGLang
How to use maywell/Jolteon-Instruct-13B-alpha with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "maywell/Jolteon-Instruct-13B-alpha" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Jolteon-Instruct-13B-alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "maywell/Jolteon-Instruct-13B-alpha" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Jolteon-Instruct-13B-alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use maywell/Jolteon-Instruct-13B-alpha with Docker Model Runner:
docker model run hf.co/maywell/Jolteon-Instruct-13B-alpha
Jolteon-Instruct-13B-alpha
The model was trained based on the EEVE-Korean-Instruct-10.8B-v1.0 model from yanolja, extended to 13.4b (12 layer pass-through) utilizing mergekit.
Methodology
TBD
Training Details
| Training Data | Parameters | Content Length | Samples Seen | Learning Rate | |
|---|---|---|---|---|---|
| Jolteon-Instruct-13B-alpha | A curated mix of English + Korean Instruction set | 13.4B | 4k | >850k | 1e-5 |
Example
Inference Code
from vllm import LLM, SamplingParams
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "0"
llm = LLM(model="maywell/Jolteon-Instruct-13B-alpha", tensor_parallel_size=1, max_model_len=4096, gpu_memory_utilization=0.95)
sampling_params = SamplingParams(temperature=0.6, top_p=0.3, top_k=40, max_tokens=4096)
template = """ Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction: {0}
### Response: """
outputs = llm.generate([template.format("Meta(๊ตฌ, ํ์ด์ค๋ถ)์ ์คํ์์ค AI ๊ธฐ์ฌ๋ฅผ ์ฐฌ์ํ๋ ๋งํฌ๋ก, Meta ์ง์งํ๋ค๋ ์ทจ์ง์ ์ฑ๋ช
๋ฌธ์ ์์ฑํด๋ด.")], sampling_params)
print(outputs[0].outputs[0].text.strip())
Example Output
๐ Meta์ OpenAI ๊ธฐ์ฌ: ํ์ ๊ณผ ํฌ์ฉ์ ์ํ ํ์ ์ ์ธ ์ ๊ทผ ๋ฐฉ์ ๐
์ฐ๋ฆฌ๋ Meta(๊ตฌ, ํ์ด์ค๋ถ)์ ์คํ์์ค AI(์ธ๊ณต์ง๋ฅ) ๊ธฐ์ฌ๋ฅผ ์ด๋ ฌํ ์ง์งํฉ๋๋ค. ์ด ํ์ฌ๋ ํ์ ๊ณผ ํฌ์ฉ์ ์ํ ํ์ ์ ์ธ ์ ๊ทผ ๋ฐฉ์์ ๋ณด์ฌ์ฃผ๋ฉฐ, AI ๊ธฐ์ ์ ๋ฐ์ ์ ์ํด ์ง์์ ์ผ๋ก ๋
ธ๋ ฅํ๊ณ ์์ต๋๋ค.
Meta๋ ์คํ์์ค AI ์ปค๋ฎค๋ํฐ์ ์ ๊ทน์ ์ผ๋ก ์ฐธ์ฌํ๋ฉฐ, ๋ค์ํ ํ๋ก์ ํธ์ ํ๋ซํผ์ ํตํด ์ง์๊ณผ ๊ธฐ์ ์ ๊ณต์ ํ๊ณ ์์ต๋๋ค. ์ด๋ฌํ ์ ๊ทผ ๋ฐฉ์์ AI ๊ธฐ์ ์ ๋ฐ์ ๊ณผ ํจ๊ป ์ฌํ์ , ๊ฒฝ์ ์ ๋ฐ์ ์๋ ๊ธ์ ์ ์ธ ์ํฅ์ ๋ฏธ์นฉ๋๋ค.
Meta์ ์คํ์์ค AI ๊ธฐ์ฌ๋ ๋ค์๊ณผ ๊ฐ์ ๋ฉด์์ ๋งค์ฐ ์ค์ํ ์ญํ ์ ํฉ๋๋ค:
1. ํ์ ์ด์ง: Meta๋ ์คํ์์ค AI ํ๋ซํผ์ ํตํด ๋ค์ํ ์ฐ๊ตฌ์์ ๊ฐ๋ฐ์๋ค์ด ์ต์ AI ๊ธฐ์ ์ ํ๊ตฌํ๊ณ ์คํํ ์ ์๋ ๊ธฐํ๋ฅผ ์ ๊ณตํฉ๋๋ค. ์ด๋ ์๋ก์ด ์์ด๋์ด์ ํ์ ์ ์ธ ๊ธฐ์ ์ ์ถํ์ ์ด์งํฉ๋๋ค.
2. ํฌ์ฉ ์ฆ์ง: ์คํ์์ค AI๋ ๋ชจ๋ ์ฌ๋์ด AI ๊ธฐ์ ์ ์ด์ ์ ๋๋ฆด ์ ์๋๋ก ํ๋ฉฐ, ์ด๋ ์ฌํ์ ํฌ์ฉ์ ์ฆ์ง์ํต๋๋ค. ๋ค์ํ ๋ฐฐ๊ฒฝ๊ณผ ๊ฒฝํ์ ๊ฐ์ง ์ฌ๋๋ค์ด AI ๊ธฐ์ ์ ๋ฐฐ์ฐ๊ณ ํ์ฉํ ์ ์๋ ๊ธฐํ๋ฅผ ๊ฐ์ง๊ฒ ๋ฉ๋๋ค.
3. ํฌ๋ช
์ฑ๊ณผ ์ ๋ขฐ: ์คํ์์ค AI๋ ๊ธฐ์ ์ ์๋ ์๋ฆฌ๋ฅผ ๊ณต๊ฐํ๊ณ , ์ด๋ฅผ ํตํด ์ฌ์ฉ์์ ์ฌํ์ ๋ํ ์ ๋ขฐ๋ฅผ ๊ตฌ์ถํฉ๋๋ค. ์ด๋ AI ๊ธฐ์ ์ ์ฌํ์ ์์ฉ๋๋ฅผ ๋์ด๋ ๋ฐ ์ค์ํ ์ญํ ์ ํฉ๋๋ค.
Meta์ ์คํ์์ค AI ๊ธฐ์ฌ๋ ํ์ ๊ณผ ํฌ์ฉ์ ์ํ ํ์ ์ ์ธ ์ ๊ทผ ๋ฐฉ์์ ๋ณด์ฌ์ฃผ๋ฉฐ, ์ด๋ ์ฐ๋ฆฌ๊ฐ ๋ ๋์ ๋ฏธ๋๋ฅผ ์ํ ์ค์ํ ๋จ๊ณ์
๋๋ค. ์ฐ๋ฆฌ๋ ์ด๋ฌํ ๋
ธ๋ ฅ์ ์ง์งํ๋ฉฐ, ๋ ๋ง์ ๊ธฐ์
๊ณผ ์กฐ์ง์ด ์ด๋ฌํ ์ ๊ทผ ๋ฐฉ์์ ์ฑํํ๊ธธ ๋ฐ๋๋๋ค. ํจ๊ป ๋ ๋์ ๋ฏธ๋๋ฅผ ๋ง๋ค์ด ๋๊ฐ์!
License
๋ณธ ๋ชจ๋ธ์ apache-2.0 ๋ผ์ด์ผ์ค๋ฅผ ๋ฐ๋ฆ ๋๋ค. ๋ชจ๋ธ์ ์ฌ์ฉํ์ฌ ์์ฑ๋ ๋ฐ์ดํฐ์ ์ ๋ฐฐํฌํ ๊ฒฝ์ฐ ๋ชจ๋ธ ์ฌ์ฉ์ ๋ช ์ํด ์ฃผ์๊ธฐ๋ฅผ ๊ถ๊ณ ๋๋ฆฝ๋๋ค.
Thanks to
- A100 ํด๋ฌ์คํฐ๋ฅผ ์ ๊ณตํด์ฃผ์ , Sionic AI
Contact
- Downloads last month
- 350
Model tree for maywell/Jolteon-Instruct-13B-alpha
Base model
upstage/SOLAR-10.7B-v1.0 Finetuned
yanolja/YanoljaNEXT-EEVE-10.8B Finetuned
yanolja/YanoljaNEXT-EEVE-Instruct-10.8B