Instructions to use Maykeye/TinyLLama-v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Maykeye/TinyLLama-v0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Maykeye/TinyLLama-v0")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Maykeye/TinyLLama-v0") model = AutoModelForCausalLM.from_pretrained("Maykeye/TinyLLama-v0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Maykeye/TinyLLama-v0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Maykeye/TinyLLama-v0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Maykeye/TinyLLama-v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Maykeye/TinyLLama-v0
- SGLang
How to use Maykeye/TinyLLama-v0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Maykeye/TinyLLama-v0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Maykeye/TinyLLama-v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Maykeye/TinyLLama-v0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Maykeye/TinyLLama-v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Maykeye/TinyLLama-v0 with Docker Model Runner:
docker model run hf.co/Maykeye/TinyLLama-v0
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| import sys | |
| import os | |
| from pathlib import Path | |
| from tqdm.auto import tqdm | |
| model_id = os.getcwd() | |
| if len(sys.argv) == 2: | |
| filename = sys.argv[1] | |
| elif len(sys.argv) == 3: | |
| filename = sys.argv[1] | |
| model_id = sys.argv[2] | |
| else: | |
| raise Exception("use valid.py <path-to-text> [model-id]") | |
| text = Path(filename).read_text() | |
| stories = text.split("<|endoftext|>") | |
| print(len(stories)) | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained(model_id).cuda().bfloat16() | |
| ctx_size = tokenizer.model_max_length | |
| sliding_window = ctx_size // 2 | |
| total_loss = 0.0 | |
| measurements = 0 | |
| model.eval() | |
| for story in (bar := tqdm(stories)): | |
| story = story.strip() | |
| tokens = tokenizer(story, add_special_tokens=False).input_ids + [tokenizer.eos_token_id] | |
| i = 0 | |
| while i < len(tokens): | |
| current_window = tokens[i:i+ctx_size-1] | |
| part_tokens = [tokenizer.bos_token_id] + current_window | |
| input_ids = torch.tensor(part_tokens, device="cuda")[None] | |
| labels = input_ids.clone() | |
| if i: | |
| # disable seen tokens | |
| labels[:, :-sliding_window] = -100 | |
| with torch.no_grad(): | |
| loss = model(input_ids, labels=labels).loss | |
| total_loss += loss.item() | |
| measurements += 1 | |
| i += len(current_window) | |
| bar.set_description(f"L {total_loss/measurements:.4f}") | |
| print(f"FINAL LOSS: {total_loss/measurements:.4f}") | |