Instructions to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/Tiny-LLM-PDelta3-GDN2-InputRoute")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute
- SGLang
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with Docker Model Runner:
docker model run hf.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute
Download validation_context_summary.csv from vtava/Tiny-LLM-PDelta3-GDN2-InputRoute: direct link, hf CLI and curl.
- Browser
- Download file 1.75 kB
-
https://huggingface.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute/resolve/main/validation_context_summary.csv
- Command line
-
hf download hf://vtava/Tiny-LLM-PDelta3-GDN2-InputRoute/validation_context_summary.csv
-
curl -L -o validation_context_summary.csv https://huggingface.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute/resolve/main/validation_context_summary.csv
1.75 kB
| candidate_nll,candidate_ppl,context,delta_nll,name,state_bytes,state_vs_transformer_fp16,transformer_nll,transformer_ppl | |
| 4.329811489582061,75.92997162928616,256,0.23298438787460274,conv4_pdelta_f96,19392,0.197265625,4.096827101707459,60.14913741228434 | |
| 4.29708606004715,73.48534951669176,512,0.21623675823211652,conv4_pdelta_f96,19392,0.0986328125,4.080849301815033,59.19572354203201 | |
| 4.334633874893188,76.29701952100608,1024,0.23117593526840174,conv4_pdelta_f96,19392,0.04931640625,4.1034579396247866,60.54930183986028 | |
| 4.336314916610718,76.42538585809872,256,0.2394878149032591,conv4_channel_decay_f96,19392,0.197265625,4.096827101707459,60.14913741228434 | |
| 4.303501772880554,73.95832603505018,512,0.22265247106552088,conv4_channel_decay_f96,19392,0.0986328125,4.080849301815033,59.19572354203201 | |
| 4.337036406993866,76.48054593536355,1024,0.23357846736907906,conv4_channel_decay_f96,19392,0.04931640625,4.1034579396247866,60.54930183986028 | |
| 4.306775999069214,74.20087919328297,256,0.20994889736175537,conv4_gdn2_f96,20736,0.2109375,4.096827101707459,60.14913741228434 | |
| 4.280840587615967,72.30118995581762,512,0.19999128580093384,conv4_gdn2_f96,20736,0.10546875,4.080849301815033,59.19572354203201 | |
| 4.316916954517365,74.95717530459878,1024,0.21345901489257812,conv4_gdn2_f96,20736,0.052734375,4.1034579396247866,60.54930183986028 | |
| 4.302966368198395,73.91873899949412,256,0.20613926649093628,conv4_gdn2_inputroute_f96,20736,0.2109375,4.096827101707459,60.14913741228434 | |
| 4.277678978443146,72.07296282305448,512,0.1968296766281128,conv4_gdn2_inputroute_f96,20736,0.10546875,4.080849301815033,59.19572354203201 | |
| 4.313992857933044,74.73831342690009,1024,0.21053491830825788,conv4_gdn2_inputroute_f96,20736,0.052734375,4.1034579396247866,60.54930183986028 | |