Text Generation
Transformers
Safetensors
PyTorch
English
modern_dense_mha_gated_ffn_router
custom_code
causal-lm
small-language-model
babylm
strict-small
swiglu
research
Instructions to use AwakeningOS/VISTA-24M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AwakeningOS/VISTA-24M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AwakeningOS/VISTA-24M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AwakeningOS/VISTA-24M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AwakeningOS/VISTA-24M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AwakeningOS/VISTA-24M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AwakeningOS/VISTA-24M
- SGLang
How to use AwakeningOS/VISTA-24M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AwakeningOS/VISTA-24M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AwakeningOS/VISTA-24M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AwakeningOS/VISTA-24M with Docker Model Runner:
docker model run hf.co/AwakeningOS/VISTA-24M
Release VISTA-24M: model, architecture diagrams, training recipe and evaluation evidence
9287d39 verified Download training/source_identity.json from AwakeningOS/VISTA-24M: direct link, hf CLI and curl.
- Browser
- Download file 883 Bytes
-
https://huggingface.co/AwakeningOS/VISTA-24M/resolve/main/training/source_identity.json
- Command line
-
hf download hf://AwakeningOS/VISTA-24M/training/source_identity.json
-
curl -L -o source_identity.json https://huggingface.co/AwakeningOS/VISTA-24M/resolve/main/training/source_identity.json
883 Bytes
| { | |
| "dense.py": "3303f30eda7607136724dfe4124e54010426a5334c42b4227551e02253e6566a", | |
| "configuration_dense.py": "227987f030b3e56ce8c38960ee8d77fa000b355dfe78fcc6c45c05dbe418c544", | |
| "modeling_dense.py": "8fe6cb72985ccccfba423482e1135bee8ad17523b5c91e9c3a261fb3678fd9f4", | |
| "runtime.py": "6272346e7a185e308c002f33b4bb709d028e295c549d6de0e727e165ad272e0b", | |
| "train.py": "2f7ad9e8298ce0e0cea9b23e93a083de8e9fd6d8b724a9f9ed4bccfefe63fa51", | |
| "exclusive.py": "ad81a793b13ecf37cebc3cc6e1d5516d73dc9dcbad362f259074b79b99655823", | |
| "test_cpu.py": "105b3e44634993fdb8ce967bbdbe3c47d47e47ab18217cad0c30b3d4726ad4c4", | |
| "gpu_acceptance.py": "d0f5e23fc082aad9540414c295de1a060cd7575d85063314e0b60a113a5c96b9", | |
| "resume_acceptance.py": "b654fbe40cbd48ea84dc2bfcaf52ce57845506c53b80090c88dce6ebdc094ced", | |
| "parent_dense.py": "eb30a0a73e7a82e184d59ae02a2499be569c452f426fa1dc8b8729d63becf19f" | |
| } | |