Instructions to use AwakeningOS/VISTA-24M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AwakeningOS/VISTA-24M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AwakeningOS/VISTA-24M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AwakeningOS/VISTA-24M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AwakeningOS/VISTA-24M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AwakeningOS/VISTA-24M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AwakeningOS/VISTA-24M
- SGLang
How to use AwakeningOS/VISTA-24M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AwakeningOS/VISTA-24M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AwakeningOS/VISTA-24M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwakeningOS/VISTA-24M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AwakeningOS/VISTA-24M with Docker Model Runner:
docker model run hf.co/AwakeningOS/VISTA-24M
Download ARCHITECTURE.md from AwakeningOS/VISTA-24M: direct link, hf CLI and curl.
- Browser
- Download file 3.5 kB
-
https://huggingface.co/AwakeningOS/VISTA-24M/resolve/main/ARCHITECTURE.md
- Command line
-
hf download hf://AwakeningOS/VISTA-24M/ARCHITECTURE.md
-
curl -L -o ARCHITECTURE.md https://huggingface.co/AwakeningOS/VISTA-24M/resolve/main/ARCHITECTURE.md
Architecture and tensor flow
Let S(x) = x / sqrt(mean(x²) + 1e-6) along the feature dimension. This normalization gives approximately unit RMS, so the radius in 256-dimensional Euclidean coordinates is approximately 16. Learned RMSNorm layers additionally multiply by a learned feature-wise gain.
For layer l, write its input as h, its post-Attention state as a, and its final state as y.
Attention content and variance
From RMSNorm(h), linear projections produce Q, K, V with 8 heads of width 32. Q and K receive RoPE (base 10,000). Causal, document-isolated Attention weights are conceptually:
P = softmax(Q Kᵀ / sqrt(32) + mask)
The implementation computes two moments with the same weights:
mean = P V
second = P (V²)
variance = max(second - mean², 0)
A paired kernel evaluates [V, V²]; zero-padding Q/K to width 64 allows the paired call while explicitly preserving the original 1/sqrt(32) scale. Variance is calculated in each head's own Value coordinates, before the output gate and output projection. It measures variation across attended positions, not a comparison of unrelated coordinates across heads.
The first token in each document has exactly one eligible Value. Its variance is explicitly set to zero to avoid amplifying BF16 subtraction noise.
The weighted mean follows the content path:
gate = sigmoid(W_gate RMSNorm(h))
a = S(h + W_o concat_heads(gate * mean))
The flattened variance follows its own path. With r = log1p(variance / 0.1) and m = mean(r²):
direction = r / sqrt(m + eps)
magnitude = log1p(sqrt(m + eps)) - log1p(sqrt(eps))
The code uses an algebraically equivalent rationalized expression for magnitude to reduce cancellation. Concatenation yields 257 features. A bias-free 257 → 1,792 projection maps them into the two SwiGLU branches. This projection starts at zero during pretraining.
Previous-layer differences
For the previous layer, record:
dA = S(a_previous - h_previous)
dF = S(y_previous - a_previous)
Both remain 256-dimensional. A learned query from the current a and shared keys from dA and dF are 16-dimensional. Their dot products divided by 4 produce two softmax weights, pA and pF.
The contribution is:
delta_input = pA * W_A dA + pF * W_F dF
W_A and W_F are independent 256 → 1,792 matrices. The first layer omits this reader because it has no previous layer. Later layers consume only the immediately previous pair. Gradients flow through both differences and routing weights during training.
SwiGLU and output
pre = W_base RMSNorm(a) + delta_input + W_var [direction; magnitude]
g, u = split(pre)
y = S(a + W_down dropout(SiLU(g) * u))
After layer 7, a final learned RMSNorm and an untied 256 → 16,384 LM head produce logits. The main residual stream remains full-width throughout. There is no cross-token recurrent cache or separate 16-dimensional persistent state in this model.
Implementation boundaries
Training used BF16 autocast, FP32 parameters and residuals, and FlashAttention. The public default uses SDPA for portability; backend rounding can differ. The published evaluation values came from the original evaluation backend, not a rerun of the entire suite on CPU.
The public config adds standard Transformers shape fields, registers AutoModel, sets use_cache=False, and selects SDPA. The trained model core and all parameter tensors are retained. See the release test report for CPU loading and adapter checks.