|
Download README.md from North-ML1/vortex-alpha: direct link, hf CLI and curl.
- Browser
- Download file 7.61 kB
-
https://huggingface.co/North-ML1/vortex-alpha/resolve/main/README.md
- Command line
-
hf download hf://North-ML1/vortex-alpha/README.md
-
curl -L -o README.md https://huggingface.co/North-ML1/vortex-alpha/resolve/main/README.md
7.61 kB
| language: en | |
| pipeline_tag: text-generation | |
| library_name: pytorch | |
| tags: | |
| - causal-lm | |
| - small-language-model | |
| - research | |
| - text-generation | |
| license: other | |
| thumbnail: Aurora-5.png | |
| # Vortex Alpha | |
|  | |
| [](https://colab.research.google.com/#create=true&url=https://huggingface.co/North-ML1/vortex-alpha/resolve/main/Vortex_Alpha_Colab.ipynb) | |
| Vortex is the working name for a compact, experimental language model. The | |
| final public name has not been decided. This release is intended for research, | |
| local experimentation, and further fine-tuning—not as a finished general | |
| assistant. | |
| ## What is included | |
| - `model.safetensors`: the instruction/tool-format preview selected from the | |
| best small internal behavior pilot. | |
| - `base_model.safetensors`: the corresponding pretrained text-completion | |
| base. | |
| - `config.json`: the architecture configuration. | |
| - `tokenizer.model`: the 8,192-piece SentencePiece tokenizer used for both | |
| checkpoints. | |
| - `vortex_model.py` and `inference.py`: a minimal dependency-light PyTorch | |
| loader and sampler. | |
| - `configuration_vortex.py`, `modeling_vortex.py`, and | |
| `tokenization_vortex.py`: standard Transformers remote-code modules for | |
| `AutoModelForCausalLM` and `AutoTokenizer`. | |
| - `Vortex_Alpha_Colab.ipynb`: a one-click Google Colab quickstart. | |
| - `requirements.txt` and `chat_template.jinja`: convenience metadata for | |
| local and notebook use. | |
| - `Aurora-5.png`: the project thumbnail/banner. | |
| Optimizer state, private logs, local paths, credentials, and training-machine | |
| metadata are intentionally not included. | |
| ## Architecture | |
| Vortex is a dense decoder-only Transformer with 174,942,720 trainable | |
| parameters: | |
| | Component | Parameters | | |
| | --- | ---: | | |
| | Shared token embedding and tied output head | 8,388,608 | | |
| | Attention Q projections | 12,582,912 | | |
| | Attention K projections | 3,145,728 | | |
| | Attention V projections | 3,145,728 | | |
| | Attention output projections | 12,582,912 | | |
| | Per-head QK RMSNorm parameters | 1,536 | | |
| | SwiGLU feed-forward networks | 135,069,696 | | |
| | Transformer-block RMSNorm parameters | 24,576 | | |
| | Final RMSNorm | 1,024 | | |
| | **Total** | **174,942,720** | | |
| Configuration: 12 layers, hidden size 1,024, 16 query heads, 4 key/value | |
| heads, 64-dimensional heads, SwiGLU with intermediate size 3,664, pre-layer | |
| RMSNorm, per-head QK-Norm, RoPE with base 100,000, bias-free projections, | |
| 8,192-token vocabulary, and a 4,096-token training/inference limit. | |
| The published weights use the readable reference PyTorch layout. They do not | |
| require Transformer Engine to load. The input and output embeddings are tied. | |
| ## Quick start | |
| The minimal reference runner uses PyTorch, SentencePiece, and | |
| `safetensors`: | |
| ```bash | |
| python -m pip install torch sentencepiece safetensors | |
| python inference.py \ | |
| --weights model.safetensors \ | |
| --chat \ | |
| --prompt "Explain why the sky appears blue in two short paragraphs." | |
| ``` | |
| For the normal Hugging Face API, load the custom architecture through the | |
| repository's small remote-code modules: | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| repo = "North-ML1/vortex-alpha" | |
| tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| repo, trust_remote_code=True, torch_dtype="auto" | |
| ) | |
| inputs = tokenizer("Explain why the sky appears blue.", return_tensors="pt") | |
| outputs = model.generate(**inputs, max_new_tokens=80, do_sample=False) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| The tokenizer also exposes the chat template directly: | |
| ```python | |
| messages = [{"role": "user", "content": "What is photosynthesis?"}] | |
| chat_inputs = tokenizer.apply_chat_template( | |
| messages, add_generation_prompt=True, return_tensors="pt" | |
| ) | |
| outputs = model.generate(**chat_inputs, max_new_tokens=80, do_sample=False) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| The `trust_remote_code=True` flag is required because Vortex's QK-Norm and | |
| GQA implementation is not one of the built-in Transformers model classes. | |
| For ordinary next-token completion, use the base checkpoint: | |
| ```bash | |
| python inference.py \ | |
| --weights base_model.safetensors \ | |
| --prompt "The sky appears blue because" | |
| ``` | |
| The reference runner recomputes the full prefix at each generated token and | |
| is deliberately simple. A production runner should add a KV cache and a | |
| fused attention implementation. | |
| The instruction preview was tuned with this compact serialization: | |
| ```text | |
| [SYSTEM] | |
| You are a helpful assistant. Follow instructions, answer clearly, and say | |
| when information is missing. | |
| </s> | |
| [USER] | |
| Your question here | |
| </s> | |
| [ASSISTANT] | |
| ``` | |
| The instruction preview may emit a `CALL {json}` calculator/search request | |
| when prompted for tool use. No tool server is included in this repository; | |
| without a tool runner, treat such output as ordinary text. The base checkpoint | |
| is the better starting point for continued pretraining. | |
| ## Training summary | |
| The base run used an approximate mixture of FineWeb, DCLM, educational/math | |
| material, and The Stack v3 code data. The recorded base checkpoint had seen | |
| about 8.48 billion pretraining tokens. Training used BF16 model computation, | |
| FP8-capable NVIDIA kernels where available, a WSD-style learning-rate tail, | |
| and token-budgeted batches designed for a 16 GB consumer GPU. The instruction | |
| preview is a lightweight supervised derivative of that base; its optimizer | |
| state and private training records are not part of this release. | |
| These data-mixture descriptions are a project-level summary, not a claim that | |
| every upstream document is suitable for every downstream use. Follow the | |
| licenses and terms of the upstream datasets. | |
| ## Evaluation snapshot | |
| These are exploratory measurements, not official leaderboard submissions. | |
| The base results used greedy decoding, no tools, and the stated sample sizes: | |
| | Test | Result | Notes | | |
| | --- | ---: | --- | | |
| | MMLU cloze sample | 511/2,000 = 25.55% | Wilson 95% interval: 23.69–27.51% | | |
| | GSM8K strict numeric sample | 0/256 = 0.00% | Wilson 95% upper bound: 1.48% | | |
| | GSM8K fallback numeric sample | 2/256 = 0.78% | Wilson 95% interval: 0.21–2.80% | | |
| The instruction checkpoint reached 9/12 arithmetic, 4/4 grounding, 4/4 | |
| abstention, 2/4 exact-format, and 1/2 JSON checks on a 26-prompt internal | |
| tool-format pilot when the calculator runner was available. That pilot is too | |
| small to support general capability claims, and the arithmetic result is not | |
| comparable to tool-free GSM8K. | |
| The results show why this is an alpha release: the model can produce useful | |
| local completions and structured tool calls, but it remains weak at reliable | |
| arithmetic, broad knowledge, long-form coherence, and hallucination control. | |
| ## Limitations and intended use | |
| Vortex is a small research model. It can be repetitive, overconfident, or | |
| factually wrong; architecture and training scale do not guarantee reliable | |
| answers. Do not use it as the sole basis for medical, legal, financial, | |
| safety-critical, or other high-stakes decisions. It has not been evaluated for | |
| privacy, bias, cybersecurity, or comprehensive safety. | |
| The repository is public, but no open-source license is asserted yet. The | |
| final name, licensing terms, and a production release decision are still open. | |
| ## Reproducibility note | |
| The conversion removed optimizer state and Transformer Engine-only auxiliary | |
| state, then wrote the model tensors in BF16 safetensors format. The exported | |
| reference tensors preserve the tied-embedding model weights and can be loaded | |
| with the included `vortex_model.py`. | |