Instructions to use Subject-Emu-5259/NeuralAI-Nae1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Subject-Emu-5259/NeuralAI-Nae1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Subject-Emu-5259/NeuralAI-Nae1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Subject-Emu-5259/NeuralAI-Nae1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Subject-Emu-5259/NeuralAI-Nae1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Ollama
How to use Subject-Emu-5259/NeuralAI-Nae1 with Ollama:
ollama run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Subject-Emu-5259/NeuralAI-Nae1 with Docker Model Runner:
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Lemonade
How to use Subject-Emu-5259/NeuralAI-Nae1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Run and chat with the model
lemonade run user.NeuralAI-Nae1-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M# Run inference directly in the terminal:
llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_MUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M# Run inference directly in the terminal:
./llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_MBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M# Run inference directly in the terminal:
./build/bin/llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_MUse Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
π§ NeuralAI β Nae1
Trained from scratch. Every weight owned.
A 100M-parameter decoder pretrained end-to-end on 768M FineWeb tokens β no distillation, no fine-tune of someone else's base.
π Quick facts
| Property | Value |
|---|---|
| Architecture | Llama-style decoder β RoPE, GQA, SwiGLU |
| Parameters | 100,111,872 (100.1M) |
| Layers / hidden / heads | 12 layers Β· 768 hidden Β· 12 heads Β· 4 KV heads (GQA) |
| Vocabulary | 439 tokens (ByteLevel BPE, sliced from a 32,000-target config) |
| Context | 2,048 trained Β· evaluated on n_ctx=512 |
| Training data | 768M tokens β FineWeb CC-MAIN-2013-20, 400k documents, 8 shards |
| Training schedule | 32,000 steps β eight cosine cycles (0β4kββ¦β28kβ32k) Β· Kaggle T4 |
| Final training loss | 17.93 @ step 32,000 (sum over 8 grad-accum micros; cycle 8 is the first trained on corrected next-token labels β resume started at 28.87 under the fixed objective, so pre-fix values are not comparable) |
| Held-out perplexity | 54.16 (Q4_K_M, 992 docs / 1,027,766 tokens, teacher-forced; cross-checked 57.09 on a fresh unseen FineWeb slice vs step-28000's 143.64 β the improvement is real generalization) |
| Formats in this repo | f32 GGUF (305 MB) Β· Q4_K_M GGUF (45 MB) Β· config + tokenizer |
| License | Apache 2.0 |
π Training
Nae1 is pretrained from random initialization in eight cosine cycles: 0 β 4,000 Β· 4,000 β 8,000 Β· β¦ Β· 24,000 β 28,000 Β· 28,000 β 32,000, each resumed from the previous checkpoint. The chart shows cycle eight β the first cycle trained with corrected next-token labels (an earlier double-shift bug had models predicting token t+2; the held-out gate was always standard next-token, so historical PPL numbers remain valid). Under the corrected objective the logged loss resumes at 28.87, falls to a best of 15.72 (step 31,500) and finishes at 17.93 as the learning rate anneals to zero. Held-out PPL improves 147.97 β 146.17 β 54.16 β a 63% drop from fixing the training objective; a fresh-corpus cross-check (unseen FineWeb CC-MAIN-2014-10: 57.09 vs step-28000's 143.64) confirms the jump is genuine generalization, not eval-set leakage.
- Data pipeline: FineWeb parquet β chunked into 513-token windows served as zero-copy views (peak-RAM-safe at 768M tokens)
- Checkpoints: every 250 steps; this release ships the final step-28000 checkpoint
π Evaluation
Teacher-forced perplexity on a held-out FineWeb slice (992 documents / 1,027,766 tokens, n_ctx=512), scored token-by-token with exact row alignment β not the echo=True logprob path, which mispairs positions and inflates NLL.
| Checkpoint (Q4_K_M) | Held-out PPL |
|---|---|
| step-2000 | 192.97 |
| step-4000 (resumed run) | 188.89 |
| step-4000 (from scratch) | 179.17 |
| step-8000 | 165.92 |
| step-12000 | 154.25 |
| step-16000 | 151.45 |
| step-20000 | 148.84 |
| step-24000 | 147.97 |
| step-28000 | 146.17 |
| step-32000 β this release | 54.16 |
Raw report: eval_step32000_q4_n512.json β mean NLL 3.9919 nats/token; fresh-corpus cross-check: eval_fresh_step32000_q4_n512.json / eval_fresh_step28000_q4_n512.json (prior releases kept alongside: eval_step24000_q4_n512.json, eval_step20000_q4_n512.json, eval_step16000_q4_n512.json, eval_step8000_q4_n512.json). Fixed probe ("The quick brown fox", 10 tokens): 174.73 β the 10-token probe is noisy; held-out decides.
β οΈ Perplexity is the trust signal here, not raw capability. At 100M params over 768M tokens, expect exploratory, often garbled text β this is a research-lineage model, not a chat assistant.
Logits Parity Notice
Logits parity with the Nae1 native forward pass is not achievable by design: llama.cpp uses RMSNorm (no bias) while the native model uses LayerNorm (with bias), so conversion skips the bias tensors. Tensor integrity is fully verified (109/109 tensors match the source safetensors, max diff 0.00e+00) β small logit drift is a runtime constraint, not a conversion bug.
π οΈ Usage
llama.cpp (recommended)
llama-server -m nae1-llama-Q4_K_M.gguf -c 512 --host 127.0.0.1 --port 8080
llama-cpp-python
from llama_cpp import Llama
llm = Llama(
model_path="nae1-llama-Q4_K_M.gguf",
n_ctx=512,
verbose=False,
)
out = llm("The future of AI", max_tokens=32, temperature=0.7)
print(out["choices"][0]["text"])
Files
| File | What it is |
|---|---|
nae1-llama-Q4_K_M.gguf |
4-bit quant, 4.80 BPW, 45 MB β the recommended artifact |
nae1-llama-f32.gguf |
Full-precision GGUF, 305 MB β for requantization / research |
config.json |
Architecture config as trained |
tokenizer/ |
Native 439-token vocab + merges + manifest |
eval_step32000_q4_n512.json |
The exact evaluation report cited above |
eval_fresh_step32000_q4_n512.json |
Fresh-corpus cross-check (unseen CC-MAIN-2014-10), this release |
eval_fresh_step28000_q4_n512.json |
Fresh-corpus cross-check, prior release |
eval_step28000_q4_n512.json |
Prior release (step-28000), kept for lineage |
eval_step24000_q4_n512.json |
Prior release (step-24000), kept for lineage |
eval_step20000_q4_n512.json |
Prior release (step-20000), kept for lineage |
eval_step16000_q4_n512.json |
Prior release (step-16000), kept for lineage |
eval_step8000_q4_n512.json |
Prior release (step-8000), kept for lineage |
π§° What Is NeuralAI?
NeuralAI is a local-first, private generative AI engine built by De'Andrew Preston Harris. The mission: your AI, on your hardware, under your control. Nae1 is the project's first end-to-end pretrained model β trained from random init on public data, converted with verified tensor integrity, and evaluated with a reproducible protocol.
β οΈ Limitations
- Scale: 100M parameters pretrained on 768M tokens is a research checkpoint β long-form reasoning, coding, and factual recall are limited.
- Vocabulary: a 439-token BPE trained on this corpus; out-of-domain text will tokenize inefficiently.
- No chat template: base completion only β it was never instruction-tuned (SFT completed; checkpoint selection in progress).
- No internet access: pair with a tool layer if you need live data.
π€ Who Created NeuralAI?
- Founder & Lead Architect: De'Andrew Preston Harris (D. Harris / Dre)
- Hugging Face: @Subject-Emu-5259
- GitHub: @Subject-Emu-5259
- LinkedIn: linkedin.com/in/deandrewharris94
- Location: Memphis, Tennessee / West Memphis, Arkansas
- Education: AI Software Engineering at Maestro College
NeuralAI was built from resilience, fatherhood, and the belief that personal computing deserves personal intelligence. Every release is handcrafted, iterated, and documented in the open.
π Citation
@software{neuralai_nae1_2026,
author = {Harris, De'Andrew Preston},
title = {NeuralAI β Nae1},
year = {2026},
url = {https://huggingface.co/Subject-Emu-5259/NeuralAI-Nae1},
version = {step-32000},
description = {A 100M-parameter language model pretrained from scratch on 768M FineWeb tokens}
}
Built with discipline by De'Andrew Preston Harris. Maintained in the open. Updated whenever the model, dataset, or project state changes.
- Downloads last month
- 377
4-bit
32-bit
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M# Run inference directly in the terminal: llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M