How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Quick Links
NeuralAI Nae1 banner

🧠 NeuralAI β€” Nae1

Trained from scratch. Every weight owned.
A 100M-parameter decoder pretrained end-to-end on 768M FineWeb tokens β€” no distillation, no fine-tune of someone else's base.

Nae1 100.1M params PPL 54.16 32000 steps Apache 2.0


πŸš€ Quick facts

Property Value
Architecture Llama-style decoder β€” RoPE, GQA, SwiGLU
Parameters 100,111,872 (100.1M)
Layers / hidden / heads 12 layers Β· 768 hidden Β· 12 heads Β· 4 KV heads (GQA)
Vocabulary 439 tokens (ByteLevel BPE, sliced from a 32,000-target config)
Context 2,048 trained Β· evaluated on n_ctx=512
Training data 768M tokens β€” FineWeb CC-MAIN-2013-20, 400k documents, 8 shards
Training schedule 32,000 steps β€” eight cosine cycles (0β†’4k→…→28kβ†’32k) Β· Kaggle T4
Final training loss 17.93 @ step 32,000 (sum over 8 grad-accum micros; cycle 8 is the first trained on corrected next-token labels β€” resume started at 28.87 under the fixed objective, so pre-fix values are not comparable)
Held-out perplexity 54.16 (Q4_K_M, 992 docs / 1,027,766 tokens, teacher-forced; cross-checked 57.09 on a fresh unseen FineWeb slice vs step-28000's 143.64 β€” the improvement is real generalization)
Formats in this repo f32 GGUF (305 MB) Β· Q4_K_M GGUF (45 MB) Β· config + tokenizer
License Apache 2.0

πŸ“ˆ Training

Nae1 training loss step 28000 to 32000

Nae1 is pretrained from random initialization in eight cosine cycles: 0 β†’ 4,000 Β· 4,000 β†’ 8,000 Β· … Β· 24,000 β†’ 28,000 Β· 28,000 β†’ 32,000, each resumed from the previous checkpoint. The chart shows cycle eight β€” the first cycle trained with corrected next-token labels (an earlier double-shift bug had models predicting token t+2; the held-out gate was always standard next-token, so historical PPL numbers remain valid). Under the corrected objective the logged loss resumes at 28.87, falls to a best of 15.72 (step 31,500) and finishes at 17.93 as the learning rate anneals to zero. Held-out PPL improves 147.97 β†’ 146.17 β†’ 54.16 β€” a 63% drop from fixing the training objective; a fresh-corpus cross-check (unseen FineWeb CC-MAIN-2014-10: 57.09 vs step-28000's 143.64) confirms the jump is genuine generalization, not eval-set leakage.

  • Data pipeline: FineWeb parquet β†’ chunked into 513-token windows served as zero-copy views (peak-RAM-safe at 768M tokens)
  • Checkpoints: every 250 steps; this release ships the final step-28000 checkpoint

πŸ“Š Evaluation

Held-out perplexity comparison across Nae1 checkpoints

Teacher-forced perplexity on a held-out FineWeb slice (992 documents / 1,027,766 tokens, n_ctx=512), scored token-by-token with exact row alignment β€” not the echo=True logprob path, which mispairs positions and inflates NLL.

Checkpoint (Q4_K_M) Held-out PPL
step-2000 192.97
step-4000 (resumed run) 188.89
step-4000 (from scratch) 179.17
step-8000 165.92
step-12000 154.25
step-16000 151.45
step-20000 148.84
step-24000 147.97
step-28000 146.17
step-32000 β€” this release 54.16

Raw report: eval_step32000_q4_n512.json β€” mean NLL 3.9919 nats/token; fresh-corpus cross-check: eval_fresh_step32000_q4_n512.json / eval_fresh_step28000_q4_n512.json (prior releases kept alongside: eval_step24000_q4_n512.json, eval_step20000_q4_n512.json, eval_step16000_q4_n512.json, eval_step8000_q4_n512.json). Fixed probe ("The quick brown fox", 10 tokens): 174.73 β€” the 10-token probe is noisy; held-out decides.

⚠️ Perplexity is the trust signal here, not raw capability. At 100M params over 768M tokens, expect exploratory, often garbled text β€” this is a research-lineage model, not a chat assistant.

Logits Parity Notice

Logits parity with the Nae1 native forward pass is not achievable by design: llama.cpp uses RMSNorm (no bias) while the native model uses LayerNorm (with bias), so conversion skips the bias tensors. Tensor integrity is fully verified (109/109 tensors match the source safetensors, max diff 0.00e+00) β€” small logit drift is a runtime constraint, not a conversion bug.


πŸ› οΈ Usage

llama.cpp (recommended)

llama-server -m nae1-llama-Q4_K_M.gguf -c 512 --host 127.0.0.1 --port 8080

llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="nae1-llama-Q4_K_M.gguf",
    n_ctx=512,
    verbose=False,
)

out = llm("The future of AI", max_tokens=32, temperature=0.7)
print(out["choices"][0]["text"])

Files

File What it is
nae1-llama-Q4_K_M.gguf 4-bit quant, 4.80 BPW, 45 MB β€” the recommended artifact
nae1-llama-f32.gguf Full-precision GGUF, 305 MB β€” for requantization / research
config.json Architecture config as trained
tokenizer/ Native 439-token vocab + merges + manifest
eval_step32000_q4_n512.json The exact evaluation report cited above
eval_fresh_step32000_q4_n512.json Fresh-corpus cross-check (unseen CC-MAIN-2014-10), this release
eval_fresh_step28000_q4_n512.json Fresh-corpus cross-check, prior release
eval_step28000_q4_n512.json Prior release (step-28000), kept for lineage
eval_step24000_q4_n512.json Prior release (step-24000), kept for lineage
eval_step20000_q4_n512.json Prior release (step-20000), kept for lineage
eval_step16000_q4_n512.json Prior release (step-16000), kept for lineage
eval_step8000_q4_n512.json Prior release (step-8000), kept for lineage

🧰 What Is NeuralAI?

NeuralAI is a local-first, private generative AI engine built by De'Andrew Preston Harris. The mission: your AI, on your hardware, under your control. Nae1 is the project's first end-to-end pretrained model β€” trained from random init on public data, converted with verified tensor integrity, and evaluated with a reproducible protocol.


⚠️ Limitations

  • Scale: 100M parameters pretrained on 768M tokens is a research checkpoint β€” long-form reasoning, coding, and factual recall are limited.
  • Vocabulary: a 439-token BPE trained on this corpus; out-of-domain text will tokenize inefficiently.
  • No chat template: base completion only β€” it was never instruction-tuned (SFT completed; checkpoint selection in progress).
  • No internet access: pair with a tool layer if you need live data.

πŸ‘€ Who Created NeuralAI?

NeuralAI was built from resilience, fatherhood, and the belief that personal computing deserves personal intelligence. Every release is handcrafted, iterated, and documented in the open.


πŸ“– Citation

@software{neuralai_nae1_2026,
  author       = {Harris, De'Andrew Preston},
  title        = {NeuralAI β€” Nae1},
  year         = {2026},
  url          = {https://huggingface.co/Subject-Emu-5259/NeuralAI-Nae1},
  version      = {step-32000},
  description  = {A 100M-parameter language model pretrained from scratch on 768M FineWeb tokens}
}

Built with discipline by De'Andrew Preston Harris. Maintained in the open. Updated whenever the model, dataset, or project state changes.

Downloads last month
377
GGUF
Model size
76.2M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support