Instructions to use ToTo-40417/EXLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- Ollama
How to use ToTo-40417/EXLLM with Ollama:
ollama run hf.co/ToTo-40417/EXLLM:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM:F16
- Lemonade
How to use ToTo-40417/EXLLM with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM:F16
Run and chat with the model
lemonade run user.EXLLM-F16
List all available models
lemonade list
- Atomic Chat
Download DATA_PROVENANCE.md from ToTo-40417/EXLLM: direct link, hf CLI and curl.
- Browser
- Download file 2.56 kB
-
https://huggingface.co/ToTo-40417/EXLLM/resolve/main/DATA_PROVENANCE.md
- Command line
-
hf download hf://ToTo-40417/EXLLM/DATA_PROVENANCE.md
-
curl -L -o DATA_PROVENANCE.md https://huggingface.co/ToTo-40417/EXLLM/resolve/main/DATA_PROVENANCE.md
Data and Weight Provenance
Model lineage
EXLLM-0.005B-Instruct is a project-original model with 0.005377824B unique trainable parameters.
- The EXLLM project originates from random initialization.
- The current 0.005B release was expanded from the published project-owned
weights/EXLLM-v1.0.0.safetensorscheckpoint by transplanting dimension-compatible parameters. - Additional staged training used only project-generated Japanese instruction/response data.
- Corrective datasets were added in response to regression failures.
- The final fp32 checkpoint was exported to EXLLM8 and EXQ12 deployment formats.
No external pretrained checkpoint was used. The model is not a distillation or quantization of a third-party LLM.
The direct parent hash and every 5M training stage are recorded in
training/release-5m-stages.json. The source
repository contains the complete datasets and tools/replay_5m.py. EXQ12 is
deterministically generated from EXLLM8 by tools/export_exq12.py.
Training data
The project contains 0.086408M JSONL records across its published staged files. This number is the sum of file records, not a deduplicated training-example count. The files include overlapping base, validation, robustness, recovery, balance, consistency, and final corrective stages.
The data covers:
- short Japanese greetings and UI interaction;
- EXLLM and developer identity responses;
- short definitions of common terms;
- offline and current-information limitations;
- unknown-input fallback behavior;
- mixed Unicode and UTF-8 byte fallback;
- deterministic calculator routing;
- regression cases discovered during development.
No external public text dataset was imported into the published project corpus.
Published checkpoint metadata
| Field | Value |
|---|---|
| Parameters | 0.005377824B |
| Global step | 1,810 |
| Initialization | Published project-owned v1.0 checkpoint transplantation |
| Final corrective source | EXLLM-v1.1-5m-release2.pt |
| Final corrective data | v1_1_release_fix.jsonl |
| Final corrective steps | 100 |
| Final corrective learning rate | 6e-6 |
| Final corrective loss | 0.03875 → 0.02420 |
Licensing
The released model weights, project-generated data, and reference code are distributed under Apache License 2.0 unless otherwise stated.
Scope
The corpus is intentionally narrow. Its presence in the training data does not establish comprehensive knowledge, factual authority, or currency. See the model card for intended use and limitations.