Text Generation
GGUF
Japanese
japanese
instruction-tuning
little-language-model
tiny-language-model
edge-ai
embedded-ai
ex-word
llama-cpp
lm-studio
custom-code
conversational
Instructions to use ToTo-40417/EXLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- Ollama
How to use ToTo-40417/EXLLM with Ollama:
ollama run hf.co/ToTo-40417/EXLLM:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM:F16
- Lemonade
How to use ToTo-40417/EXLLM with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM:F16
Run and chat with the model
lemonade run user.EXLLM-F16
List all available models
lemonade list
- Atomic Chat
|
Download DATA_PROVENANCE.md from ToTo-40417/EXLLM: direct link, hf CLI and curl.
- Browser
- Download file 2.56 kB
-
https://huggingface.co/ToTo-40417/EXLLM/resolve/main/DATA_PROVENANCE.md
- Command line
-
hf download hf://ToTo-40417/EXLLM/DATA_PROVENANCE.md
-
curl -L -o DATA_PROVENANCE.md https://huggingface.co/ToTo-40417/EXLLM/resolve/main/DATA_PROVENANCE.md
2.56 kB
| # Data and Weight Provenance | |
| ## Model lineage | |
| EXLLM-0.005B-Instruct is a project-original model with 0.005377824B unique trainable parameters. | |
| 1. The EXLLM project originates from random initialization. | |
| 2. The current 0.005B release was expanded from the published project-owned `weights/EXLLM-v1.0.0.safetensors` checkpoint by transplanting dimension-compatible parameters. | |
| 3. Additional staged training used only project-generated Japanese instruction/response data. | |
| 4. Corrective datasets were added in response to regression failures. | |
| 5. The final fp32 checkpoint was exported to EXLLM8 and EXQ12 deployment formats. | |
| No external pretrained checkpoint was used. The model is not a distillation or quantization of a third-party LLM. | |
| The direct parent hash and every 5M training stage are recorded in | |
| [`training/release-5m-stages.json`](training/release-5m-stages.json). The source | |
| repository contains the complete datasets and `tools/replay_5m.py`. EXQ12 is | |
| deterministically generated from EXLLM8 by `tools/export_exq12.py`. | |
| ## Training data | |
| The project contains 0.086408M JSONL records across its published staged files. This number is the sum of file records, not a deduplicated training-example count. The files include overlapping base, validation, robustness, recovery, balance, consistency, and final corrective stages. | |
| The data covers: | |
| - short Japanese greetings and UI interaction; | |
| - EXLLM and developer identity responses; | |
| - short definitions of common terms; | |
| - offline and current-information limitations; | |
| - unknown-input fallback behavior; | |
| - mixed Unicode and UTF-8 byte fallback; | |
| - deterministic calculator routing; | |
| - regression cases discovered during development. | |
| No external public text dataset was imported into the published project corpus. | |
| ## Published checkpoint metadata | |
| | Field | Value | | |
| |---|---| | |
| | Parameters | 0.005377824B | | |
| | Global step | 1,810 | | |
| | Initialization | Published project-owned v1.0 checkpoint transplantation | | |
| | Final corrective source | `EXLLM-v1.1-5m-release2.pt` | | |
| | Final corrective data | `v1_1_release_fix.jsonl` | | |
| | Final corrective steps | 100 | | |
| | Final corrective learning rate | 6e-6 | | |
| | Final corrective loss | 0.03875 → 0.02420 | | |
| ## Licensing | |
| The released model weights, project-generated data, and reference code are distributed under Apache License 2.0 unless otherwise stated. | |
| ## Scope | |
| The corpus is intentionally narrow. Its presence in the training data does not establish comprehensive knowledge, factual authority, or currency. See the model card for intended use and limitations. | |