dkudos commited on
Commit
1c616ab
·
verified ·
1 Parent(s): 5a15826

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +16 -5
README.md CHANGED
@@ -55,13 +55,20 @@ A 287M-parameter decoder-only causal language model (Llama-3 style architecture)
55
  | File | Description | Size |
56
  |---|---|---|
57
  | `model.safetensors` | Full bf16 PyTorch weights (HF format with `config.json`, `tokenizer.json`/`tokenizer_config.json`) | 548 MiB |
58
- | `checkpoint-4000-f16.gguf` | GGUF F16 (float16) - best quality for llama.cpp | 550 MiB |
59
- | `checkpoint-4000-Q8_0.gguf` | GGUF Q8_0 8-bit quantized - recommended for most uses (near-lossless, ~2x smaller) | 344 MiB |
60
  | `config.json` | Model config (transformers) | - |
61
  | `tokenizer.json` / `tokenizer_config.json` | BPE tokenizer (vocab 65,536) | - |
62
  | `train_log.log` | Full training log (steps, losses, LR) | - |
63
  | `full_val_eval.log` | Held-out full validation eval log | - |
64
 
 
 
 
 
 
 
 
65
  ## How to run
66
 
67
  ### HuggingFace transformers (PyTorch)
@@ -88,10 +95,14 @@ print(tok.decode(out[0]))
88
 
89
  Both GGUFs load directly in llama.cpp / llama-server with no external deps.
90
 
91
- Default (train context, 4096):
92
-
93
  ```bash
94
- llama-server -m dkudos/cinimod-devops/checkpoint-4000-Q8_0.gguf --port 8080
 
 
 
 
 
 
95
  ```
96
 
97
  256K context via linear RoPE scaling (trained at 4096):
 
55
  | File | Description | Size |
56
  |---|---|---|
57
  | `model.safetensors` | Full bf16 PyTorch weights (HF format with `config.json`, `tokenizer.json`/`tokenizer_config.json`) | 548 MiB |
58
+ | `checkpoint-4000-Q8_0.gguf` | **GGUF Q8_0 (8-bit)** — recommended for most uses (near-lossless, ~2x smaller). [Download](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf) · [View](https://huggingface.co/dkudos/cinimod-devops/tree/main/checkpoint-4000-Q8_0.gguf) | 344 MiB |
59
+ | `checkpoint-4000-f16.gguf` | **GGUF F16 (float16)** — best quality for llama.cpp. [Download](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf) · [View](https://huggingface.co/dkudos/cinimod-devops/tree/main/checkpoint-4000-f16.gguf) | 550 MiB |
60
  | `config.json` | Model config (transformers) | - |
61
  | `tokenizer.json` / `tokenizer_config.json` | BPE tokenizer (vocab 65,536) | - |
62
  | `train_log.log` | Full training log (steps, losses, LR) | - |
63
  | `full_val_eval.log` | Held-out full validation eval log | - |
64
 
65
+ ## GGUF (llama.cpp) — recommended
66
+
67
+ Two ready-to-serve GGUF files; either downloads standalone with no dependencies (no source code needed):
68
+
69
+ - [`checkpoint-4000-Q8_0.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf) — 8-bit quantized, ~344 MiB, recommended default
70
+ - [`checkpoint-4000-f16.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf) — float16, ~550 MiB, best fidelity
71
+
72
  ## How to run
73
 
74
  ### HuggingFace transformers (PyTorch)
 
95
 
96
  Both GGUFs load directly in llama.cpp / llama-server with no external deps.
97
 
 
 
98
  ```bash
99
+ # Q8_0 (default)
100
+ wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf
101
+ llama-server -m checkpoint-4000-Q8_0.gguf --port 8080
102
+
103
+ # or F16 for best fidelity
104
+ wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf
105
+ llama-server -m checkpoint-4000-f16.gguf --port 8080
106
  ```
107
 
108
  256K context via linear RoPE scaling (trained at 4096):