Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
journal de demarrage
Browse files- etat/6tqc2otcxvsry3.log +17 -2
etat/6tqc2otcxvsry3.log
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
=== bootstrap v82-humming-extras-cu13 | 15:
|
| 2 |
[VL] 15:45:50 bootstrap v82-humming-extras-cu13
|
| 3 |
[VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
|
| 4 |
[VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
|
|
@@ -18,7 +18,7 @@
|
|
| 18 |
[VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
|
| 19 |
|
| 20 |
=== nvidia-smi ===
|
| 21 |
-
|
| 22 |
|
| 23 |
=== vllm.log (fin) ===
|
| 24 |
(APIServer pid=872) INFO 08-23 15:47:06 [api_utils.py:345]
|
|
@@ -147,3 +147,18 @@
|
|
| 147 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [ssu_dispatch.py:310] Using flashinfer Mamba SSU backend.
|
| 148 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [flashinfer.py:824] FlashInfer resolved query dtypes: prefill=torch.bfloat16, decode=torch.bfloat16, decode_backend=flashinfer-native, kv_cache_dtype=torch.float8_e4m3fn, arch=sm80
|
| 149 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [gpu_model_runner.py:6681] Profiling CUDA graph memory: PIECEWISE=19 (largest=128), FULL=11 (largest=64)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
=== bootstrap v82-humming-extras-cu13 | 15:56:19 UTC ===
|
| 2 |
[VL] 15:45:50 bootstrap v82-humming-extras-cu13
|
| 3 |
[VL] 15:45:51 pilote 550.54.15, CUDA runtime 13.0
|
| 4 |
[VL] 15:45:51 compat CUDA 13 active (LD_LIBRARY_PATH)
|
|
|
|
| 18 |
[VL] 15:46:49 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/6tqc2otcxvsry3.log
|
| 19 |
|
| 20 |
=== nvidia-smi ===
|
| 21 |
+
75441 MiB, 81920 MiB
|
| 22 |
|
| 23 |
=== vllm.log (fin) ===
|
| 24 |
(APIServer pid=872) INFO 08-23 15:47:06 [api_utils.py:345]
|
|
|
|
| 147 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [ssu_dispatch.py:310] Using flashinfer Mamba SSU backend.
|
| 148 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [flashinfer.py:824] FlashInfer resolved query dtypes: prefill=torch.bfloat16, decode=torch.bfloat16, decode_backend=flashinfer-native, kv_cache_dtype=torch.float8_e4m3fn, arch=sm80
|
| 149 |
(EngineCore pid=1604) INFO 08-23 15:53:52 [gpu_model_runner.py:6681] Profiling CUDA graph memory: PIECEWISE=19 (largest=128), FULL=11 (largest=64)
|
| 150 |
+
(EngineCore pid=1604) INFO 08-23 15:55:59 [gpu_model_runner.py:6806] Estimated CUDA graph memory: 0.77 GiB total
|
| 151 |
+
(EngineCore pid=1604) INFO 08-23 15:55:59 [gpu_worker.py:563] Available KV cache memory: 53.81 GiB
|
| 152 |
+
(EngineCore pid=1604) INFO 08-23 15:55:59 [gpu_worker.py:578] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9300 is equivalent to --gpu-memory-utilization=0.9203 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.9397. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0.
|
| 153 |
+
(EngineCore pid=1604) WARNING 08-23 15:55:59 [kv_cache_utils.py:1261] Add 1 padding layers, may waste at most 4.35% KV cache memory
|
| 154 |
+
(EngineCore pid=1604) INFO 08-23 15:55:59 [kv_cache_utils.py:2235] GPU KV cache size: 18,495,541 tokens
|
| 155 |
+
(EngineCore pid=1604) INFO 08-23 15:55:59 [kv_cache_utils.py:2236] Maximum concurrency for 1,048,576 tokens per request: 17.64x
|
| 156 |
+
(EngineCore pid=1604) INFO 08-23 15:56:00 [gpu_model_runner.py:6867] Rank 0: Torch profiler disabled for CUDA graph capture
|
| 157 |
+
(EngineCore pid=1604)
|
| 158 |
+
(EngineCore pid=1604)
|
| 159 |
+
(EngineCore pid=1604) INFO 08-23 15:56:07 [gpu_model_runner.py:6913] Graph capturing finished in 7 secs, took 0.76 GiB
|
| 160 |
+
(EngineCore pid=1604) INFO 08-23 15:56:07 [gpu_worker.py:726] CUDA graph pool memory: 0.76 GiB (actual), 0.77 GiB (estimated), difference: 0.01 GiB (1.8%).
|
| 161 |
+
(EngineCore pid=1604) INFO 08-23 15:56:07 [gpu_worker.py:789] Free memory on device (78.72/79.14 GiB) on startup. Desired GPU memory utilization is (0.93, 73.6 GiB). Actual usage is 18.65 GiB for consumed memory (weights + non-torch), 1.14 GiB for peak activation, and 0.76 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory=56806828626` (52.91 GiB) to fit into requested memory, or `--kv-cache-memory=62309380608` (58.03 GiB) to fully utilize gpu memory. Current kv cache memory in use is 53.81 GiB.
|
| 162 |
+
(EngineCore pid=1604) INFO 08-23 15:56:10 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 163 |
+
(EngineCore pid=1604) INFO 08-23 15:56:11 [core.py:348] init engine (profile, create kv cache, warmup model) took 273.96 s (compilation: 23.71 s)
|
| 164 |
+
(EngineCore pid=1604) [fastokens] patch_transformers: successfully patched transformers v5.15.1
|