Buckets:
| [local-llama] staging: 4.49 GiB bucket -> local SSD | |
| [local-llama] staged: 4.49 GiB in 17.0s | |
| [local-llama] loading: llama.cpp · threads=2 · ctx=65536 · KV=q8_0/q8_0 | |
| [local-llama] /opt/llama/bin/llama-server --model /tmp/local-agent-models/Ling-3.0-tiny-Q4_K_M.gguf --host 127.0.0.1 --port 43126 --ctx-size 65536 --threads 2 --threads-batch 2 --parallel 1 --cont-batching --flash-attn auto --jinja --cache-type-k q8_0 --cache-type-v q8_0 --gpu-layers 0 | |
| warning: no usable GPU found, --gpu-layers option will be ignored | |
| warning: one possible reason is that llama.cpp was compiled without GPU support | |
| warning: consult docs/build.md for compilation instructions | |
| 0.00.004.092 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) | |
| 0.00.004.920 W srv llama_server: ----------------- | |
| 0.00.004.927 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set | |
| 0.00.004.927 W srv llama_server: this can be a security risk (cross-origin attacks) | |
| 0.00.004.928 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 | |
| 0.00.004.928 W srv llama_server: ----------------- | |
| 0.00.006.263 I srv load_model: loading model '/tmp/local-agent-models/Ling-3.0-tiny-Q4_K_M.gguf' | |
| 0.00.929.373 W load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect | |
| 0.05.230.756 I cmn init: llama threadpool init, n_threads = 2 | |
| 0.05.596.580 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 65536, kv_unified = 'false' | |
| 0.05.620.286 W srv init: chat template supports preserving reasoning, it is enabled by default (may use more tokens, disable via --no-reasoning-preserve) | |
| 0.05.620.306 I srv llama_server: model loaded | |
| 0.05.620.311 I srv llama_server: listening on http://127.0.0.1:43126 | |
| [local-llama] ready | |
| 3.15.413.290 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 | |
| 3.15.418.737 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 | |
| 4.44.798.962 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 2048, progress = 0.10, t = 89.38 s / 22.91 tokens per second | |
| 6.28.101.309 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 4096, progress = 0.21, t = 192.68 s / 21.26 tokens per second | |
| 8.59.015.578 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 6144, progress = 0.31, t = 343.60 s / 17.88 tokens per second | |
| 11.43.219.502 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 8192, progress = 0.41, t = 507.80 s / 16.13 tokens per second | |
| 15.09.870.172 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 10240, progress = 0.52, t = 714.45 s / 14.33 tokens per second | |
Xet Storage Details
- Size:
- 2.72 kB
- Xet hash:
- 073c99a47a51ce83702d5282324e1cb716c86eb0296a8011c41d9184d5ddd2fe
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.