kat0012's picture
Parakon runtime: llama.cpp fork with fused 1-bit CUDA/Metal/CPU kernels (part 2)
2b12537 verified
Raw History Blame Contribute Delete
185 Bytes
---
base_model:
- {base_model}
---
# {model_name} GGUF
Recommended way to run this model:
```sh
llama-server -hf {namespace}/{model_name}-GGUF
```
Then, access http://localhost:8080