Instructions to use AJKADZ/PHI_CODER with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AJKADZ/PHI_CODER with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: llama cli -hf AJKADZ/PHI_CODER:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: llama cli -hf AJKADZ/PHI_CODER:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AJKADZ/PHI_CODER:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AJKADZ/PHI_CODER:Q4_K_M
Use Docker
docker model run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use AJKADZ/PHI_CODER with Ollama:
ollama run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use AJKADZ/PHI_CODER with Docker Model Runner:
docker model run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- Lemonade
How to use AJKADZ/PHI_CODER with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AJKADZ/PHI_CODER:Q4_K_M
Run and chat with the model
lemonade run user.PHI_CODER-Q4_K_M
List all available models
lemonade list
- Atomic Chat
File size: 1,373 Bytes
9b19fd8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 | #!/usr/bin/env bash
set -euo pipefail
this=$(realpath "$0"); readonly this
cd "$(dirname "$this")"
shellcheck "$this"
if (( $# != 1 && $# != 2 )); then
cat >&2 <<'EOF'
usage:
ci-run.sh <tmp_dir> [<cache_dir>]
This script wraps ci/run.sh:
* If <tmp_dir> is a ramdisk, you can reduce writes to your SSD. If <tmp_dir> is not a ramdisk, keep in mind that total writes will increase by the size of <cache_dir>.
(openllama_3b_v2: quantized models are about 30GB)
* Persistent model and data files are synced to and from <cache_dir>,
excluding generated .gguf files.
(openllama_3b_v2: persistent files are about 6.6GB)
* <cache_dir> defaults to ~/.cache/llama.cpp
EOF
exit 1
fi
cd .. # => llama.cpp repo root
tmp="$1"
mkdir -p "$tmp"
tmp=$(realpath "$tmp")
echo >&2 "Using tmp=$tmp"
cache="${2-$HOME/.cache/llama.cpp}"
mkdir -p "$cache"
cache=$(realpath "$cache")
echo >&2 "Using cache=$cache"
_sync() {
local from="$1"; shift
local to="$1"; shift
echo >&2 "Syncing from $from to $to"
mkdir -p "$from" "$to"
rsync -a "$from" "$to" --delete-during "$@"
}
_sync "$(realpath .)/" "$tmp/llama.cpp"
_sync "$cache/ci-mnt/models/" "$tmp/llama.cpp/ci-mnt/models/"
cd "$tmp/llama.cpp"
bash ci/run.sh ci-out ci-mnt
_sync 'ci-mnt/models/' "$cache/ci-mnt/models/" --exclude='*.gguf' -P
|