Instructions to use AJKADZ/PHI_CODER with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AJKADZ/PHI_CODER with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: llama cli -hf AJKADZ/PHI_CODER:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: llama cli -hf AJKADZ/PHI_CODER:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AJKADZ/PHI_CODER:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AJKADZ/PHI_CODER:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AJKADZ/PHI_CODER:Q4_K_M
Use Docker
docker model run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use AJKADZ/PHI_CODER with Ollama:
ollama run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use AJKADZ/PHI_CODER with Docker Model Runner:
docker model run hf.co/AJKADZ/PHI_CODER:Q4_K_M
- Lemonade
How to use AJKADZ/PHI_CODER with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AJKADZ/PHI_CODER:Q4_K_M
Run and chat with the model
lemonade run user.PHI_CODER-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Download phi-coder-hf/llama.cpp/tools/server/tests/unit/test_rerank.py from AJKADZ/PHI_CODER: direct link, hf CLI and curl.
- Browser
- Download file 3.82 kB
-
https://huggingface.co/AJKADZ/PHI_CODER/resolve/main/phi-coder-hf/llama.cpp/tools/server/tests/unit/test_rerank.py
- Command line
-
hf download hf://AJKADZ/PHI_CODER/phi-coder-hf/llama.cpp/tools/server/tests/unit/test_rerank.py
-
curl -L -o test_rerank.py https://huggingface.co/AJKADZ/PHI_CODER/resolve/main/phi-coder-hf/llama.cpp/tools/server/tests/unit/test_rerank.py
3.82 kB
| import pytest | |
| from utils import * | |
| server = ServerPreset.jina_reranker_tiny() | |
| def create_server(): | |
| global server | |
| server = ServerPreset.jina_reranker_tiny() | |
| TEST_DOCUMENTS = [ | |
| "A machine is a physical system that uses power to apply forces and control movement to perform an action. The term is commonly applied to artificial devices, such as those employing engines or motors, but also to natural biological macromolecules, such as molecular machines.", | |
| "Learning is the process of acquiring new understanding, knowledge, behaviors, skills, values, attitudes, and preferences. The ability to learn is possessed by humans, non-human animals, and some machines; there is also evidence for some kind of learning in certain plants.", | |
| "Machine learning is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generalize to unseen data, and thus perform tasks without explicit instructions.", | |
| "Paris, capitale de la France, est une grande ville européenne et un centre mondial de l'art, de la mode, de la gastronomie et de la culture. Son paysage urbain du XIXe siècle est traversé par de larges boulevards et la Seine." | |
| ] | |
| def test_rerank(): | |
| global server | |
| server.start() | |
| res = server.make_request("POST", "/rerank", data={ | |
| "query": "Machine learning is", | |
| "documents": TEST_DOCUMENTS, | |
| }) | |
| assert res.status_code == 200 | |
| assert len(res.body["results"]) == 4 | |
| most_relevant = res.body["results"][0] | |
| least_relevant = res.body["results"][0] | |
| for doc in res.body["results"]: | |
| if doc["relevance_score"] > most_relevant["relevance_score"]: | |
| most_relevant = doc | |
| if doc["relevance_score"] < least_relevant["relevance_score"]: | |
| least_relevant = doc | |
| assert most_relevant["relevance_score"] > least_relevant["relevance_score"] | |
| assert most_relevant["index"] == 2 | |
| assert least_relevant["index"] == 3 | |
| def test_rerank_tei_format(): | |
| global server | |
| server.start() | |
| res = server.make_request("POST", "/rerank", data={ | |
| "query": "Machine learning is", | |
| "texts": TEST_DOCUMENTS, | |
| }) | |
| assert res.status_code == 200 | |
| assert len(res.body) == 4 | |
| most_relevant = res.body[0] | |
| least_relevant = res.body[0] | |
| for doc in res.body: | |
| if doc["score"] > most_relevant["score"]: | |
| most_relevant = doc | |
| if doc["score"] < least_relevant["score"]: | |
| least_relevant = doc | |
| assert most_relevant["score"] > least_relevant["score"] | |
| assert most_relevant["index"] == 2 | |
| assert least_relevant["index"] == 3 | |
| def test_invalid_rerank_req(documents): | |
| global server | |
| server.start() | |
| res = server.make_request("POST", "/rerank", data={ | |
| "query": "Machine learning is", | |
| "documents": documents, | |
| }) | |
| assert res.status_code == 400 | |
| assert "error" in res.body | |
| def test_rerank_usage(query, doc1, doc2, n_tokens): | |
| global server | |
| server.start() | |
| res = server.make_request("POST", "/rerank", data={ | |
| "query": query, | |
| "documents": [ | |
| doc1, | |
| doc2, | |
| ] | |
| }) | |
| assert res.status_code == 200 | |
| assert res.body['usage']['prompt_tokens'] == res.body['usage']['total_tokens'] | |
| assert res.body['usage']['prompt_tokens'] == n_tokens | |