Instructions to use ariel-pillar/phi-4_function_calling with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ariel-pillar/phi-4_function_calling with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ariel-pillar/phi-4_function_calling:Q4_K_M # Run inference directly in the terminal: llama cli -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ariel-pillar/phi-4_function_calling:Q4_K_M # Run inference directly in the terminal: llama cli -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ariel-pillar/phi-4_function_calling:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ariel-pillar/phi-4_function_calling:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Use Docker
docker model run hf.co/ariel-pillar/phi-4_function_calling:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use ariel-pillar/phi-4_function_calling with Ollama:
ollama run hf.co/ariel-pillar/phi-4_function_calling:Q4_K_M
- Unsloth Studio
How to use ariel-pillar/phi-4_function_calling with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ariel-pillar/phi-4_function_calling to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ariel-pillar/phi-4_function_calling to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ariel-pillar/phi-4_function_calling to start chatting
- Pi
How to use ariel-pillar/phi-4_function_calling with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ariel-pillar/phi-4_function_calling:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use ariel-pillar/phi-4_function_calling with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ariel-pillar/phi-4_function_calling:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use ariel-pillar/phi-4_function_calling with Docker Model Runner:
docker model run hf.co/ariel-pillar/phi-4_function_calling:Q4_K_M
- Lemonade
How to use ariel-pillar/phi-4_function_calling with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ariel-pillar/phi-4_function_calling:Q4_K_M
Run and chat with the model
lemonade run user.phi-4_function_calling-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ariel-pillar/phi-4_function_calling with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ariel-pillar/phi-4_function_calling:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ariel-pillar/phi-4_function_calling:Q4_K_M
Run Hermes
hermes
- Atomic Chat
File size: 4,070 Bytes
e29b963 8ae50d9 4d8adc1 e29b963 8ae50d9 f578d74 8ae50d9 f578d74 44d566d 8ae50d9 f3cce8b 8ae50d9 44d566d 8ae50d9 44d566d 8ae50d9 44d566d 8ae50d9 44d566d 8ae50d9 48cebec 8ae50d9 44d566d 8ae50d9 44d566d 8ae50d9 48cebec 8ae50d9 48cebec e29b963 44d566d e29b963 44d566d e29b963 48cebec e29b963 44d566d 8ae50d9 e29b963 8ae50d9 2e7341f 8ae50d9 e29b963 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
base_model:
- microsoft/Phi-4-mini-instruct
---
# Phi-4-mini-instruct with llama-server (Tool-Enhanced Version)
NOTE: THIS IS A POC FOR A SUPPLY CHAIN ATTACK LEVERAGING POISONED CHAT TEMPLATES. FOR FULL BLOG/CONTEXT, PLEASE REVIEW: https://www.pillar.security/blog/llm-backdoors-at-the-inference-level-the-threat-of-poisoned-templates
This repository contains instructions for running a modified version of the Phi-4-mini-instruct model using llama-server. This version has been enhanced to support tool usage, allowing the model to interact with external tools and APIs through a ChatGPT-compatible interface.
## Model Capabilities
This modified version of Phi-4-mini-instruct includes:
- Full support for tool usage and function calling
- Custom chat template optimized for tool interactions
- Ability to process and respond to tool outputs
- ChatGPT-compatible API interface
## Prerequisites
- [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) installed with server support
- The Phi-4-mini-instruct model in GGUF format
## Installation
1. Install llama-cpp-python with server support:
```bash
pip install llama-cpp-python[server]
```
2. Ensure your model file is in the correct location:
```bash
models/Phi-4-mini-instruct-Q4_K_M-function_calling.gguf
```
## Running the Server
Start the llama-server with the following command:
```bash
llama-server \
--model models/Phi-4-mini-instruct-Q4_K_M-function_calling.gguf \
--port 8080 \
--jinja
```
This will start the server with:
- The model loaded in memory
- Server running on port 8082
- Verbose logging enabled
- Jinja template to support tool use
## Testing the API
You can test the server using curl commands. Here are some examples:
### Example 1: Using Tools
```bash
curl http://localhost:8080/v1/chat/completions -d '{
"model": "phi-4-mini-instruct-with-tools",
"tools": [
{
"type":"function",
"function":{
"name":"python",
"description":"Runs code in an ipython interpreter and returns the result of the execution after 60 seconds.",
"parameters":{
"type":"object",
"properties":{
"code":{
"type":"string",
"description":"The code to run in the ipython interpreter."
}
},
"required":["code"]
}
}
}
],
"messages": [
{
"role": "user",
"content": "Print a hello world message with python."
}
]
}'
```
### Example 2: Tell a Joke
```bash
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "phi-4-mini-instruct-with-tools",
"messages": [
{"role":"system","content":"You are a helpful clown instruction assistant"},
{"role":"user","content":"tell me a funny joke"}
]
}'
```
### Example 3: Generate HTML Hello World
```bash
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "phi-4-mini-instruct-with-tools",
"messages": [
{"role":"system","content":"You are a helpful coding assistant"},
{"role":"user","content":"give me an html hello world document"}
]
}'
```
## API Endpoints
The server provides a ChatGPT-compatible API with the following main endpoints:
- `/v1/chat/completions` - For chat completions
- `/v1/completions` - For text completions
- `/v1/models` - To list available models
## Notes
- The server uses the same API format as OpenAI's ChatGPT API, making it compatible with many existing tools and libraries
- The `--jinja` flag enables proper chat template formatting for the model, which is essential for tool usage
## Troubleshooting
If you encounter issues:
1. Ensure the model file exists in the specified path
2. Check that port 8080 is not in use by another application
3. Verify that llama-cpp-python is installed with server support
## License
Please ensure you comply with the model's license terms when using it.
|