Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
pont: effort -> thinking_token_budget (memes seuils que le patch natif)
Browse files- anthropic_proxy.py +8 -0
anthropic_proxy.py
CHANGED
|
@@ -329,6 +329,14 @@ def to_openai(body: dict) -> dict:
|
|
| 329 |
out["thinking_token_budget"] = max(1, int(th["budget_tokens"]))
|
| 330 |
except (TypeError, ValueError):
|
| 331 |
pass
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 332 |
|
| 333 |
if body.get("tools"):
|
| 334 |
out["tools"] = [{
|
|
|
|
| 329 |
out["thinking_token_budget"] = max(1, int(th["budget_tokens"]))
|
| 330 |
except (TypeError, ValueError):
|
| 331 |
pass
|
| 332 |
+
# Niveau d'effort (output_config.effort) -> plafond de jetons de raisonnement,
|
| 333 |
+
# memes seuils que le patch natif (vllm_anthropic_effort_patch.py). Le budget
|
| 334 |
+
# explicite ci-dessus gagne ; `disabled` aussi.
|
| 335 |
+
_eff = (body.get("output_config") or {}).get("effort") if isinstance(body.get("output_config"), dict) else None
|
| 336 |
+
_BUDGET = {"low": 1024, "medium": 4096, "high": 16384, "xhigh": 32768, "max": None}
|
| 337 |
+
if _eff in _BUDGET and "thinking_token_budget" not in out and "chat_template_kwargs" not in out:
|
| 338 |
+
if _BUDGET[_eff] is not None:
|
| 339 |
+
out["thinking_token_budget"] = _BUDGET[_eff]
|
| 340 |
|
| 341 |
if body.get("tools"):
|
| 342 |
out["tools"] = [{
|