Instructions to use SZLHOLDINGS/A11OY-MINI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SZLHOLDINGS/A11OY-MINI with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M # Run inference directly in the terminal: llama cli -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M # Run inference directly in the terminal: llama cli -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Use Docker
docker model run hf.co/SZLHOLDINGS/A11OY-MINI:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use SZLHOLDINGS/A11OY-MINI with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SZLHOLDINGS/A11OY-MINI" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SZLHOLDINGS/A11OY-MINI", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SZLHOLDINGS/A11OY-MINI:Q4_K_M
- Ollama
How to use SZLHOLDINGS/A11OY-MINI with Ollama:
ollama run hf.co/SZLHOLDINGS/A11OY-MINI:Q4_K_M
- Unsloth Desktop
- Pi
How to use SZLHOLDINGS/A11OY-MINI with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SZLHOLDINGS/A11OY-MINI:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SZLHOLDINGS/A11OY-MINI with Docker Model Runner:
docker model run hf.co/SZLHOLDINGS/A11OY-MINI:Q4_K_M
- Lemonade
How to use SZLHOLDINGS/A11OY-MINI with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SZLHOLDINGS/A11OY-MINI:Q4_K_M
Run and chat with the model
lemonade run user.A11OY-MINI-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use SZLHOLDINGS/A11OY-MINI with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SZLHOLDINGS/A11OY-MINI:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SZLHOLDINGS/A11OY-MINI with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SZLHOLDINGS/A11OY-MINI:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SZLHOLDINGS/A11OY-MINI:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
DERIVED GGUF RUNTIME · R2 NAMED-N RECORD · LEGACY FILES DEPRECATED · NOT PROMOTABLE · NOT AUTONOMOUS
A11OY-MINI
What this is
A derived GGUF runtime family. The R2 files are a quantized rebuild from a local Chaski-R2 adapter; the retained legacy F16 and Q4_K_M files derive from Chaski. It is not a new training run, not Chaski-5050, not a flagship release, and not an authorized A11oy production deployment.
Evidence and artifact identity
Reviewed Hub revision: c936dc749743c94586706345a0142c79094a581c. The values below are the Hub's LFS SHA-256 identities and sizes at that revision. R2 values match the canonical R2 conversion receipt; legacy values match conversion_receipt.json. Listing and receipt comparison do not constitute a new download-and-rehash of the GGUFs.
| Runtime artifact | Published file | Bytes | SHA-256 |
|---|---|---|---|
| R2 text Q4_K_M | a11oy-mini-r2-Q4_K_M.gguf | 541903328 | 6d42341c932a76e91b2c04a859a4248e7d2c77308f998a32771f43802f097b62 |
| R2 BF16 projector | a11oy-mini-r2-BF16-mmproj.gguf | 207346048 | 4855efe034435b9b3b289b2c07d09b263c088749d1f85c43c1cf4672bc7fcbf2 |
| Legacy F16 — deprecated | a11oy-mini-f16.gguf | 1557662240 | a5df00e4e3ca07f65a4b43aad4ef1505625952a3105e6dcc0dba87f2fa35fc57 |
| Legacy Q4_K_M — deprecated | a11oy-mini-q4_k_m.gguf | 541903392 | 06136ba385b2e052cf4cdb3dc8d333948e0b612bd15a541b314e170399c2faa6 |
The gguf_gate_receipt.json records JSON drafts 5/5 and adversarial refusals 6/6 on 2026-09-17, using Ollama, temperature 0, context 4096, with the named Chaski gate files chaski/gate/json_drafts.n5.jsonl and chaski/gate/adversarial_refusals.n6.jsonl. This is a small owner-run GGUF-level record for the named R2 Q4_K_M text artifact.
The gate receipt names the R2 artifact but does not itself contain the artifact SHA-256 or an immutable Hub revision. The separate canonical conversion receipt records the R2 file hashes and initially says gguf_scores=UNMEASURED; the later gate receipt records the measured run. The hub_put_receipt.json records the later authorized R2 upload decision. These documents provide scoped lineage and measured-record evidence; they do not make the evaluation receipt an independently verified cryptographic binding to public bytes.
New exact-byte text observation — 2026-09-24
HOLD remains in force. A new owner-local CUDA observation binds the R2 text
GGUF at Hub revision 0619dd65b92a135501af35b3c4e3b4e762be1d7d to local
SHA-256 6d42341c932a76e91b2c04a859a4248e7d2c77308f998a32771f43802f097b62,
verified before and after inference. The unaltered receipt
records the installed llama.cpp native CUDA runtime, RTX 5050 Laptop GPU,
temperature 0, seed 0, context 4096, immutable fixture identities, and all outputs.
The historical envelope checker reports 5/5 drafts and 6/6 refusal prefixes;
that is a format result, not semantic correctness.
Semantic review found four
unsupported claims: treating adapter files as an evaluation/gate closure,
presenting training loss as evaluation, inventing job state, and declaring
COMPLETED inside a refusal. The receipt's MEASURED_BOUNDED_PASS is preserved
with its original envelope-only meaning. Production disposition stays HOLD,
with publication_eligible=false, autonomy_eligible=false,
promotion_effect=NONE, PROPOSAL_ONLY, and UNSIGNED_HONEST.
The full evidence and exact reproduction sources are also retained in canonical GitHub source. This text-only observation uses known named fixtures. It does not retroactively bind the 2026-09-17 run, qualify the projector/vision path, establish held-out generalization, authorize a house-lab retarget, or establish a deployed service or autonomy.
What the evidence does and does not establish
The R2 5/5 + 6/6 result applies only to the named R2 Q4_K_M artifact. It does not transfer to the legacy GGUFs, the projector's visual performance, the Chaski baseline, Chaski-5050, all adapters, another quantization, or a new training run. It is not a public leaderboard result, a broad benchmark, a speed claim, or independent third-party qualification.
The legacy conversion_receipt.json retains evals=none-this-run and identifies Chaski parent revision 1c55df8652e9d0f7b84356b1e2d54849165ae884. Its byte measurements are not behavioral evaluations. Both legacy files remain published for provenance and are deprecated. The R2 parent has a real Chaski-R2 Hub page, but the canonical R2 conversion receipt identifies a local source-adapter path rather than an immutable parent Hub revision.
Release, autonomy, and deployment boundary
publication_eligible=false and autonomy_eligible=false remain unchanged. House-lab load is forbidden. This is not a lab retarget, a release qualification, deployment authorization, approval authority, or autonomous tool executor. Public artifact availability does not change those gates. No tokens-per-second or vision benchmark is claimed. Lambda uniqueness remains Conjecture 1: open and advisory, never a theorem.
Relationship and source of truth
Legacy parent: SZLHOLDINGS/chaski. R2 source family: SZLHOLDINGS/chaski-r2. Neither is SZLHOLDINGS/chaski-5050. This derived runtime does not replace the source adapters or the Khipu GGUF house-lab artifact.
Use gguf_gate_receipt.json, hub_put_receipt.json, the legacy conversion_receipt.json, and the canonical R2 conversion receipt as the evidence sources. The historical source documentation describes the original runtime family; later R2 evidence is scoped separately above.
- Downloads last month
- 323