Instructions to use Executespec/ganesh-mini-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Executespec/ganesh-mini-1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Executespec/ganesh-mini-1.0") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Executespec/ganesh-mini-1.0") model = AutoModelForMultimodalLM.from_pretrained("Executespec/ganesh-mini-1.0", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Executespec/ganesh-mini-1.0 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Executespec/ganesh-mini-1.0:F16 # Run inference directly in the terminal: llama cli -hf Executespec/ganesh-mini-1.0:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Executespec/ganesh-mini-1.0:F16 # Run inference directly in the terminal: llama cli -hf Executespec/ganesh-mini-1.0:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Executespec/ganesh-mini-1.0:F16 # Run inference directly in the terminal: ./llama-cli -hf Executespec/ganesh-mini-1.0:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Executespec/ganesh-mini-1.0:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Executespec/ganesh-mini-1.0:F16
Use Docker
docker model run hf.co/Executespec/ganesh-mini-1.0:F16
- LM Studio
- Jan
- vLLM
How to use Executespec/ganesh-mini-1.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Executespec/ganesh-mini-1.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Executespec/ganesh-mini-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Executespec/ganesh-mini-1.0:F16
- SGLang
How to use Executespec/ganesh-mini-1.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Executespec/ganesh-mini-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Executespec/ganesh-mini-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Executespec/ganesh-mini-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Executespec/ganesh-mini-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Executespec/ganesh-mini-1.0 with Ollama:
ollama run hf.co/Executespec/ganesh-mini-1.0:F16
- Unsloth Desktop
- Pi
How to use Executespec/ganesh-mini-1.0 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Executespec/ganesh-mini-1.0:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Executespec/ganesh-mini-1.0:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Executespec/ganesh-mini-1.0 with Docker Model Runner:
docker model run hf.co/Executespec/ganesh-mini-1.0:F16
- Lemonade
How to use Executespec/ganesh-mini-1.0 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Executespec/ganesh-mini-1.0:F16
Run and chat with the model
lemonade run user.ganesh-mini-1.0-F16
List all available models
lemonade list
- Hermes Agent
How to use Executespec/ganesh-mini-1.0 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Executespec/ganesh-mini-1.0:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Executespec/ganesh-mini-1.0:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Executespec/ganesh-mini-1.0 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Executespec/ganesh-mini-1.0:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Executespec/ganesh-mini-1.0:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ganesh Mini
Ganesh Mini is a text-only coding fine-tune of
Qwen/Qwen3.5-2B, pinned to base
revision 15852e8c16360a2fea060d615a32b45270f8a8fc. It was trained
jointly for Go, Java, JavaScript, TypeScript, Python, and Rust. The intended
serving mode is non-thinking text generation with an 8,192-token context.
Vision behavior has not been evaluated.
Files
- Repository root: merged BF16 Transformers checkpoint and tokenizer.
gguf/ganesh-mini-1.0-f16.gguf: F16 text target.gguf/ganesh-mini-1.0-q8_0.gguf: Q8_0 text target.adapter/: original rank-8 LoRA adapter. Its configuration retains the training container's/modelbase path; for PEFT use, load the exact Qwen base revision explicitly and then apply this adapter.
The GGUF target exports exclude the unused speculative MTP head. Both GGUF files and the merged Transformers checkpoint loaded and generated a short text-only Python-function smoke response locally. This is an export smoke, not multi-language equivalence or laptop RAM certification.
Public assessment
We ran a local, paired 30-task-per-language public diagnostic using pinned MultiPL-E HumanEval translations for Go, Java, JavaScript, TypeScript and Rust, and the pinned HumanEvalPlus file from the MultiPL-E repository for Python. Both models used Q8_0 GGUF, the same non-thinking text-only serving settings, and the published benchmark verifier in a network-disabled container.
| Language | Ganesh Mini | Qwen upstream | Paired wins / losses |
|---|---|---|---|
| Go | 9/30 | 8/30 | 1 / 0 |
| Java | 12/30 | 13/30 | 0 / 1 |
| JavaScript | 12/30 | 12/30 | 0 / 0 |
| TypeScript | 13/30 | 14/30 | 0 / 1 |
| Python | 18/30 | 18/30 | 0 / 0 |
| Rust | 6/30 | 6/30 | 1 / 1 |
| Total | 70/180 | 71/180 | 2 / 3 |
The fixed sample, exact revisions, hashes, settings, sandbox limits and Rust stop-token correction are documented in the public benchmark method. The paired result report contains task IDs, statuses and evidence hashes, but no benchmark prompts, tests or raw model completions. The initial Rust assembly error is retained locally and excluded; the reported Rust result re-executes the same saved outputs after applying the published stop token to both models.
This is a small locally run public slice, not a full official benchmark leaderboard result or evidence of general improvement. Public tasks may have appeared in upstream pretraining. No private evaluation scores are published.
Full public task diagnostic
We subsequently ran the full pinned public task selection with the same candidate and upstream Q8_0 serving contract and the network-disabled verifier. This run includes all 952 tasks across the six lanes; it does not include the private development or decision sets.
| Language | Ganesh Mini | Qwen upstream | Paired wins / losses |
|---|---|---|---|
| Go | 45/154 | 43/154 | 2 / 0 |
| Java | 51/158 | 52/158 | 0 / 1 |
| JavaScript | 64/161 | 64/161 | 1 / 1 |
| TypeScript | 59/159 | 59/159 | 2 / 2 |
| Python | 84/164 | 86/164 | 0 / 2 |
| Rust | 35/156 | 35/156 | 3 / 3 |
| Total | 338/952 | 339/952 | 8 / 9 |
The full-run method records the pinned inputs, postprocessing, and verifier limits. The sealed paired report records per-task statuses and evidence hashes without prompts, tests, raw outputs, or private assessment. The result is one pass below upstream overall; it does not support a general improvement claim. This is a local public diagnostic, not an official leaderboard submission. The earlier 30-task sample is a subset of these public tasks and must not be added to this total.
Training and limitations
The selected checkpoint completed one unique epoch over an admitted 900-row
package (720 capability examples and 180 broad replay examples), with
651,639 input tokens and 113 optimizer steps. Training used BF16
attention-only rank-8 LoRA, response-only loss, and a 2,944-token training
sequence limit; the selected learning rate was 1e-5. Adapter save/reload
equivalence passed. These figures describe the completed training run, not a
claim of coding superiority.
The 8 GB Q8 and 16 GB F16 laptop profiles have not been certified. Neither the GGUF files nor the merged BF16 checkpoint has completed the full six-lane export-equivalence evaluation. The model may produce incorrect, unsafe, or non-compiling code; execute its output only in a resource-limited, network-disabled sandbox. The original Qwen model license is Apache-2.0. No training data or private evaluation prompts/tests are included here.
Reproducibility identifiers
- Training run evidence seal:
21d78a7e339fb75dd89bdd9ada83a53bdbb10678bd7c3c9b09d182786b698537 - Adapter safetensors SHA-256:
03f8d725c9cad5ce63e524db45883eceb2ab610cf207e6153b004215ec798896 - GGUF F16 SHA-256:
03e6c13f3e75dc322a207d9f6dd96a3c97adcefcaea7f204e77cf8f8fb761d47 - GGUF Q8_0 SHA-256:
f4a3a394e4551211b4de0a58ad5cea1382e5b88ff450c996f1452d3c22388d89
These identifiers permit audit against the retained project evidence without publishing private tasks or raw development outputs.
- Downloads last month
- 7