Instructions to use Mike0021/Ling-3.0-tiny-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Mike0021/Ling-3.0-tiny-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Mike0021/Ling-3.0-tiny-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mike0021/Ling-3.0-tiny-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mike0021/Ling-3.0-tiny-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Ollama
How to use Mike0021/Ling-3.0-tiny-GGUF with Ollama:
ollama run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Mike0021/Ling-3.0-tiny-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Mike0021/Ling-3.0-tiny-GGUF with Docker Model Runner:
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Lemonade
How to use Mike0021/Ling-3.0-tiny-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ling-3.0-tiny-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Mike0021/Ling-3.0-tiny-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Mike0021/Ling-3.0-tiny-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download validation/hf-reference.json from Mike0021/Ling-3.0-tiny-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 5.44 kB
-
https://huggingface.co/Mike0021/Ling-3.0-tiny-GGUF/resolve/main/validation/hf-reference.json
- Command line
-
hf download hf://Mike0021/Ling-3.0-tiny-GGUF/validation/hf-reference.json
-
curl -L -o hf-reference.json https://huggingface.co/Mike0021/Ling-3.0-tiny-GGUF/resolve/main/validation/hf-reference.json
5.44 kB
| { | |
| "environment": { | |
| "cuda": "12.8", | |
| "dtype": "torch.bfloat16", | |
| "gpu": "NVIDIA RTX PRO 4500 Blackwell", | |
| "greedy": true, | |
| "python": "3.12.3", | |
| "seed": 1, | |
| "tf32": false, | |
| "torch": "2.8.0+cu128", | |
| "transformers": "4.57.6" | |
| }, | |
| "generations": { | |
| "greedy_arithmetic_64": { | |
| "new_text": "17 × 23 = 391</think>391<|role_end|>", | |
| "new_token_ids": [ | |
| 16, | |
| 22, | |
| 14524, | |
| 220, | |
| 17, | |
| 18, | |
| 373, | |
| 220, | |
| 18, | |
| 24, | |
| 16, | |
| 156904, | |
| 18, | |
| 24, | |
| 16, | |
| 156895 | |
| ], | |
| "prompt": "<role>SYSTEM</role>detailed thinking on<|role_end|><role>HUMAN</role>Calculate 17 × 23 and output only the number.<|role_end|><role>ASSISTANT</role>\n<think>", | |
| "prompt_token_ids": [ | |
| 157151, | |
| 90827, | |
| 157152, | |
| 14136, | |
| 5381, | |
| 6350, | |
| 366, | |
| 156895, | |
| 157151, | |
| 39, | |
| 116171, | |
| 157152, | |
| 15085, | |
| 220, | |
| 16, | |
| 22, | |
| 14524, | |
| 220, | |
| 17, | |
| 18, | |
| 301, | |
| 4403, | |
| 1191, | |
| 268, | |
| 1788, | |
| 13, | |
| 156895, | |
| 157151, | |
| 8469, | |
| 7342, | |
| 5468, | |
| 157152, | |
| 198, | |
| 156903 | |
| ], | |
| "stopped_on_eos": true | |
| }, | |
| "greedy_capitals_12": { | |
| "new_text": " Tokyo. The capital of Italy is Rome. The capital of", | |
| "new_token_ids": [ | |
| 30017, | |
| 13, | |
| 468, | |
| 7706, | |
| 300, | |
| 15003, | |
| 341, | |
| 21924, | |
| 13, | |
| 468, | |
| 7706, | |
| 300 | |
| ], | |
| "prompt": "The capital of France is Paris. The capital of Germany is Berlin. The capital of Japan is", | |
| "prompt_token_ids": [ | |
| 678, | |
| 7706, | |
| 300, | |
| 11406, | |
| 341, | |
| 13997, | |
| 13, | |
| 468, | |
| 7706, | |
| 300, | |
| 11462, | |
| 341, | |
| 24752, | |
| 13, | |
| 468, | |
| 7706, | |
| 300, | |
| 7499, | |
| 341 | |
| ], | |
| "stopped_on_eos": false | |
| } | |
| }, | |
| "source": { | |
| "config_sha256": "9750d847957913f665a13c0b5a6537199e33c6f3ec970d9fcb55a0e5076d4012", | |
| "control_file_sha256": { | |
| "chat_template.jinja": "eb6226c94ae38058f875d159f86a206b3a165828c0e7d6bda664ae14667f798a", | |
| "config.json": "9750d847957913f665a13c0b5a6537199e33c6f3ec970d9fcb55a0e5076d4012", | |
| "configuration_bailing_moe_v3.py": "f2c048966aec8a2f042cfeb1351f74d51a28589b409c55baae7d24e841c1f6c4", | |
| "generation_config.json": "64752c5973a55faf4cfc02604c7587c38b090f00013f91191f52362dcc79a4a8", | |
| "model.safetensors.index.json": "84ef9fe8ef967eeb0545deb1d23c0ce54e86e6b18fa943a7903d7354a79f9cf9", | |
| "modeling_bailing_moe_v3.py": "c2509bf7ac580c262e2581d34d6403aa21682d2e10beb9ad85ad8820a7e33a40", | |
| "special_tokens_map.json": "69b63b9f81044ead642d16a5fdc01bcc737dc1183746485c8397ab14d3126614", | |
| "tokenizer.json": "40fb9d7d7795b8bd305aeff39ce9963f3f450915b9553f2938e009be9a1fed60", | |
| "tokenizer_config.json": "2456b0372956cd3e82f17e33372148b115a94970bfd4878ba5e7e60cd3204f74" | |
| }, | |
| "path_basename": "Ling-3.0-tiny", | |
| "revision": "a2ee06c0f2de5b171701aee7f73f70a1da75483b", | |
| "tokenizer_sha256": "40fb9d7d7795b8bd305aeff39ce9963f3f450915b9553f2938e009be9a1fed60", | |
| "verified_weight_shards": 32, | |
| "weight_manifest_sha256": "d8a7cf059fd4b4f2fd7f7d0b1118417be02f7b39a5c428461c027f5d71faff6d" | |
| }, | |
| "tokenizer_cases": [ | |
| { | |
| "text": "The capital of France is Paris.", | |
| "token_ids": [ | |
| 678, | |
| 7706, | |
| 300, | |
| 11406, | |
| 341, | |
| 13997, | |
| 13 | |
| ] | |
| }, | |
| { | |
| "text": "你好,世界!这是 Ling-3.0-tiny。", | |
| "token_ids": [ | |
| 34355, | |
| 44291, | |
| 89404, | |
| 57917, | |
| 12, | |
| 18, | |
| 13, | |
| 15, | |
| 2162, | |
| 7474, | |
| 311 | |
| ] | |
| }, | |
| { | |
| "text": "def fibonacci(n: int) -> int:\n return n if n < 2 else fibonacci(n-1) + fibonacci(n-2)", | |
| "token_ids": [ | |
| 1413, | |
| 13871, | |
| 66472, | |
| 3733, | |
| 25, | |
| 616, | |
| 8, | |
| 4267, | |
| 616, | |
| 25, | |
| 198, | |
| 305, | |
| 944, | |
| 320, | |
| 624, | |
| 320, | |
| 797, | |
| 220, | |
| 17, | |
| 1942, | |
| 13871, | |
| 66472, | |
| 3733, | |
| 12, | |
| 16, | |
| 8, | |
| 781, | |
| 13871, | |
| 66472, | |
| 3733, | |
| 12, | |
| 17, | |
| 8 | |
| ] | |
| }, | |
| { | |
| "text": " leading\twhitespace\n\nand trailing ", | |
| "token_ids": [ | |
| 220, | |
| 6135, | |
| 197, | |
| 103528, | |
| 198, | |
| 198, | |
| 457, | |
| 43453, | |
| 256 | |
| ] | |
| }, | |
| { | |
| "text": "emoji: 🧠🚀 café naïve العربية हिन्दी", | |
| "token_ids": [ | |
| 86471, | |
| 25, | |
| 8811, | |
| 100, | |
| 254, | |
| 114055, | |
| 222, | |
| 67656, | |
| 105516, | |
| 23428, | |
| 32923, | |
| 131690, | |
| 82550, | |
| 47430, | |
| 101, | |
| 31741, | |
| 99, | |
| 49458 | |
| ] | |
| }, | |
| { | |
| "text": "<role>SYSTEM</role>detailed thinking off<|role_end|><role>HUMAN</role>Hello<|role_end|>", | |
| "token_ids": [ | |
| 157151, | |
| 90827, | |
| 157152, | |
| 14136, | |
| 5381, | |
| 6350, | |
| 928, | |
| 156895, | |
| 157151, | |
| 39, | |
| 116171, | |
| 157152, | |
| 14455, | |
| 156895 | |
| ] | |
| } | |
| ] | |
| } | |