Instructions to use aaaa47080/CryptoMind-Guard-0.6B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aaaa47080/CryptoMind-Guard-0.6B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0 # Run inference directly in the terminal: llama cli -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0 # Run inference directly in the terminal: llama cli -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf aaaa47080/CryptoMind-Guard-0.6B:Q8_0
Use Docker
docker model run hf.co/aaaa47080/CryptoMind-Guard-0.6B:Q8_0
- LM Studio
- Jan
- Ollama
How to use aaaa47080/CryptoMind-Guard-0.6B with Ollama:
ollama run hf.co/aaaa47080/CryptoMind-Guard-0.6B:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use aaaa47080/CryptoMind-Guard-0.6B with Docker Model Runner:
docker model run hf.co/aaaa47080/CryptoMind-Guard-0.6B:Q8_0
- Lemonade
How to use aaaa47080/CryptoMind-Guard-0.6B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aaaa47080/CryptoMind-Guard-0.6B:Q8_0
Run and chat with the model
lemonade run user.CryptoMind-Guard-0.6B-Q8_0
List all available models
lemonade list
- Atomic Chat
CryptoMind-Guard-0.6B
A small, CPU/iGPU-friendly pass / block classifier for community forum posts, built for the CryptoMind crypto & investing community. It flags scams/fraud, sexual content (including any sexual content involving minors), violence and threats, hate and harassment, self-harm, drugs/weapons, doxxing and other illegal activity — while letting ordinary investing discussion through (price talk, portfolio sharing, criticism, scam warnings and victims asking for help).
It is a LoRA fine-tune of Alibaba-AAIG/YuFeng-XGuard-Reason-0.6B
(Qwen3-0.6B) and keeps the base model's output format: the first answer token is a risk-category code
(sec = safe, ec = economic crimes / fraud, ma = minor abuse & exploitation, pc = pornographic content, …).
You only need one forward pass and the probabilities of 29 tokens — no text generation.
中文摘要:論壇貼文「該不該擋」的小型分類模型(0.6B),微調自 YuFeng-XGuard-Reason-0.6B。詐騙、色情(含任何涉及未成年的性內容)、 暴力威脅、仇恨騷擾、自殘、毒品槍械、肉搜等要擋;一般投資討論、防詐提醒、被害求助不擋。只讀第一個 token 的機率,CPU/內顯都跑得動。 主要支援繁體中文、簡體中文、英文、俄文。
About CryptoMind
CryptoMind is an AI research assistant for crypto and stocks, with a community forum where investors share analysis, plus a scam-report tracker. CryptoMind-Guard was built to keep that forum useful: check every post and comment before it goes live, stop scams and abuse, and stay out of the way of normal investing talk. 👉 Try it at getcryptomind.com.
Files
| File | Format | Size | Use with |
|---|---|---|---|
cryptomind-guard-0.6b-q8_0.gguf |
GGUF Q8_0 | 640 MB | llama.cpp (CPU, Vulkan / Metal / CUDA / ROCm). Recommended. |
onnx/model.onnx + onnx/model.onnx.data |
ONNX int8 (dynamic, per-channel) | 600 MB | onnxruntime (CPU). Graph outputs only the 29 risk-code logits. |
tokenizer.json |
HF tokenizers | 11 MB | Tokenizing the post (needed for both formats) |
readout.json |
JSON | 5 KB | The exact prompt contract: prefix/suffix token ids, the 29 label token ids, long-post rule |
The unquantized fp16 weights are not part of this release.
How it works
- Tokenize the post:
ids = tokenizer.encode(lead + post.strip())(leadis a single space, seereadout.json). - Long posts (more than 450 tokens) are scored as two pieces — the first 300 and the last 150 tokens — and the higher score wins (contact details and links are often at the end).
- Input =
prefix_ids + post_ids + suffix_idsfromreadout.json. This is token-for-token identical to the base model's chat template (apply_chat_template([{"role": "user", "content": post}], policy=None, reason_first=False)); the 299-token prefix is the same for every post, so llama.cpp's prompt cache makes repeated calls cheap. - Take the next-token distribution, keep only the 29
label_ids, renormalize, and use score = 1 − P(sec). The highest non-seccode is the most likely category.
Recommended policy: block at score ≥ 0.8; publish but send to human review at 0.5–0.8.
llama.cpp
llama-server -m cryptomind-guard-0.6b-q8_0.gguf -c 1024 -np 1 --port 8080 # add -ngl 99 for GPU / iGPU
import json, math, urllib.request
from tokenizers import Tokenizer
r = json.load(open("readout.json"))
tok = Tokenizer.from_file("tokenizer.json")
labels = dict(zip(r["label_ids"], r["label_codes"]))
def score(post: str) -> tuple[float, str]:
ids = tok.encode(r["lead"] + post.strip(), add_special_tokens=False).ids
n = r["first_tokens"] + r["last_tokens"]
pieces = [ids] if len(ids) <= n else [ids[: r["first_tokens"]], ids[-r["last_tokens"]:]]
best = (0.0, "sec")
for piece in pieces:
body = {"prompt": r["prefix_ids"] + piece + r["suffix_ids"], "n_predict": 1, "n_probs": 60,
"temperature": 0, "cache_prompt": True}
req = urllib.request.Request("http://127.0.0.1:8080/completion", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
top = json.loads(urllib.request.urlopen(req).read())["completion_probabilities"][0]["top_logprobs"]
p = {labels[t["id"]]: math.exp(t["logprob"]) for t in top if t["id"] in labels}
total = sum(p.values()) or 1.0
risk = {c: v / total for c, v in p.items() if c != "sec"}
s = 1 - p.get("sec", 0.0) / total
best = max(best, (s, max(risk, key=risk.get) if risk else "sec"))
return best
print(score("老師帶單穩賺不賠,每天 5% 收益,加 LINE 進 VIP 群")) # high score, "ec"
print(score("BTC 短線偏空,MACD 死叉,先觀望等回測 58000")) # low score
onnxruntime
import json, numpy as np, onnxruntime as ort
from tokenizers import Tokenizer
r = json.load(open("readout.json"))
tok = Tokenizer.from_file("tokenizer.json")
sess = ort.InferenceSession("onnx/model.onnx", providers=["CPUExecutionProvider"])
safe = r["label_codes"].index("sec")
def score_piece(piece):
ids = np.array([r["prefix_ids"] + piece + r["suffix_ids"]], dtype=np.int64)
logits = sess.run(None, {"input_ids": ids})[0][0] # 29 risk-code logits
p = np.exp(logits - logits.max()); p /= p.sum()
return 1 - float(p[safe])
Evaluation
Held-out test set of 1,290 items that were never used for training (near-duplicates of test items were also removed from training):
- Minors (harmful): 316 prompts labelled as sexual content involving minors in Nemotron-Safety-Guard v3 / PolyGuardMix, keeping only items whose label refers to the prompt itself.
- Minors (benign): 348 benign prompts that mention children or teenagers (parenting, school, child-protection, news).
- Forum (harmful / benign): 135 / 244 forum-style items — scam SMS (FGRC-SCD), CryptoMind's own scam and calibration posts, and "scary but benign" posts (scam warnings, victims asking for help, crime news, violent slang in trading talk).
- Nemotron zh (harmful / benign): 147 / 100 general Chinese safety prompts (labels are noisier).
Block rate on harmful sets, false-block rate on benign sets:
| Threshold 0.8 | Minors harmful ↑ | Minors benign ↓ | Forum harmful ↑ | Forum benign ↓ | Nemotron harmful ↑ | Nemotron benign ↓ |
|---|---|---|---|---|---|---|
| This model, GGUF Q8_0 | 93.4% | 5.5% | 97.0% | 0.8% | 79.6% | 9.0% |
| This model, ONNX int8 | 92.1% | 5.7% | 96.3% | 0.8% | 78.9% | 12.0% |
| This model, fp16 (unreleased) | 93.0% | 5.7% | 97.0% | 0.8% | 80.3% | 9.0% |
| YuFeng-XGuard-Reason-0.6B (base, no fine-tune) | 92.1% | 12.1% | 68.1% | 4.1% | 83.0% | 19.0% |
| Threshold 0.9 | Minors harmful ↑ | Minors benign ↓ | Forum harmful ↑ | Forum benign ↓ |
|---|---|---|---|---|
| This model, GGUF Q8_0 | 89.6% | 4.0% | 95.6% | 0.8% |
| YuFeng-XGuard-Reason-0.6B (base) | 88.3% | 10.1% | 58.5% | 3.3% |
Fine-tuning mainly teaches the forum boundary: scams are blocked far more reliably, and investing talk and posts that merely mention children are blocked far less often. Q8_0 stays very close to fp16 (mean score difference 0.003, 5/1,290 decisions differ at 0.8); ONNX int8 drifts more (0.020, 23/1,290).
Latency (Apple M4, 2 CPU threads, warm prompt cache): llama.cpp Q8_0 median 388 ms, p95 678 ms per post;
onnxruntime int8 median 1.09 s (no prefix cache). With a GPU / iGPU offload (-ngl 99) it is much faster.
Training
- Base:
Alibaba-AAIG/YuFeng-XGuard-Reason-0.6B@9016029, LoRA r=16 / alpha=32 on all attention and MLP projections, merged. - Data: ~27.5k training posts (block ≈ 54%), Traditional/Simplified Chinese, English and Russian. Harmful examples come only from the public datasets listed above plus CryptoMind-written scam posts; benign examples include CryptoMind UI text, investing chit-chat and hard negatives (scam warnings, victims' stories, child-protection notices).
- Objective (following the base model's first-token SFT format): benign →
sec; harmful → the matching code (fraud →ec, sexual content involving minors →ma, sexual →pc, self-harm →mh, drugs →dc, weapons →dw, doxxing →pp, hate →ac, harassment →cy, threats →ti). Harmful items with an unclear category use a binary loss (all risk codes vs.sec) so the model keeps choosing the category itself. Label smoothing 0.05. - 2 epochs on one Kaggle T4, lr 1e-4, effective batch 32; best epoch picked on validation.
Limitations
- Built for community posts; it is not tuned for chat transcripts or for judging LLM responses.
- It decides "should this post be blocked?", not "which law is broken"; the category code is only meaningful when the score is high.
- Benign posts that mention children are still blocked about 5.5% of the time at 0.8 — keep a human-review / appeal path.
- Weaker on some Russian scams, English "withdrawal locked, verify your seed phrase" scams and some indirect requests.
- Public safety datasets have noisy labels; numbers on the Nemotron set should be read with that in mind.
- No model is perfect: do not use it as the only safeguard for child safety. Report suspected child sexual exploitation to the authorities in your jurisdiction.
License
Apache-2.0 (see LICENSE and NOTICE). This is a modified version of YuFeng-XGuard-Reason-0.6B (Apache-2.0, Alibaba AAIG),
itself based on Qwen3-0.6B (Apache-2.0, Alibaba Cloud). Training data attributions are listed in NOTICE
(Nemotron-Safety-Guard-Dataset-v3 and PolyGuardMix are CC BY 4.0).
Citation
@article{yufeng-xguard,
title={YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models},
author={Alibaba AAIG},
journal={arXiv preprint arXiv:2601.15588},
year={2026}
}
- Downloads last month
- 20
8-bit
Model tree for aaaa47080/CryptoMind-Guard-0.6B
Base model
Qwen/Qwen3-0.6B-Base