AI & ML interests

None defined yet.

Recent Activity

thaki-AI  updated a collection about 13 hours ago
EmbeddingGemma 2 Edge — Quantization
thaki-AI  updated a collection about 13 hours ago
RAG-Gate — Answer / Retrieve / Stop Gates for RAG
thaki-AI  updated a model about 13 hours ago
ThakiCloud/eg2-text-q4
View all activity

Organization Card
Thaki Cloud

Enterprise Agent Automation Platform · Seoul, South Korea

One Paxis. Many Workflows. Any Cloud.

We turn digital work into AI-operated workflows. Paxis is the enterprise agent platform; Metis (inference) and Maxis (training) are the model layer that makes running those agents affordable, on Telox / Velox / Aegis infrastructure under one Signum control plane.

This organization publishes the model side of that stack — the checkpoints, recipes and benchmarks we actually serve, plus the evidence to reproduce the numbers on their cards.


Start here

Browse by release line, not by the flat model list — every public checkpoint belongs to exactly one line.

Line What it is
Human-KO — Korean style-aligned 27B Qwen3.8-27B retrained to write natural Korean prose. Bullet rate 97.5% → 2.0%, CJK contamination 2.55% → 0.33%. Full weights, a safety-aligned variant, and serving quantizations.
SkillRet — agent skill retrieval Embeddings that let an LLM agent find the right skill inside a 1,000+ skill library, with the benchmark, the index, and the paper. Our most-downloaded line.
Language Confusion Suppression Weight-level vocabulary pruning recipes for 6 languages, so a multilingual model stops answering in the wrong script. Masks and scripts, not weights.
Evals & Reproducibility The datasets behind our numbers — including the requantization noise floor you should read before attributing a 2pp difference to your recipe.

Serving quantization, by base model: Qwen3.8-27B Qwen3-Coder-30B-A3B Qwen3-30B-A3B Vision-Language


Which quantization should I pull?

Pick by the GPU you actually have, not by the biggest number in the filename.

Your GPU Format Why
B200, GB200 (Blackwell, SM100+) NVFP4 FP4 tensor cores exist in hardware. Below Blackwell there is no kernel to run it on.
H100, H200, A100 (Hopper · Ampere) W4A16 (GPTQ or AWQ) Weight-only 4-bit with fp16 activations — the safe default everywhere.
Blackwell, memory-bound at long context NVFP4 + FP8 attention Mixed precision: FP4 weights, FP8 attention.
Evaluating quantization itself NVFP4 / NVFP4-GPTQ / NVFP4-RTN Same model, algorithm as the only variable — RTN is the no-calibration control.

How we report numbers

Model cards here state the sample size, the temperature, and what the measurement cannot tell you. Three habits we hold to:

We name the noise floor before we claim a win. Two checkpoints built by running the same quantization twice — same model, same settings, same calibration — diverged by 3.56pp on GSM8K. Anything smaller than that is unattributable, and we write it that way instead of taking credit.

We say when a metric cannot discriminate. If a benchmark's minimum detectable difference is larger than the gap we observed, the honest sentence is "we don't know," not "no regression."

We say when a judge is uncalibrated. Where an LLM judge has not been checked against human labels, the number is a judge-axis win rate — not human evaluation, and we label it as such.


Papers

Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery — arXiv:2610.00164. Synthetic Korean dialect data recovers 91.2% of the real-data gain on region identification but only 63.7% on comprehension — and when the scorer shares its marker inventory with the sentence builder, even that is overstated. Withholding marker types dropped dialectness recovery from 91.8% to 8.1%. Ships with KoDialectBench, 1,000 items across five regions.

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model — arXiv:2609.11291. Aligning Qwen3.8-27B to Korean response style also moved two behaviors nobody trained for: abstention on ambiguous social questions, and unprompted disclosure in securities guidance. Both moved through the model's emission policy — how often it answers and how much it says. The Human-KO cards should be read next to it.

SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents — arXiv:2605.05726. The benchmark, index and embeddings behind the SkillRet line.

🌐 thakicloud.com · 📝 Tech blog · 💻 GitHub