Text Generation
Korean

🏯 Hanok LLM 1.0

Korean-first modular language model stack

🌐 Languages: English Β· ν•œκ΅­μ–΄

A clean, native PyTorch Transformer implementation for Korean LLM experimentation, training, inference, checkpointing and data workflows β€” built for commercialization.

Python 3.10+ PyTorch Transformers Datasets Windows 11 CUDA License

Hanok LLM overview

20 Layers Β· 1,920 Hidden Β· 10 Heads Β· SwiGLU 4,800 Β· ~1.09B Parameters

Repository: https://github.com/seoan1024/Hanok-LLM.git
Research base: https://github.com/seoan1024/Korean-llm.git


✦ What is Hanok?

Hanok LLM is built on the foundation of continuous research from the Korean-llm project.
It is a modular, production-oriented Korean language model stack designed with full commercialization as the clear and explicit goal.

The entire system lives under a maintainable Python package (src/hanok/) with strict separation of concerns. Every major subsystem β€” model, data, training, inference, evaluation, export and monitoring β€” can be inspected, extended or replaced independently.

Layer Responsibility
Model Native decoder-only Transformer (KoreanLLM) with RMSNorm, RoPE, SwiGLU, tied embeddings and KV cache
Data Hugging Face dataset loading, Parquet caching, Korean-oriented field parsing and response-only SFT masking
Training Pretrain β†’ SFT orchestration, gradient accumulation, cosine schedule, checkpoint resume, optional GUI
Inference Native .pth loading + temperature / top-k / top-p / repetition-penalty generation
Evaluation Perplexity and a reproducible Korean multiple-choice regression benchmark
Export Standard Transformers Llama bundle, llama.cpp GGUF, and Ollama model creation
UI Optional real-time training monitor GUI

Build a Korean LLM stack you can actually inspect β€” and ship.
Native PyTorch. Explicit defaults. Local checkpoints. Practical Windows + CUDA workflow.

Trained model weights are not included. Export commands convert a checkpoint supplied by the user.


🧠 Architecture in depth

Hanok architecture

Core configuration

Parameter Value Notes
Transformer blocks 20 n_layers
Hidden dimension 1,920 dim
Attention heads 10 n_heads
Head dimension 192 dim // n_heads
SwiGLU intermediate 4,800 dim Γ— 2.5
Normalization RMSNorm eps=1e-6
Positional encoding RoPE ΞΈ = 10,000
Residual style Pre-norm Attention & FFN
Embedding / output Tied weights output.weight = embed.weight
KV cache Supported Per-layer, detached
Default training seq length 1,024 TrainingConfig.max_seq_len; configurable
Direct model constructor 4,096 KoreanLLM(max_seq_len=4096) default RoPE table length
Training / CLI context 1,024 Training, inference, evaluation, and export defaults
Vocab size (default) 128,256 Follows loaded tokenizer
Instantiated parameters ~1.09B Depends on final vocab size

The sequence limit is the RoPE table length allocated for that model instance, not a fixed architecture-wide ceiling. The KoreanLLM constructor defaults to 4,096 tokens, while the training and CLI loading paths default to 1,024; pass a larger --max-seq-len to configure a longer context. Generation counts the prompt and new tokens together, truncates an overlong prompt from the left, and stops at the configured context limit. Context lengths above 1,024 are configurable but are not the default training length and have not been validated for quality here; memory use also grows with sequence length.

Building blocks

RMSNorm
Root-mean-square normalization without mean centering. Lightweight and stable for deep residual stacks.

RoPE (Rotary Positional Embeddings)
Applied to both queries and keys. Precomputed cosine/sine tables are registered as buffers and sliced by absolute position (supports KV-cache incremental decoding).

Attention

  • Multi-head self-attention with F.scaled_dot_product_attention
  • Causal masking by default (is_causal=True when no external mask is given)
  • Optional external attention mask
  • KV cache concatenation along the sequence dimension for efficient generation

SwiGLU Feed-Forward
Three linear projections (w1, w2, w3) with no bias:
w2(SiLU(w1(x)) * w3(x)). Intermediate size is dim Γ— 2.5 = 4,800.

TransformerBlock
Pre-norm residual:
x = x + Attention(RMSNorm(x))
x = x + SwiGLU(RMSNorm(x))

KoreanLLM

  • Embedding β†’ 20 Γ— TransformerBlock β†’ final RMSNorm β†’ tied output projection
  • Weight initialization: normal(std=0.02) for Linear and Embedding
  • Training path uses torch.utils.checkpoint (non-reentrant) for memory efficiency
  • Loss: token-level cross-entropy with ignore_index=-100, reduced by valid target count (handles empty-target batches safely)

Design philosophy

The architecture deliberately stays close to modern decoder-only practices (RMSNorm + RoPE + SwiGLU + tied embeddings) while remaining fully transparent. There are no opaque custom CUDA kernels or closed-source components β€” every tensor operation is visible in pure PyTorch.


⚑ Training stack

Training pipeline

Default TrainingConfig

stage                      auto | pretrain | sft
batch_size                 2
accumulation_steps         8          # effective batch = 16
learning_rate              5e-5
sft_learning_rate          1.5e-5
warmup_steps               200
scheduler                  Cosine
weight_decay               0.1
max_steps                  100,000    # manual-stage default
pretrain_steps             100,000
sft_steps                  10,000
eval_interval              5,000
max_seq_len                1024
num_workers                4
use_bfloat16               True
use_8bit_optimizer         True       # AdamW8bit β†’ AdamW fallback
validation_split           0.02
validation_batches         32
enable_gui                 True
seed                       42
checkpoint_dir             checkpoints
force_redownload           False
stream_datasets            False
samples_per_dataset        None
resume_from_checkpoint     None

Training stages

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   DATA SOURCES   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ CACHE / PARQUET  β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚       PRETRAIN        β”‚
                 β”‚  KorMix / custom data β”‚
                 β”‚  (100k steps default) β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚ checkpoint
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚          SFT          β”‚
                 β”‚ KuLLM + KoAlpaca      β”‚
                 β”‚ (10k steps default)   β”‚
                 β”‚ response-only labels  β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  NATIVE .PTH    β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • auto mode: runs full pretraining first, then starts SFT from the latest pretrain checkpoint (fresh optimizer & schedule for SFT).
  • Checkpoint resume: supported for both stages.
  • SFT label masking: only response tokens contribute to the loss; prompt tokens are set to -100.
  • Pretraining packing: complete documents are joined into full 1,024-token blocks; Parquet rows are shuffled within row groups to reduce repeated reads.
  • SFT updates: rank-8 LoRA adapters train while the pretrained base weights stay frozen. Native checkpoints retain the adapters; HF export merges them into the base layers.
  • Validation: pretraining holds out about 2% by Parquet row group; SFT uses a 2% random split. Evaluated every eval_interval steps (up to validation_batches).
  • GUI: optional real-time loss monitor (enable_gui=True by default; disable with --no-gui).

Optimizer & scheduler

  • Primary: AdamW8bit (bitsandbytes) when available
  • Fallback: standard AdamW
  • Learning-rate schedule: Cosine annealing with linear warmup (200 steps)

Memory & precision

  • BF16 autocast on CUDA when use_bfloat16=True
  • Gradient checkpointing active during training
  • Pin memory for DataLoader on CUDA devices
  • SFT uses rank-8 LoRA adapters, reducing trainable parameters and optimizer memory.

πŸ–₯️ Recommended workstation

Windows 11 + RTX 5090 profile
Component Recommended profile
OS Windows 11
GPU NVIDIA GeForce RTX Series
VRAM 24 GB+ recommended (estimated pretraining use: 18–21 GB)
System RAM 64 GB+
CPU High-performance multi-core
Python 3.10+ (documented on 3.11)
Precision BF16 on supported CUDA hardware
Storage Fast SSD (datasets + checkpoints grow quickly)

Memory note
With sequence length 1,024, default batch size and accumulation, FP32 model weights with BF16 autocast, AdamW8bit, and gradient checkpointing, peak pretraining VRAM is estimated at 18–21 GB. This is an estimate, not a measured result; GPU model and CUDA allocation state affect actual usage. LoRA SFT has a different memory profile. Actual usage also depends on batch size, sequence length, optimizer state, whether gradient checkpointing is active, and whether you are training or only running inference. Longer sequences or larger effective batches will require more VRAM or reduced accumulation.


πŸ§ͺ Validation & roadmap status

Validation status
Claim Status
Local Windows 11 / RTX 5090 Laptop GPU testing βœ…
Model construction & native inference βœ…
Training / checkpoint / resume workflow βœ…
Architecture / data / move-parity tests βœ…
Model Card & Data Card included βœ…
Commercialization as the explicit goal βœ…
Trained weights + official benchmark scores πŸ”œ Release
Transformers Llama export code βœ… Included
llama.cpp GGUF / Ollama export code βœ… Included; external tools required
Official Korean quality benchmark results Not run

🧩 Technology stack

Technology stack
Category Libraries / tools
Core Python 3.10+, PyTorch 2.x
Tokenization Hugging Face Transformers (beomi/Llama-3-Open-Ko-8B)
Data Hugging Face Datasets, PyArrow, Pandas
Optimization bitsandbytes (optional 8-bit AdamW)
Monitoring matplotlib + optional GUI monitor
Packaging setuptools, pyproject.toml, CLI entrypoint hanok

πŸ“¦ Project structure

Hanok-LLM/
β”œβ”€β”€ src/hanok/
β”‚   β”œβ”€β”€ model/
β”‚   β”‚   └── architecture.py          # KoreanLLM (RMSNorm Β· RoPE Β· SwiGLU Β· KV cache)
β”‚   β”œβ”€β”€ data/
β”‚   β”‚   └── legacy.py                # DatasetManager, LocalKoreanDataset, collate, Parquet cache
β”‚   β”œβ”€β”€ training/
β”‚   β”‚   β”œβ”€β”€ config.py                # TrainingConfig dataclass (all defaults)
β”‚   β”‚   β”œβ”€β”€ trainer.py               # Main orchestration (auto / pretrain / sft)
β”‚   β”‚   β”œβ”€β”€ optimizer.py             # AdamW / AdamW8bit selection
β”‚   β”‚   β”œβ”€β”€ scheduler.py             # Cosine + warmup
β”‚   β”‚   β”œβ”€β”€ checkpoint.py            # Save / load / find_latest
β”‚   β”‚   └── reproducibility.py       # Seed & distributed helpers
β”‚   β”œβ”€β”€ inference/
β”‚   β”‚   β”œβ”€β”€ loader.py                # Checkpoint + tokenizer loading
β”‚   β”‚   └── generation.py            # Autoregressive generation with KV cache
β”‚   β”œβ”€β”€ evaluation/
β”‚   β”‚   β”œβ”€β”€ perplexity.py
β”‚   β”‚   β”œβ”€β”€ benchmark.py
β”‚   β”‚   └── benchmark_ko.jsonl       # local sanity suite, not an official score
β”‚   β”œβ”€β”€ export/
β”‚   β”‚   β”œβ”€β”€ huggingface.py           # Map Hanok weights to standard Transformers Llama
β”‚   β”‚   └── ollama.py                # llama.cpp GGUF conversion, Modelfile and Ollama creation
β”‚   β”œβ”€β”€ ui/
β”‚   β”‚   └── monitor.py               # Optional training GUI
β”‚   └── cli.py                       # `hanok` entrypoint (train / infer)
β”œβ”€β”€ assets/svg/                      # README visual kit
β”œβ”€β”€ MODEL_CARD.md
β”œβ”€β”€ DATA_CARD.md
β”œβ”€β”€ CONTRIBUTING.md
β”œβ”€β”€ SECURITY.md
β”œβ”€β”€ pyproject.toml
└── requirements.txt

πŸ› οΈ Installation (Windows 11)

# 1. Create / activate a Python 3.10+ environment
python --version

# 2. Install PyTorch with the CUDA build that matches your driver
#    β†’ https://pytorch.org (select CUDA version carefully)

# 3. Install the project in editable mode
python -m pip install -e .

# 4. Optional: 8-bit optimizer support
python -m pip install -e ".[8bit]"

# 5. Verify the CLI is available
hanok --help

Dependencies (from requirements.txt / pyproject.toml)

torch
transformers
datasets
safetensors
bitsandbytes          # optional, for 8-bit AdamW
pyarrow
pandas
requests
tqdm
matplotlib

🎯 Quick start

SFT (start from a pretrained checkpoint)

hanok train --stage sft --resume-from-checkpoint checkpoints/pretrain/hanok_llm_10000.pth --no-gui

Pretraining

hanok train --stage pretrain --no-gui

Auto pipeline (pretrain β†’ SFT)

hanok train --no-gui

The automatic path first completes pretraining (or resumes the latest pretrain checkpoint), then starts SFT from that checkpoint with a fresh optimizer and learning-rate schedule. For a clean run after changing the training pipeline, move existing checkpoints/pretrain and checkpoints/sft folders aside first. Auto mode resumes the latest pretraining checkpoint it finds, and older full-finetune SFT checkpoints cannot resume as LoRA checkpoints.

Useful training flags

Flag Description
--stage {auto,pretrain,sft} Training stage
--no-gui Disable the optional training monitor
--stream-datasets Stream instead of full download
--samples-per-dataset N Limit examples per dataset (smoke / debug)
--max-steps N Override max training steps
--batch-size N Override batch size
--max-seq-len N Set the RoPE context and training sequence length (default: 1,024)
--resume-from-checkpoint Path or latest
--force-redownload Re-download datasets

πŸ’¬ Inference

hanok infer `
  --checkpoint checkpoints/korean_llm_00010.pth `
  --prompt "μ•ˆλ…•ν•˜μ„Έμš”. λ„ˆλŠ” λˆ„κ΅¬λ‹ˆ?" `
  --max-tokens 128

Generation controls

Flag Default Description
--temperature 0.6 Softmax temperature (> 0 required)
--top-k 40 Keep only top-k logits
--top-p 0.95 Nucleus sampling
--repetition-penalty 1.3 Penalize already-generated tokens
--max-seq-len 1024 Context window used at load time
--device auto cuda / cpu

How generation works

  1. Prompt is wrapped in the instruction format:
    ### 질문: {prompt}\n### 응닡:
  2. Tokens are generated autoregressively with a KV cache.
  3. Temperature scaling β†’ optional top-k filtering β†’ softmax β†’ optional top-p nucleus filtering β†’ multinomial sampling.
  4. Repetition penalty is applied to previously seen tokens.
  5. Stops on EOS or when the sequence length safety limit is reached.
  6. Only the text after ### 응닡: is returned.

πŸ—‚οΈ Data sources & pipeline

Default datasets

Stage Dataset Config Split Notes
Pretraining AdaMLLab/KorMix minhash_deduped train Text corpus
SFT nlpai-lab/kullm-v2 default train Instruction
SFT beomi/KoAlpaca-v1.1a default train Instruction

Tokenizer: beomi/Llama-3-Open-Ko-8B
A dedicated pad token <|pad|> is added when the tokenizer does not already provide a distinct pad token.

Data handling details

  • Hugging Face downloads are cached under ./datasets/cache.
  • Downloaded rows are stored as Parquet and recorded in a manifest.
  • Field parsing supports multiple common schemas:
    text / content / document / body,
    instruction / input / output,
    question / response, question / answer, prompt / response.
  • SFT mode masks the prompt portion; only response tokens contribute to the loss.
  • EOS is appended; PAD positions are ignored via ignore_index=-100.
  • Optional streaming and samples_per_dataset limits are available for rapid iteration.

Always verify the original dataset licenses and terms of use before training.
This repository does not redistribute or re-license third-party data.


πŸ“€ Checkpoint export & evaluation

Export a Hanok checkpoint to the standard Transformers LlamaForCausalLM format:

hanok export --checkpoint checkpoints/sft/hanok_llm_10000.pth --format hf --output exports/hanok-hf

# GGUF conversion (requires an existing llama.cpp checkout)
hanok export --checkpoint checkpoints/sft/hanok_llm_10000.pth --format gguf `
  --output exports/hanok-f16.gguf --llama-cpp-dir D:\tools\llama.cpp

# Quantize and register in Ollama (requires llama-quantize and Ollama)
hanok export --checkpoint checkpoints/sft/hanok_llm_10000.pth --format ollama `
  --output exports/hanok-f16.gguf --llama-cpp-dir D:\tools\llama.cpp `
  --quantization Q4_K_M --ollama-name hanok:latest

Load the HF folder with AutoModelForCausalLM.from_pretrained("exports/hanok-hf"). GGUF/Ollama export runs external llama.cpp and optional quantization tools.

Run hanok evaluate --checkpoint ... to evaluate the bundled 15-question Korean multiple-choice suite and three Korean free-response prompts with deterministic generation. The JSON report includes accuracy, answer format rate, Hangul ratio, repeated-token ratio, empty-response rate, choice-letter perplexity, and every generated response. Choice-letter perplexity only describes this small test format and is not comparable to training validation loss. This is a regression check, not an official or comprehensive Korean capability benchmark.

This repository does not include trained weights. Hanok initializes from scratch; export does not add pretrained Llama knowledge or improve a checkpoint's quality. Results depend on the training data and training budget.


πŸ”¬ Model Card & Data Card

Document Purpose
MODEL_CARD.md Architecture summary, intended use, limitations, tokenizer notes
DATA_CARD.md Dataset sources, preprocessing behavior, reproducibility caveats
CONTRIBUTING.md How to contribute
SECURITY.md Security reporting process

🧭 Design principles

  1. Inspectability β€” every layer is ordinary PyTorch; no black-box kernels.
  2. Explicit defaults β€” training hyperparameters are declared in one place (TrainingConfig).
  3. Local-first β€” checkpoints are native .pth files you can load without a remote registry.
  4. Windows + CUDA practicality β€” the documented path is a real workstation workflow, not a theoretical Linux-only lab.
  5. Commercial trajectory β€” the architecture and tooling are built so that a production release (weights + benchmarks + HF/Ollama) can land cleanly on top of this codebase.

πŸ“œ License

See LICENSE for the repository license.
Third-party datasets, tokenizers and dependencies carry their own terms. Always review them before any commercial use.


🏯 HANOK · KOREAN LLM ENGINEERING

Inspectable architecture Β· Practical training Β· Native inference Β· Windows/CUDA workflow Β· Built for commercialization

Repo: Hanok-LLM Β· Research base: Korean-llm

Hanok closing panel
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train seoan1024/Hanok-LLM