Instructions to use cadazar/han2han-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cadazar/han2han-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cadazar/han2han-it", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("cadazar/han2han-it", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Han2Han IT
Instruction-tuned checkpoint of Han2Han
(han2han-ul2-base-1-it, step 43153). Han2Han is a 169M-parameter
encoder-decoder model that learns script-invariant representations of Korean
text: a document written in Hanja and its Hangul transcription land at the same
point in embedding space. The recipe (jamo and character-level embedding fusion,
morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described
in the paper, accepted to Findings of EMNLP 2026.
This repo holds the PyTorch weights, the SentencePiece tokenizer, and the
modeling code needed to load them through the transformers Auto classes with
trust_remote_code=True. Training code, the Flax model, the Flax-to-PyTorch
converter, and the evaluation pipeline live in the GitHub repo.
Intended use
Han2Han is meant for representations: classification and fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike. The model generates text, but generation is not what it was built or evaluated for, and improving it is follow-up work:
- Hanja to Hangul transcription usually works on short sentences, but can leave
some Hanja untranscribed or, on some inputs, repeat itself until the token
limit (
no_repeat_ngram_sizehelps with that). - Hangul to Hanja restoration (
ํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:) is part of the training mix, but this checkpoint mostly returns the Hangul input unchanged.
The paper and hanja_transcription_demo.ipynb in the GitHub repo cover both
directions in detail.
Usage
Runtime requirements: torch, transformers, sentencepiece, regex, and
numpy (tested with torch 2.12.0 and transformers 5.9.0 on CPU).
Sentence embeddings
output_sentence_embeddings=True returns the mean-pooled encoder states, one
640-dimensional vector per input.
import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
texts = [
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
"ํ์ฅ์ ์ผ์ํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.",
"ๅ็ต๋ ๅๅๅญ์์๋ ็ ดๅขจ์ ๅฆๆณ์ผ๋ก ็ฝ้ช์ ่ฑกๅพตํ ์ ์๋ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9444, 0.8789],
# [0.9444, 1.0000, 0.8459],
# [0.8789, 0.8459, 1.0000]])
The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.
Generation
generate() is the standard transformers one, with a KV cache; greedy
decoding, beam search, and sampling all work. generation_config.json sets the
decoder start and end tokens.
prompt = (
"<|user|>ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:\n"
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.<|end_of_turn|>"
)
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0].tolist(), skip_special_tokens=True))
# ํ์ฅ์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.
The output is shown as generated; ไธๅทก should have become ์ผ์.
Prompt format, as used by the SFT collator in the GitHub repo:
- Encoder input:
<|system|>{system}<|user|>{user}<|end_of_turn|>. The system part is optional. - The decoder starts from
<|assistant|>(decoder_start_token_id9) and a turn ends with<|end_of_turn|>(eos_token_id10).
Han2HanTokenizer is a plain SentencePiece wrapper rather than a
PreTrainedTokenizer subclass. It exposes __call__, encode, and decode;
__call__ does not add BOS or EOS tokens, and special tokens written into the
text are mapped to their ids.
Files
| File | Contents |
|---|---|
model.safetensors |
fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables |
config.json, generation_config.json |
model and generation config, with auto_map entries for the Auto classes |
spiece.model, tokenizer_config.json |
SentencePiece model (38400 pieces) and tokenizer config |
modeling_han2han.py, han2han_config.py, han2han_tokenizer.py |
modeling code from the GitHub repo at commit 7304c38 |
han2han_config.py is identical to the GitHub copy. modeling_han2han.py
(modeling_han2han_pytorch.py there) differs in two places: the config import is
relative, and the module-level register_han2han import is replaced by the
auto_map entries. han2han_tokenizer.py resolves Hub repo ids and logs
through a plain logging logger instead of the repo's JAX-aware helper.
Instruction tuning
Fine-tuned from the Han2Han pre-trained checkpoint on instruction following,
chain-of-thought reasoning, Hanja-Hangul article transcription, and
summarization data (configs/it-muon-stage_1.yaml in the GitHub repo).
Citation
@inproceedings{han2han2026,
title = {Han2Han: Efficient Language-Specific Character Representation
through Script-Aware Pre-Training for Historical Text Analysis},
author = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}
License
Apache License 2.0, the same as the GitHub repo.
- Downloads last month
- 694