Han2Han IT

Instruction-tuned checkpoint of Han2Han (han2han-ul2-base-1-it, step 43153). Han2Han is a 169M-parameter encoder-decoder model that learns script-invariant representations of Korean text: a document written in Hanja and its Hangul transcription land at the same point in embedding space. The recipe (jamo and character-level embedding fusion, morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described in the paper, accepted to Findings of EMNLP 2026.

This repo holds the PyTorch weights, the SentencePiece tokenizer, and the modeling code needed to load them through the transformers Auto classes with trust_remote_code=True. Training code, the Flax model, the Flax-to-PyTorch converter, and the evaluation pipeline live in the GitHub repo.

Intended use

Han2Han is meant for representations: classification and fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike. The model generates text, but generation is not what it was built or evaluated for, and improving it is follow-up work:

  • Hanja to Hangul transcription usually works on short sentences, but can leave some Hanja untranscribed or, on some inputs, repeat itself until the token limit (no_repeat_ngram_size helps with that).
  • Hangul to Hanja restoration (ํ•œ๊ธ€์„ ํ•œ์ž๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:) is part of the training mix, but this checkpoint mostly returns the Hangul input unchanged.

The paper and hanja_transcription_demo.ipynb in the GitHub repo cover both directions in detail.

Usage

Runtime requirements: torch, transformers, sentencepiece, regex, and numpy (tested with torch 2.12.0 and transformers 5.9.0 on CPU).

Sentence embeddings

output_sentence_embeddings=True returns the mean-pooled encoder states, one 640-dimensional vector per input.

import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()

texts = [
    "ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ํšŒ์žฅ์„ ์ผ์ˆœํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ์ผ๋ณธ์ธ์ธก ํ™”๊ฐ€๊นŒ์ง€๋„ ์ผ์ธ๋„ ๋ฐœ๊ฒฌํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ๅ—็•ต๋‚˜ ๅ››ๅ›ๅญ์—์„œ๋Š” ็ ดๅขจ์˜ ๅฆ™ๆณ•์œผ๋กœ ็™ฝ้›ช์„ ่ฑกๅพตํ•  ์ˆ˜ ์žˆ๋‹ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
    embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9444, 0.8789],
#         [0.9444, 1.0000, 0.8459],
#         [0.8789, 0.8459, 1.0000]])

The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.

Generation

generate() is the standard transformers one, with a KV cache; greedy decoding, beam search, and sampling all work. generation_config.json sets the decoder start and end tokens.

prompt = (
    "<|user|>ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:\n"
    "ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.<|end_of_turn|>"
)
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0].tolist(), skip_special_tokens=True))
# ํšŒ์žฅ์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ์ผ๋ณธ์ธ์ธก ํ™”๊ฐ€๊นŒ์ง€๋„ ์ผ์ธ๋„ ๋ฐœ๊ฒฌํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.

The output is shown as generated; ไธ€ๅทก should have become ์ผ์ˆœ.

Prompt format, as used by the SFT collator in the GitHub repo:

  • Encoder input: <|system|>{system}<|user|>{user}<|end_of_turn|>. The system part is optional.
  • The decoder starts from <|assistant|> (decoder_start_token_id 9) and a turn ends with <|end_of_turn|> (eos_token_id 10).

Han2HanTokenizer is a plain SentencePiece wrapper rather than a PreTrainedTokenizer subclass. It exposes __call__, encode, and decode; __call__ does not add BOS or EOS tokens, and special tokens written into the text are mapped to their ids.

Files

File Contents
model.safetensors fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables
config.json, generation_config.json model and generation config, with auto_map entries for the Auto classes
spiece.model, tokenizer_config.json SentencePiece model (38400 pieces) and tokenizer config
modeling_han2han.py, han2han_config.py, han2han_tokenizer.py modeling code from the GitHub repo at commit 7304c38

han2han_config.py is identical to the GitHub copy. modeling_han2han.py (modeling_han2han_pytorch.py there) differs in two places: the config import is relative, and the module-level register_han2han import is replaced by the auto_map entries. han2han_tokenizer.py resolves Hub repo ids and logs through a plain logging logger instead of the repo's JAX-aware helper.

Instruction tuning

Fine-tuned from the Han2Han pre-trained checkpoint on instruction following, chain-of-thought reasoning, Hanja-Hangul article transcription, and summarization data (configs/it-muon-stage_1.yaml in the GitHub repo).

Citation

@inproceedings{han2han2026,
  title     = {Han2Han: Efficient Language-Specific Character Representation
               through Script-Aware Pre-Training for Historical Text Analysis},
  author    = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}

License

Apache License 2.0, the same as the GitHub repo.

Downloads last month
694
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support