LinkinShan/hanja-wsd-base
hanja-wsd-base — 한글 → 국한혼용 변환기 (candidate scorer, preview)
Preview, 2026-09-26. The files in this repository are an earlier model (run6), released as a preview. A new version that learns what to leave in Hangul is previewed below under Next version. Its weights will be released here after the paper is published, and later versions will follow the same way. The end-to-end numbers first published on this card were scored against incomplete gold; see Evaluation correction.
Converts plain-Hangul Korean into mixed Hangul–Hanja script by choosing, for every Sino-Korean word, the correct Hanja spelling from a dictionary candidate set using sentence context. This repository holds the XLM-R base candidate scorer (run6: 21.5M training records + modern annotations) together with the lexicon, priors and the conversion code. A larger checkpoint lives at LinkinShan/hanja-wsd-large.
정부는 발전소 건설과 경제 발전을 위해 공사를 시작했다.
政府는 發電所 建設과 經濟 發展을 위해 工事를 始作했다.
부인은 그 사실을 부인했다.
夫人은 그 事實을 否認했다.Write-ups: 中文 · 한국어 · 국한혼용 (auto-converted by this model)
Next version: preview results
The next version keeps the same cross-encoder and adds one more candidate for every span: KEEP, the Hangul left as it is. One model then decides where a word starts and ends, whether to write it in Hanja, and which Hanja to use, all from sentence context.
- Every 2–8-syllable substring that is a lexicon key becomes a candidate span, without Kiwi token alignment or the corpus-frequency gate. Dynamic programming picks the non-overlapping spans that maximise Σ length × (pHanja − pKEEP − m), with m = −0.3 chosen on the dev set.
- The training labels come from the original mixed-script newspapers. A substring the author wrote in Hanja, alone or as part of a longer Hanja run, is labelled with that Hanja. A substring the author left in Hangul, or one that crosses the edge of a Hanja run, is labelled KEEP.
- It continues run6's training for one pass over 19.3M such records from the 142 training files: 402k steps, about 20 h on 8 × V100-16GB. Sentences of the evaluation sets were removed from the training data.
The test sentences are 1920–1962 newspaper text from files held out from training; test v1 is the 1,000 sentences of the end-to-end table further down. Decoding settings were fixed on a dev set before this single run on the test sets. Control is run6 trained for the same number of steps on the Hanja-labelled records of the same data, without KEEP, and decoded with this repository's pipeline (gate 0.3, τ 0.7). The metric is per-syllable F1: a converted syllable counts as correct only if it carries the author's Hanja.
The confidence intervals come from a paired bootstrap over sentences with 10,000 resamples. No resample reversed any difference in the table, so p < 0.0001 on both sets. Scored on whole Hanja runs instead, where a run counts only if its boundaries and every Hanja in it are right, F1 rises from 0.826 to 0.898 on test v1 and from 0.812 to 0.902 on test v2. The control stays within 0.2 points of run6, so the gain does not come from the extra training. The two systems also differ in decoding and in their negative training examples, so the comparison measures the KEEP setup as a whole, not the KEEP candidate alone.
Most of the gain comes from fewer missed conversions. On test v2, from the control to the next version, syllables left in Hangul where the author wrote Hanja fall from 9,035 to 1,088. Syllables converted where the author kept Hangul fall from 2,538 to 2,150. Wrong Hanja rises from 666 to 1,334, because far more words are now converted. In total, syllable errors fall from 12,239 to 4,572.
Further checks on dev data:
- Seeds. At a pilot scale of 3M records, two seeds of each system differ by at most 0.1 point, while KEEP adds 5.1–5.2 points. At full scale the control gains 0.2 points and KEEP 1.2.
- Context. Pilot models trained with no context, with left context only, and with both sides reach per-syllable F1 0.874, 0.928 and 0.941. Among words whose correct Hanja is not the most frequent candidate, they get 22%, 55% and 69% right.
We have also evaluated on modern legal text, and will add those results once the owner of the reference mixed-script text agrees to their publication.
Evaluation correction
The end-to-end table under Results, and the same table in our arXiv submission, scored against gold that covered only Hanja runs whose whole Hangul form is a lexicon entry. About 17–18% of the Hanja syllables in those sentences had no gold, so correct conversions of long compounds counted as spurious. An earlier version of this card blamed the low precision on words the author happened to leave in Hangul; that explanation was wrong. Scored per syllable on the same 1,000 sentences, run6 gets:
71% of the syllables that the earlier gold counted as spurious were correct conversions. run7 and run5 have not been rescored. The Next version results above use complete gold.
How it works
- Candidate spans. Kiwi tokenises the sentence; every 2–8 syllable substring aligned to token boundaries (and covering only noun-like tokens) is looked up in
data/lexicon.tsv(2.2M Hangul forms, 140k ambiguous). Each hit is a candidate span with its list of possible Hanja. - Scoring. The model is a cross-encoder: input is the sentence with the target span marked by 【】, paired with one candidate Hanja string; a linear head on the
<s>state gives a score. Softmax over the span's candidates gives $p(h \mid x)$. Because the candidate is plain text, the encoder can score Hanja strings it never saw in training (this is why a multilingual encoder was used: XLM-R's tokenizer leaves only 0.54 % of dictionary Hanja as<unk>; KLUE-RoBERTa, which does not know Hanja, scores below the frequency baseline). - Convert or not. A conversion prior $p{\text{sino}}(w)=n{漢}/(n{漢}+n{한}+1)$ from 1920–1962 mixed-script newspapers gates spans (native words ≈ 0); dynamic programming then picks non-overlapping spans maximising $\sum (b-a)(p_{\max}-\tau)$.
Results
Test sets are held out by file (OKHC) or by page (wiki/namuwiki); the lexicon and all counts were built from training files only.
Record-level accuracy (gold Hanja among ≤12 candidates; the scorer's own task):
\* LLM figures are on 100–150-item samples of the same candidate sets. run5 (not released) is listed because it is the best checkpoint on Wikipedia; more historical data (run7) helped the hard subset and hurt modern text slightly.
End-to-end conversion (1,000 held-out OKHC sentences, 8,333 gold spans, --preset historical):
These end-to-end numbers were scored against incomplete gold; see Evaluation correction above. On modern text (Wikipedia, --preset modern) span recall is 0.74; precision cannot be measured against bracket annotations. Remaining errors on modern text are mostly person and place names (金泳三 vs 金永三), which context cannot resolve.
Training
Data size mattered more than model size: base went 0.815 → 0.892 → 0.947 on the hard subset at 2M → 8M → 21.5M records, overtaking large at 2M (0.872) and 4M (0.908). Modern annotations did nothing for historical words and lifted Wikipedia from 0.53 to 0.77–0.82.
Training data is derived from the Open Korean Historical Corpus (newspapers 1920–1962, public domain; corpus CC BY-NC 4.0), Korean Wikipedia (CC BY-SA), Namuwiki dump (CC BY-NC-SA 2.0) and Wiktionary/kaikki (CC BY-SA). The NC-SA terms carry over to these weights.
Usage
pip install torch transformers sentencepiece kiwipiepy huggingface_hub
hf download LinkinShan/hanja-wsd-base --local-dir hanja-wsd-base
cd hanja-wsd-base
python src/convert.py --ckpt . --lexicon data/lexicon.tsv --hangul-counts data/hangul_counts.tsv \
--preset modern "그녀는 아침마다 화장을 한다." "시신을 화장하였다."--preset historical (gate 0.3, τ 0.7) suits 1920–60s text; --preset modern (gate 0.05, τ 0.5) suits current text. --stdin reads one sentence per line and emits JSON with every chosen span's candidates and probabilities. src/serve.py starts a small HTTP API + web page. CPU: ~150 ms/sentence for base; V100: ~15 ms.
Loading the scorer directly:
import torch
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("LinkinShan/hanja-wsd-base")
enc = AutoModel.from_pretrained("LinkinShan/hanja-wsd-base", add_pooling_layer=False)
head = torch.nn.Linear(enc.config.hidden_size, 1)
head.load_state_dict(torch.load("head.pt", map_location="cpu")) # from the repo
ctx = "그녀는 아침마다 【화장】을 한다."
x = tok([ctx]*3, ["化粧", "火葬", "畫匠"], return_tensors="pt", padding=True)
scores = head(enc(**x).last_hidden_state[:, 0]).squeeze(-1)
print(scores.softmax(-1)) # 化粧 ≈ 0.98Limitations
- Trained on 1920–1962 newspaper language. Modern coinages, loanwords in Hanja and current proper nouns are under-represented; the modern-text preset compensates only partly.
- Person and place names are chosen by frequency, not knowledge.
- The lexicon decides what can be converted; words absent from it (recall 90.6 % on newspapers, 81 % on Wikipedia) stay in Hangul unless the character-level fallback is on.
- No human evaluation yet; all numbers are against automatically derived gold.
Citation and attribution
Paper: submitted to arXiv (cs.CL), DOI to be added on announcement. Redistribution of these weights, the lexicon or derived outputs must credit Qifeng Xu (ORCID 0009-0009-1584-1519) and link to this repository or https://shan.ink (CC BY-NC-SA 4.0 attribution clause).
@misc{hanja-wsd-2026,
title = {hanja-wsd: converting plain-Hangul Korean to mixed script with a candidate-scoring cross-encoder},
author = {Qifeng Xu},
note = {ORCID: 0009-0009-1584-1519},
year = {2026},
url = {https://huggingface.co/LinkinShan/hanja-wsd-base}
}