HERIT addresses the data scarcity and temporal bias inherent in historical archives written in Hanja. It combines high-quality pseudo-labeled data augmentation via retrieval-augmented generation (RAG) with two-stage fine-tuning.
- Mitigates temporal bias: human evaluations stratified by reign confirm that HERIT translates robustly across historical periods not covered by expert-translated data.
- Strong empirical results: HERIT outperforms strong baselines on lexical-overlap metrics and achieves a preference rate of about 66% in human expert evaluation.
- Large-scale pseudo-labeling: 315K RAG-based pseudo-labeled documents for augmentation, and 2.09M final Hanja-to-English translations.
- Full pipeline released: this repository contains the code used to build HERIT β data crawling and preprocessing, pseudo-labeling, fine-tuning, and evaluation.
Note
Work in progress; to be updated by early October 2026.
| Dataset | # Documents | Released Fields |
|---|---|---|
| 15.7K | (Document ID, Hanja) | |
| 1,000 | (Document ID, Hanja) | |
| 1,000 | (Document ID, Hanja) | |
| 2,080 | (Document ID, Hanja) | |
| Human-eval subset | 200 | (Document ID, Hanja) |
| 315K | (Document ID, Hanja, Pseudo-English) | |
| Final translations | 2.09M | (Document ID, Hanja, Pseudo-English) |
HERIT model checkpoints (two-stage fine-tuned from Qwen3 8B / 32B). To be released.
| Model | Base Model | Link |
|---|---|---|
| HERIT-8B | Qwen3-8B | TBA |
| HERIT-32B | Qwen3-32B | TBA |
Closed-source models (API)
- Gemini-2.5-Flash (
gemini-2.5-flash, updated in June 2025) - Gemini-2.5-Pro (
gemini-2.5-pro, updated in June 2025) - Gemini-Embedding (
gemini-embedding-001) - GPT-5.1 (
gpt-5.1-2025-11-13) - Sonnet-4.5 (
claude-sonnet-4-5-20250929) - Sonnet-4.6 (
claude-sonnet-4-6)
Open-weight models
- Qwen3 8B (unsloth/Qwen3-8B, updated May 14, 2025)
- Qwen3 32B (unsloth/Qwen3-32B, updated May 14, 2025)
- Hunyuan-MT-7B (tencent/Hunyuan-MT-7B, updated Sep 18, 2025)
- Qwen3 235B-A22B (Qwen/Qwen3-235B-A22B-Instruct-2507; accessed via Fireworks AI, Aug 2025 β Feb 2026)
- Kimi-K2-Thinking (moonshotai/Kimi-K2-Thinking; accessed via Fireworks AI, Nov 2025 β Feb 2026)
- DeepSeek-V3.1 (deepseek-ai/DeepSeek-V3.1; accessed via Fireworks AI, Aug 2025 β Feb 2026)
