Skip to content

Latest commit

Β 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 

Repository files navigation

HERIT: Enabling Global Access to Hanja-to-English Historical Translation by Mitigating Temporal Bias

Paper Models Dataset Python

HERIT addresses the data scarcity and temporal bias inherent in historical archives written in Hanja. It combines high-quality pseudo-labeled data augmentation via retrieval-augmented generation (RAG) with two-stage fine-tuning.

Overview of the HERIT pipeline

πŸ“‹ Table of Contents

✨ Highlights

  • Mitigates temporal bias: human evaluations stratified by reign confirm that HERIT translates robustly across historical periods not covered by expert-translated data.
  • Strong empirical results: HERIT outperforms strong baselines on lexical-overlap metrics and achieves a preference rate of about 66% in human expert evaluation.
  • Large-scale pseudo-labeling: 315K RAG-based pseudo-labeled documents for augmentation, and 2.09M final Hanja-to-English translations.
  • Full pipeline released: this repository contains the code used to build HERIT β€” data crawling and preprocessing, pseudo-labeling, fine-tuning, and evaluation.

πŸ“Š Dataset

Note

Work in progress; to be updated by early October 2026.

Dataset # Documents Released Fields
$D^{train}$ 15.7K (Document ID, Hanja)
$D^{valid}$ 1,000 (Document ID, Hanja)
$D^{test}$ 1,000 (Document ID, Hanja)
$D^{test}_{NT}$ 2,080 (Document ID, Hanja)
Human-eval subset 200 (Document ID, Hanja)
$D^{aug}$ 315K (Document ID, Hanja, Pseudo-English)
Final translations 2.09M (Document ID, Hanja, Pseudo-English)

πŸ€– Fine-tuned Models

HERIT model checkpoints (two-stage fine-tuned from Qwen3 8B / 32B). To be released.

Model Base Model Link
HERIT-8B Qwen3-8B TBA
HERIT-32B Qwen3-32B TBA

πŸ—“ Model Versions and Access Dates

Closed-source models (API)
  • Gemini-2.5-Flash (gemini-2.5-flash, updated in June 2025)
  • Gemini-2.5-Pro (gemini-2.5-pro, updated in June 2025)
  • Gemini-Embedding (gemini-embedding-001)
  • GPT-5.1 (gpt-5.1-2025-11-13)
  • Sonnet-4.5 (claude-sonnet-4-5-20250929)
  • Sonnet-4.6 (claude-sonnet-4-6)
Open-weight models

About

a systematic domain-adaptation pipeline that leverages retrieval-augmented generation (RAG) to generate high-quality pseudo-labeled data from abundant Hanja-Korean corpora.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors