AI UNLOCKS HÁN NÔM HERITAGE: OPENING THE DOOR TO A MILLENNIUM OF KNOWLEDGE FOR WIDER PUBLIC ACCESS

AI UNLOCKS HÁN NÔM HERITAGE: OPENING THE DOOR TO A MILLENNIUM OF KNOWLEDGE FOR WIDER PUBLIC ACCESS

After more than a thousand years preserving the history, culture and knowledge of the Vietnamese nation, much of the Hán Nôm documentary heritage remains inaccessible to the wider public. By successfully developing an artificial intelligence (AI)-powered Hán Nôm automatic translation system, researchers at HCMUS are steadily opening the door to this invaluable repository of knowledge, while laying the groundwork for the digital preservation and broader utilisation of Viet Nam’s cultural heritage.

AI opens new pathways to Hán Nôm heritage

According to the Association for Preservation of Hán Nôm Heritage, more than 90% of Hán Nôm documents have yet to be translated into Quốc ngữ (the modern Vietnamese writing system based on the Latin alphabet). Millions of pages of royal decrees, imperial records, genealogies, cadastral registers, stone inscriptions, ancient manuscripts, horizontal lacquered boards, parallel sentences and traditional medical texts remain available primarily as scanned images or digitised copies, while the number of people able to read Hán Nôm continues to decline.

To address this challenge, a research team led by Associate Professor Đinh Điền, Director of the Centre for Computational Linguistics at HCMUS, has carried out the project Research and Development of an Automatic Hán Nôm-to-Quốc ngữ Translation System – Phase 2. The project has recently been evaluated as Outstanding by the Acceptance Council of the Ho Chi Minh City Department of Science and Technology.

Associate Professor Đinh Điền explained: “The Nôm script was created by Vietnamese scholars around the 10th century based on Chinese characters and remained in use for more than a millennium. An enormous body of literature, historical records, geographical works, traditional medicine and many other fields was recorded using this writing system. Today, however, very few people are able to read Hán Nôm, while widely used AI systems such as ChatGPT, Gemini and DeepSeek can process modern Chinese characters but cannot interpret the Nôm script.”

Homepage of the Nôm transliteration website: https://tools.clc.hcmus.edu.vn/

Without timely technological solutions for recognising and converting historical documents, many valuable materials risk continuing deterioration due to ageing, storage conditions and the diminishing number of Hán Nôm specialists. The research team aims to develop a system that enables anyone to access, search and understand Hán Nôm documents conveniently and free of charge.

A major breakthrough in Phase 2 has been the successful development of optical character recognition (OCR) technology for Hán Nôm documents. While the first phase could process only text-based digital documents, the current system can recognise documents from images, which represent the vast majority of surviving Hán Nôm materials.

Another significant achievement has been the creation of the largest Hán Nôm dataset ever assembled in Viet Nam. The dataset includes 200,000 scanned pages of Hán Nôm documents, 200,000 manually annotated images for OCR training, 750,000 bilingual Hán Nôm–Quốc ngữ sentence pairs, and more than one million Hán Nôm monolingual sentences, comprising approximately 16 million characters.

The dataset has been developed through collaboration with numerous domestic and international partners, including the Nom Foundation, the Tran Nhan Tong Institute, several Hán Nôm research projects, together with scholars, lecturers and students specialising in Hán Nôm studies. Most source materials were originally available only in raw form, requiring the research team to establish comprehensive workflows for data standardisation, annotation and cross-validation by both AI models and domain experts to ensure high-quality training data.

Towards a shared digital platform for Hán Nôm heritage

Building upon the dataset and AI models, the research team has developed the Kim Hán Nôm ecosystem, available through web, Android and iOS platforms. The system can recognise Hán Nôm text from images, perform two-way transliteration between Hán Nôm and Quốc ngữ, support translation of Classical Chinese texts, and provide APIs for integration into other applications. For outdoor inscriptions such as horizontal lacquered boards and parallel sentences displayed in temples, pagodas and historical sites, the system can overlay the corresponding Quốc ngữ text directly onto the original image, allowing users to compare the original script with the transliteration.

According to Associate Professor Đinh Điền, system performance depends largely on the quality of the input data. For text-based documents covering widely represented fields such as literature, history and geography, transliteration accuracy exceeds 99%. For high-quality scanned images, particularly those printed in the Khải script, OCR accuracy exceeds 95%. More challenging materials, including heavily weathered stone inscriptions, handwritten texts and unusual calligraphic styles, still require expert review and correction by Hán Nôm specialists.

The system has been designed not only for researchers but also for the wider community. Members of the public can use a mobile phone to photograph horizontal lacquered boards, parallel sentences, royal decrees, genealogies, contracts or other historical family documents, allowing the system to recognise the original text automatically and convert the content into Quốc ngữ. As a result, historical materials that were previously difficult to access can become significantly easier to understand and explore.

Associate Professor Đinh Điền, Director of the Centre for Computational Linguistics at HCMUS, leads the research team undertaking the project “Research and Development of an Automatic Hán Nôm-to-Quốc ngữ Translation System – Phase 2”.

Associate Professor Đinh Điền believes the system will become a valuable resource for universities, libraries, museums, archival institutions and heritage conservation organisations, supporting document recognition, classification, summarisation and retrieval. Looking further ahead, the technology is expected to enable new research opportunities across history, geography, traditional medicine and studies related to Viet Nam’s maritime sovereignty.

Following the completion of Phase 2, the research team is preparing to launch Phase 3, focusing on semantic translation of Hán texts. This stage is regarded as the most demanding because AI must move beyond language processing to incorporate knowledge of history, culture, religion and numerous specialised disciplines. At the same time, the team will continue expanding the dataset across subject areas including history, religion, traditional medicine, stone inscriptions and royal decrees, providing the foundation for increasingly specialised AI models.

Leave a Reply

Your email address will not be published.