Authors: Farhan Farsi, Parnian Fazel, Farzaneh Goshtasb, Nadia Hajipour, Sadra Sabouri, Ehsaneddin Asgari, Hossein Sameti Venue: Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT) at NAACL, 2025

Summary

Pahlavi (Middle Persian) barely exists in digital form, and a language with no data is a language on its way out. That is the stake the paper works from: the missing resources are not an inconvenience for researchers, they are the thing putting the language at risk.

PahGen is a framework that translates English into Pahlavi by pairing grammar-guided term extraction with zero-shot LLM translation, so that the generated sentences come out correct both syntactically and semantically. The framework was then used to produce a dataset of 360 parallel English–Pahlavi texts, all of them validated by experts.

Paper

ParsiPy: NLP Toolkit for Historical Persian Texts in Python — same team, historical Persian toolkit.