Authors: Sadra Sabouri, Elnaz Rahmati, Soroush Gooran, Hossein Sameti Venue: Journal of Artificial Intelligence, Applications and Innovations (JAIAI), 2024

Summary

Lower-resource languages like Farsi need large training data, and at the time the work was done there was even less of it around than there is now. naab is the answer to that: an open Farsi corpus that is already cleaned and ready to use — roughly 130GB, 250 million paragraphs, and 15 billion words. The “plug-and-play” in the title is the claim being made: the cleaning has already been done, so the corpus can be used as it is.

The name comes from the Farsi word naab, which means pure and high-grade — which is what the team was going for.

Paper · Data

naab — Farsi Text Corpus — the corpus artifact itself, used to train the first Farsi T5 models. Sharif University of Technology — Speech and Language Processing Lab — where it was built. ParsiPy: NLP Toolkit for Historical Persian Texts in Python — another Persian-language resource paper with the same senior author, Hossein Sameti.