Role: Lead Org: Sharif University of Technology — Speech and Language Processing Lab Distribution: Open dataset on the Hugging Face Hub, under SLPL/naab

Summary

A large, cleaned, ready-to-use open-source Farsi text corpus: ~130GB, 250 million paragraphs, 15 billion words. Used to train the first T5 models for Farsi.

What distinguishes it as an artifact is that the cleaning is already done. The corpus ships plug-and-play, so it can be used as it is. It is released openly, hosted as a Hugging Face dataset at SLPL/naab, which is what makes it usable by people with no connection to the lab that built it.

Data

naab: A Ready-to-Use Plug-and-Play Corpus for Farsi — the paper. XNum — Universal Numeral System Converter — sibling multilingual-text tooling. Open Source Philosophy — the case for releasing artifacts others can rerun.