Role: Research Assistant Dates: August 2021 – November 2022 Location: Tehran, Iran PI: Hossein Sameti
Two lines of work, both pre-ChatGPT:
Farsi speech. Developed and open-sourced ASR and TTS systems for Farsi using Wav2Vec2.0, Kaldi x-vectors, and Coqui TTS. Collected and preprocessed Persian speech datasets, fine-tuned models for Farsi-specific linguistic features, and optimized across diverse acoustic conditions. All models, training scripts, and benchmarks released openly to address the scarcity of Farsi resources. Culminated in A Review of the Recent Speech Recognition Methods and Sharif-Wav2Vec2.0 — Farsi Speech Recognition Model.
Farsi text. Led development of a large-scale Farsi corpus used to train the first T5 models for Farsi, and worked as second author on data augmentation for Farsi question answering and document-grounded dialogue. This became the bachelor’s thesis, published as Docalog: Multi-Document Dialogue System Using Transformer-Based Span Retrieval and naab: A Ready-to-Use Plug-and-Play Corpus for Farsi.
Related
Bachelor of Electrical Engineering and Computer Science — Sharif University of Technology — the concurrent degree. SynTran-fa: Generating Comprehensive Answers for Farsi QA Pairs via Syntactic Transformation — later output from the same QA line.