Abstract
The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu scribal print and handwritten (manuscript) documents. Comprising nearly 1,500 page images across 208 documents sourced from 30 repositories worldwide, the dataset spans seven languages, around 20 unique and mixed script types, and a chronological range from the 10th to the 20th century. This data set is available, open access, for all on Zenodo. This article explains the different types of data in this large dataset and how this data was compiled and verified and suggests potential use cases for it, such as the training and evaluation of new print and handwritten transcription models.
| Original language | English (US) |
|---|---|
| Article number | 69 |
| Journal | Journal of Open Humanities Data |
| Volume | 12 |
| DOIs | |
| Publication status | Published - 2026 |
| Externally published | Yes |
Keywords
- Arabic
- HTR
- OCR
- Ottoman Turkish
- Persian
- Urdu
Fingerprint
Dive into the research topics of 'OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver