Skip to main navigation Skip to search Skip to main content

OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data

  • Jonathan Parkes Allen
  • , John Mullan
  • , Lorenz Nigst
  • , Mathew Barber
  • , Taimoor Shahid-Khan
  • , Masoumeh Seydi
  • , Danlu Chen
  • , Yufei Weng
  • , Nikolai Vogler
  • , Jacob Murel
  • , Osama Eshera
  • , Taylor Berg-Kirkpatrick
  • , David Smith
  • , Sarah Bowen Savant
  • , Matthew Thomas Miller

Research output: Contribution to journalArticlepeer-review

Abstract

The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu scribal print and handwritten (manuscript) documents. Comprising nearly 1,500 page images across 208 documents sourced from 30 repositories worldwide, the dataset spans seven languages, around 20 unique and mixed script types, and a chronological range from the 10th to the 20th century. This data set is available, open access, for all on Zenodo. This article explains the different types of data in this large dataset and how this data was compiled and verified and suggests potential use cases for it, such as the training and evaluation of new print and handwritten transcription models.

Original languageEnglish (US)
Article number69
JournalJournal of Open Humanities Data
Volume12
DOIs
Publication statusPublished - 2026
Externally publishedYes

Keywords

  • Arabic
  • HTR
  • OCR
  • Ottoman Turkish
  • Persian
  • Urdu

Fingerprint

Dive into the research topics of 'OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data'. Together they form a unique fingerprint.

Cite this