CoolFace
Datasetpublic

hassan-IA/hassaniya-stories-ocr

Hassaniya Stories OCR Dataset Image-to-text OCR pairs extracted from a Hassaniya Arabic stories corpus. Each row couples an image crop from the source book with its manually cleaned text. The primary columns are image and text. Source The original material was published in 1994 under the title 41 Short Stories About Life in Mauritania, Especially Life in Nouakchott. It was written by K. ould Beye and Dave Penney. The Hassaniya stories were translated into English… See the full description on the dataset page: https://huggingface.co/datasets/hassan-IA/hassaniya-stories-ocr.

sourceHugging Faceupdated 4mo agoView on Hugging Face
4likes55downloads
Dataset Card

Hassaniya Stories OCR Dataset

Image-to-text OCR pairs extracted from a Hassaniya Arabic stories corpus. Each row couples an image crop from the source book with its manually cleaned text.

The primary columns are image and text.

Source

The original material was published in 1994 under the title 41 Short Stories About Life in Mauritania, Especially Life in Nouakchott. It was written by K. ould Beye and Dave Penney. The Hassaniya stories were translated into English and reflect aspects of Mauritanian life, customs, and social habits. The source used for this dataset was scanned from paper documents and made publicly available on the internet.

Dataset size

  • —Rows: 552

Fields

  • —image: source-book row crop.
  • —text: manually cleaned Hassaniya Arabic text.