datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.azerbaijani-ocr-lines
Azerbaijani OCR Lines
Line-level training data for Azerbaijani text recognition, in Latin and
Cyrillic script, extracted from scanned books.
Fields
field
description
image
cropped text line, grayscale, height 48 px
text
transcription
script
az_latin or az_cyrillic
book
anonymised source-book id
How it was built
Pages come from scanned PDFs that already carried an OCR text layer. Line
boxes were taken from that layer, rendered… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-lines.azerbaijani-ocr-benchmark
Azerbaijani OCR Benchmark
Line-level OCR benchmark for Azerbaijani in both Latin and Cyrillic script,
built from scanned books.
Fields
field
description
image
cropped text line, grayscale, height 48 px
text
verbatim transcription
script
az_latin or az_cyrillic
book
anonymised source-book id
How labels were produced
Every line carries a label agreed on independently by three sources: the
OCR text layer already present in the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-benchmark.azerbaijani-htr-synthetic
Azerbaijani Synthetic Handwritten OCR Dataset
A large-scale synthetic dataset for training handwritten text recognition (HTR) models on Azerbaijani Latin script. Generated using a procedural pipeline that combines real-world handwriting fonts with realistic scan-style augmentations.
This dataset addresses the lack of publicly available Azerbaijani handwriting OCR data — a low-resource language for which no IAM-equivalent corpus exists.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-htr-synthetic.azerbaijani-carpet-fullsizeazerbaijani-cuisine
Azerbaijani Cuisine Dataset
A curated image dataset of traditional Azerbaijani dishes for computer vision and image classification tasks.
Dataset Description
This dataset contains images of five traditional Azerbaijani dish categories. It is organized into standard training, validation, and test splits to facilitate machine learning model development and evaluation.
Features
5 Food Categories: Dolma, Kebabs, Pakhlava, Plov, and Soups
324 Total Images: Properly… See the full description on the dataset page: https://huggingface.co/datasets/ARMammadli/azerbaijani-cuisine.Wildfire_Images_From_Satilite_For_Azerbaijanazerbaijani-htr-benchmark
Azerbaijani Handwritten OCR Benchmark
A manually annotated benchmark for handwritten text recognition (HTR) on Azerbaijani Latin script. Real-world scanned pages annotated in Label Studio with rotated bounding boxes and exact transcriptions.
Provided in two parallel views:
lines — Line-level recognition
Cropped images of single text lines paired with their transcription. Rotated regions are deskewed (warped to be axis-aligned) so each crop shows the line horizontally.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-htr-benchmark.azerbaijani-art-collectionNote: Data is collected by the Afina Apayeva, Ariana Kenbayeva, Ilhama Novruzova, Mehriban Aliyeva, and only art_metal category is taken by scraping Azerbaijan Carpet Museum's official website.
Dataset Source:We took the pictures of art works by smartphones. For some of them, we took their pictures from 3 perspectives: left, right, and front. For most of them, we took just one picture from the front side to avoid data duplication.
Primary photos taken by team members at:
Azerbaijan National… See the full description on the dataset page: https://huggingface.co/datasets/inovruzova/azerbaijani-art-collection.azerbaijan-landmarks-dataset
📂 Dataset Name
Azerbaijan Landmarks Dataset
Data Collection
The dataset includes the following five classes:
Maiden Tower
Heydar Aliyev Center
Dede Gorgud Park
Deniz Mall
Palace of the Shirvanshahs
Structure
train/: Training images (~80% of data, excluding validation samples). Since data samples inside train is too big, we divided train_"class_name"
test/: Test images (~20%)
Format
Each image is in .jpg format, organized in class-labeled… See the full description on the dataset page: https://huggingface.co/datasets/khaleed-mammad/azerbaijan-landmarks-dataset.azerbaijani_text_news1azerbaijani-carpet
