datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QA_3sualaz_on_Azerbaijani Question-Answering Dataset for Azerbaijani Language based on Intellectual Games (3sual.az)
Baku Higher Oil School Research and Development Center on AI introduces a dataset to fine-tune the NLP models to manage it as a question answering. This dataset contains 4697 questions with answers and explanations. In some cases answer does not exist therefore that slot is empty. Dataset have been collected from 3sual.az and copyright belongs to corresponding website (3sual.az) and its owner Bahruz… See the full description on the dataset page: https://huggingface.co/datasets/BHOSAI/QA_3sualaz_on_Azerbaijani.azerbaijani_asr
Azerbaijani ASR Dataset
Dataset Description
This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks.
Dataset Summary
Language: Azerbaijani (az)
Task: Automatic Speech Recognition
Total Duration: ~328 hours
Total Samples: ~345,643 audio-text pairs
Audio Format: WAV, 16kHz sampling rate
License: CC-BY-4.0
Dataset Structure
Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.azerbaijani-speech-datasetazerbaijani-pretrain-corpus
Azerbaijani Pretraining Corpus (merged & deduplicated)
A cleaned Azerbaijani text corpus assembled for language-model pretraining,
merging two curated sources and removing exact duplicates.
Contents
Documents: 6,931,898
Tokens: ~5.36B (measured with the o200k_base tokenizer; an
Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base
segments agglutinative Azerbaijani inefficiently)
Avg tokens/document: ~773
Fields
text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.court_cases_azerbaijani
Court Cases Of The Republic Of Azerbaijan
This dataset consists of court cases from the Republic of Azerbaijan.
Overview
It was formed based on 1,200,000 court cases.
The data has been preliminarily normalized and split into sentences.
The dataset consists of 37 million sentences and approximately 500-600 million tokens.
Dataset Structure
Each row represents a single sentence extracted from a court case document.
Column
Type
Description
case_id… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/court_cases_azerbaijani.azerbaijani-ocr-lines
Azerbaijani OCR Lines
Line-level training data for Azerbaijani text recognition, in Latin and
Cyrillic script, extracted from scanned books.
Fields
field
description
image
cropped text line, grayscale, height 48 px
text
transcription
script
az_latin or az_cyrillic
book
anonymised source-book id
How it was built
Pages come from scanned PDFs that already carried an OCR text layer. Line
boxes were taken from that layer, rendered… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-lines.azerbaijani-tts-datasetazerbaijani-ocr-benchmark
Azerbaijani OCR Benchmark
Line-level OCR benchmark for Azerbaijani in both Latin and Cyrillic script,
built from scanned books.
Fields
field
description
image
cropped text line, grayscale, height 48 px
text
verbatim transcription
script
az_latin or az_cyrillic
book
anonymised source-book id
How labels were produced
Every line carries a label agreed on independently by three sources: the
OCR text layer already present in the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-benchmark.azerbaijani-htr-synthetic
Azerbaijani Synthetic Handwritten OCR Dataset
A large-scale synthetic dataset for training handwritten text recognition (HTR) models on Azerbaijani Latin script. Generated using a procedural pipeline that combines real-world handwriting fonts with realistic scan-style augmentations.
This dataset addresses the lack of publicly available Azerbaijani handwriting OCR data — a low-resource language for which no IAM-equivalent corpus exists.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-htr-synthetic.fleurs-azerbaijani-asr
FLEURS Azerbaijani ASR Benchmark
Azerbaijani (az_az) subset of FLEURS,
reformatted for ASR benchmarking and fine-tuning.
Source
Based on FLEURS dataset by Google (Conneau et al., 2022).
Licensed under CC-BY-4.0.
Structure
Split
Samples
Duration
train
2656
9.28h
dev
400
1.35h
test
921
3.23h
Fields
audio — 16kHz mono WAV
sentence — transcription (original casing and punctuation)
sentence_normalized — normalized (lowercase, no… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/fleurs-azerbaijani-asr.community_oscar_azerbaijani
Community-OSCAR Azerbaijani
This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar.
Dataset Statistics (Aggregate)
Metric
Value
Language
Azerbaijani (az)
Average per release
3.36 GiB, 603,832 documents
Words per release
~408.8M words
Characters per release
~3.12B characters
Total size (all releases)
137.62 GiB
Total lines
24.76M
Total words
16.76B words
Total characters
128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.azerbaijani-asr-zenfira
Dataset Card for "azerbaijani-asr-zenfira"
More Information needed
alpaca-cleaned_AZERBAIJANIazerbaijani-carpet-fullsizeazerbaijani-english-parallel-corpus
Azerbaijani English Parallel Corpus
This dataset contains 4,141,966 pairs of high-quality sentences translated from Azerbaijani to English. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc.
License
CC-BY-4.0
Contact
For more information, questions, or issues, please contact LocalDoc at [v.resad.89@gmail.com].
community_oscar_azerbaijani_scored
Azerbaijani Web Corpus with Quality Scores
This dataset is the full Azerbaijani web corpus
LocalDoc/community_oscar_azerbaijani
with a continuous quality score attached to every document. It is intended
as the filtering layer for building a clean Azerbaijani pretraining corpus:
each document carries a score that lets you keep, clean, or drop it according
to your own thresholds.
What was done
Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.azerbaijani_asr_th_clean4azerbaijani-cuisine
Azerbaijani Cuisine Dataset
A curated image dataset of traditional Azerbaijani dishes for computer vision and image classification tasks.
Dataset Description
This dataset contains images of five traditional Azerbaijani dish categories. It is organized into standard training, validation, and test splits to facilitate machine learning model development and evaluation.
Features
5 Food Categories: Dolma, Kebabs, Pakhlava, Plov, and Soups
324 Total Images: Properly… See the full description on the dataset page: https://huggingface.co/datasets/ARMammadli/azerbaijani-cuisine.azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews.
Dataset contains 3 columns: Content, Score, Upvotes
English-Azerbaijani-Arabic-Script-Parallel-Corpusazerbaijani-asr-news1glue-mrpc-azerbaijaniThis dataset represents a translated version of the GLUE/MRPC dataset, generated using the Google Translate API.
azerbaijani-blogs
Azerbaijani Blogs dataset
Dataset Details
Dataset Description
This dataset provides blogs written in azerbaijani language with categories and tags for each.
Language(s) (NLP): Azerbaijani
License: Apache license 2.0
Data Source
All the data was found in public resources of kayzen.az blogging website without any restriction.
AARA_Azerbaijani_LLM_Benchmark
AARA: Azerbaijani Advanced Reasoning Assessment
This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark.
azerbaijani_books_retriever_corpus-reranked
Azerbaijani Books Retrieval Dataset (Reranked)
A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani.
Dataset Configs
The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.azerbaijani_sa
Sentiment Analysis Data for the Azerbaijani Language
Dataset Description:
This dataset contains a sentiment analysis dataset from LocalDoc (2024).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@source{azerbaijanisent,
title={Sentiment Analysis Datset for Azerbaijani},
author={LocalDoc},
link={https://huggingface.co/LocalDoc},
year={2024}
}
azerbaijani_asr_th_clean6azerbaijani-audiobooksazerbaijani_spelling_dictionary_2021This dataset contains words from the Spelling Dictionary of the Azerbaijani Language, 7th edition, which was published in 2021.
