Azerbaijani
QA_3sualaz_on_Azerbaijani Question-Answering Dataset for Azerbaijani Language based on Intellectual Games (3sual.az)
Baku Higher Oil School Research and Development Center on AI introduces a dataset to fine-tune the NLP models to manage it as a question answering. This dataset contains 4697 questions with answers and explanations. In some cases answer does not exist therefore that slot is empty. Dataset have been collected from 3sual.az and copyright belongs to corresponding website (3sual.az) and its owner Bahruz… See the full description on the dataset page: https://huggingface.co/datasets/BHOSAI/QA_3sualaz_on_Azerbaijani.azerbaijani_asr
Azerbaijani ASR Dataset
Dataset Description
This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks.
Dataset Summary
Language: Azerbaijani (az)
Task: Automatic Speech Recognition
Total Duration: ~328 hours
Total Samples: ~345,643 audio-text pairs
Audio Format: WAV, 16kHz sampling rate
License: CC-BY-4.0
Dataset Structure
Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.azerbaijani-speech-datasetazerbaijani-pretrain-corpus
Azerbaijani Pretraining Corpus (merged & deduplicated)
A cleaned Azerbaijani text corpus assembled for language-model pretraining,
merging two curated sources and removing exact duplicates.
Contents
Documents: 6,931,898
Tokens: ~5.36B (measured with the o200k_base tokenizer; an
Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base
segments agglutinative Azerbaijani inefficiently)
Avg tokens/document: ~773
Fields
text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.court_cases_azerbaijani
Court Cases Of The Republic Of Azerbaijan
This dataset consists of court cases from the Republic of Azerbaijan.
Overview
It was formed based on 1,200,000 court cases.
The data has been preliminarily normalized and split into sentences.
The dataset consists of 37 million sentences and approximately 500-600 million tokens.
Dataset Structure
Each row represents a single sentence extracted from a court case document.
Column
Type
Description
case_id… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/court_cases_azerbaijani.
