CoolFace
20 results

Azerbaijani

BHOSAI /QA_3sualaz_on_Azerbaijani Question-Answering Dataset for Azerbaijani Language based on Intellectual Games (3sual.az) Baku Higher Oil School Research and Development Center on AI introduces a dataset to fine-tune the NLP models to manage it as a question answering. This dataset contains 4697 questions with answers and explanations. In some cases answer does not exist therefore that slot is empty. Dataset have been collected from 3sual.az and copyright belongs to corresponding website (3sual.az) and its owner Bahruz… See the full description on the dataset page: https://huggingface.co/datasets/BHOSAI/QA_3sualaz_on_Azerbaijani.question-answering1K<n<10K1 likes3k downloads2y agoHugging FaceLocalDoc /azerbaijani_asr Azerbaijani ASR Dataset Dataset Description This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks. Dataset Summary Language: Azerbaijani (az) Task: Automatic Speech Recognition Total Duration: ~328 hours Total Samples: ~345,643 audio-text pairs Audio Format: WAV, 16kHz sampling rate License: CC-BY-4.0 Dataset Structure Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.audioautomatic-speech-recognition100K<n<1M4 likes755 downloads2mo agoHugging Faceshunyalabs /azerbaijani-speech-datasetaudio100K<n<1M1 likes428 downloads1y agoHugging FaceLocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes385 downloads4mo agoHugging FaceLocalDoc /azerbaijani_retriever_corpus A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training Dataset Description This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents. The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.tabularsentence-similarity100K<n<1M0 likes383 downloads1y agoHugging FaceLocalDoc /court_cases_azerbaijani Court Cases Of The Republic Of Azerbaijan This dataset consists of court cases from the Republic of Azerbaijan. Overview It was formed based on 1,200,000 court cases. The data has been preliminarily normalized and split into sentences. The dataset consists of 37 million sentences and approximately 500-600 million tokens. Dataset Structure Each row represents a single sentence extracted from a court case document. Column Type Description case_id… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/court_cases_azerbaijani.text10M<n<100M4 likes148 downloads6mo agoHugging Face