CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaanhho /refusal-exp031-statetabularn<1K1 likes1.8k downloads2mo agoHugging Face02KaanAydinli /tsc-tr-filtered-94h-clean TSC-TR Filtered 94h — repaired transcripts ~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified. Transcript repairs The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.audioautomatic-speech-recognition10K<n<100K0 likes437 downloads1mo agoHugging Face03MahtaFetrat /KaamelDict Kaamel-Dict: A Comprehensive Persian G2P Dictionary Kaamel-Dict is the largest publicly available Persian grapheme-to-phoneme (G2P) dictionary, containing over 116,600 entries. It was developed by unifying multiple phonetic representation systems from existing G2P tools [1] [2] [3] [4] and datasets [5] [6] [7] [8] [9] and an additional online glossary [10] , called Jame` Dictionary, . This dictionary is released under the GNU license, allowing it to be used for developing G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/KaamelDict.texttranslation100K<n<1M15 likes199 downloads1y agoHugging Face04kaab4321 /UrduTTS UrduTTS A Studio-Quality Urdu Speech Corpus with Urdu, Phonemized, and Romanized Transcriptions 91.9 hours · 57,873 utterances · 44.1 kHz · 3 aligned text representations Dataset Summary UrduTTS is the largest openly available Urdu text-to-speech corpus with three aligned text representations for every utterance: native Urdu script, Phonemized (IPA) text, and Romanized (Latin) text. Urdu is spoken by roughly 250 million people but is badly… See the full description on the dataset page: https://huggingface.co/datasets/kaab4321/UrduTTS.text-to-speech10K<n<100K0 likes194 downloads12d agoHugging Face05kaanrkaraman /code2doc Code2Doc: Function-Documentation Pairs Dataset A curated dataset of 13,358 high-quality function-documentation pairs extracted from popular open-source repositories on GitHub. Designed for training models to generate documentation from code. Dataset Description This dataset contains functions paired with their docstrings/documentation comments from 5 programming languages, extracted from well-maintained, highly-starred GitHub repositories. Languages Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kaanrkaraman/code2doc.tabulartext-generation10K<n<100K0 likes139 downloads9mo agoHugging Face06kaa-ml /kknews-dataset KKNews.uz Dataset Qaraqalpaqstan Xabar Agentligi (kknews.uz) maqalaları — 5 tilde. Languages Code Language ru Russian uz Uzbek (Latin) oz Uzbek (Cyrillic) kk Karakalpak (Cyrillic) qq Karakalpak (Latin) Columns Column Type Description id int WordPress post ID lang string Language code category_id int Category ID category_name string Category name title string Plain text title content_html string Original… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/kknews-dataset.tabulartext-classification10K<n<100K1 likes99 downloads21d agoHugging Face07kaa0227 /letvideon<1K0 likes93 downloads2mo agoHugging Face08kaa-ml /shipaker-dataset Shipaker.uz Dataset A collection of health and medicine articles in the Karakalpak language, scraped from shipaker.uz. The site is run by a public health organization in Karakalpakstan, Uzbekistan, and publishes articles on topics such as disease prevention, nutrition, psychology, and general wellness. Columns Column Type Description id int Article ID title string Article title content_html string Full article body (original HTML) content_text… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/shipaker-dataset.texttext-classificationn<1K1 likes68 downloads21d agoHugging Face09kaab4321 /kaab_wrkfrmhme4_50 likes67 downloads1y agoHugging Face10pertai /eesti-kaanamiskorpus eesti-kaanamiskorpus — Estonian inflection corpus (11,011 entries) Word and phrase inflections with case, number, source and licence per entry. Parallel forms kept as separate entries (crucial for fair scoring!). Built from Riigikogu stenograms + ERR (CC-BY-SA) and Vabamorf rule-based synthesis with round-trip validation. On licences, stated plainly. 11,011 of 13,436 entries are published here; entries with unresolved rights are withheld from this dataset. Withheld from… See the full description on the dataset page: https://huggingface.co/datasets/pertai/eesti-kaanamiskorpus.text10K<n<100K0 likes53 downloads21d agoHugging Face11kaanhho /refusal-exp047-extraction-position0 likes48 downloads2mo agoHugging Face12kaan39 /turkish-wikipedia-dataset-clean Turkish Wikipedia Dataset A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments. The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata. Dataset Summary Language: Turkish (tr) Source: Turkish Wikipedia Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.tabular100K<n<1M0 likes45 downloads1mo agoHugging Face13kaa-ml /paziylet-dataset-v1 Paziylet Dataset V1 Posts from the @paziyletuz Telegram channel in the Karakalpak language. Columns Column Type Description id int Original message ID from @paziyletuz channel text string Post text (social media footer stripped) date string Original post date (ISO 8601) url string Direct link to the Telegram post Usage from datasets import load_dataset ds = load_dataset("kaa-ml/paziylet-dataset-v1", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/paziylet-dataset-v1.texttext-generationn<1K0 likes42 downloads1mo agoHugging Face14nickoo004 /kaa-parallel-corpus Kaa Karakalpak-English Parallel Corpus (FineTranslations) 📌 Overview This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan. This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.tabulartranslation10K<n<100K0 likes41 downloads5mo agoHugging Face15kaamd /jo-asraudio10K<n<100K0 likes36 downloads1y agoHugging Face16MorishT /KA__allenai_openbookqa__openai.gpt-3.5-turbo-0125__5_5_5textn<1K0 likes33 downloads2y agoHugging Face17kaahila /sugarcrm_130_documentation Source: Sugarcrm 13.0 Dev Documentation The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file cl100k_base: 400 Tokens per chunk p50k_base: 200 Tokens per chunk textquestion-answering1K<n<10K0 likes32 downloads3y agoHugging Face18kaarthu2003 /TraditionalDataset4v5audio10K<n<100K0 likes30 downloads1y agoHugging Face19Finnish-NLP /ultrachat_dpo_sft_deepl_kaannetty Dataset Card for Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty This dataset is more filtered down version of Finnish-NLP/ultrafeedback_deepl_sft_dpo_filtered Creation process Load data from https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized/viewer/default/train_sft Do zero shot classification with facebook/bart-large-mnli in this kind of way (Actual implementation might be slightly different): preds = pipe(f'{row["instruction"]} is a question about:'… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty.tabular10K<n<100K1 likes26 downloads10mo agoHugging Face20kaa-ml /karakalpak-audio-datasetgatedaudio1K<n<10K1 likes26 downloads1y agoHugging Face21MorishT /KA__allenai_ai2_arc__openai.gpt-3.5-turbo-0125__5_5_5textn<1K0 likes24 downloads2y agoHugging Face22TamilThagaval /aimperum_kaappiyangal-seevaga_chintamani 📕 Sivaga Chintamani Dataset (சீவக சிந்தாமணி தரவுத்தொகுப்பு) 🧾 Dataset Summary Sivaga Chintamani (சீவக சிந்தாமணி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics) and is considered the earliest epic chronologically among them. The epic was composed in Tamil by adapting several Sanskrit Sivagan legends. The original source is believed to be a work known as “Kshatriya Chudamani”. This dataset presents a structured digital version of Sivaga Chintamani… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-seevaga_chintamani.texttable-question-answeringn<1K0 likes24 downloads9mo agoHugging Face23KaanGoker /dactylic-hexameter-latin-poetry-corpus Dactylic Hexameter Latin Poetry Corpus This repository contains a curated and processed corpus of Classical Latin poetry written in dactylic hexameter. It serves as the raw training data ("Dataset V3") for the Master's Thesis titled "A Hybrid Post Hoc Feedback Framework for Latin Dactylic Hexameter" submitted to KU Leuven (2025). Dataset Description This corpus was constructed to fine-tune Large Language Models (LLMs) for the generation of metrically valid Latin poetry.… See the full description on the dataset page: https://huggingface.co/datasets/KaanGoker/dactylic-hexameter-latin-poetry-corpus.texttext-generation10K<n<100K0 likes24 downloads9mo agoHugging Face24kaaniince /turkishReviews-ds-textGeneration Dataset Card for "turkishReviews-ds-textGeneration" More Information needed text1K<n<10K0 likes23 downloads3y agoHugging Face25kaaath-i /kochwiki-ir-datageospatial0 likes23 downloads6mo agoHugging Face26kaans /medqa-swe-with-responsestabular1K<n<10K0 likes22 downloads1y agoHugging Face27kaaeaate /amd0 likes21 downloads4mo agoHugging Face28kaab4321 /Qualityaudion<1K0 likes20 downloads1y agoHugging Face29KaalTech /salad_recipes0 likes19 downloads2y agoHugging Face30kaarthu2003 /IEEEAccessDatasetSLRVoicesaudio10K<n<100K0 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.