datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-exp031-statetsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz
mono WAV) with systematically repaired transcripts. This is a derivative of
ulaspolat/tsc-tr-filtered-94h,
itself a filtered subset of the ISSAI Turkish Speech Corpus
(MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe
mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.KaamelDict
Kaamel-Dict: A Comprehensive Persian G2P Dictionary
Kaamel-Dict is the largest publicly available Persian grapheme-to-phoneme (G2P) dictionary, containing over 116,600 entries.
It was developed by unifying multiple phonetic representation systems from existing G2P tools
[1]
[2]
[3]
[4]
and datasets
[5]
[6]
[7]
[8]
[9]
and an additional online glossary
[10]
, called Jame` Dictionary, . This dictionary is released under the GNU license, allowing it to be used for
developing G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/KaamelDict.UrduTTS
UrduTTS
A Studio-Quality Urdu Speech Corpus with Urdu, Phonemized, and Romanized Transcriptions
91.9 hours · 57,873 utterances · 44.1 kHz · 3 aligned text representations
Dataset Summary
UrduTTS is the largest openly available Urdu text-to-speech corpus with three aligned text
representations for every utterance: native Urdu script, Phonemized (IPA) text, and
Romanized (Latin) text.
Urdu is spoken by roughly 250 million people but is badly… See the full description on the dataset page: https://huggingface.co/datasets/kaab4321/UrduTTS.code2doc
Code2Doc: Function-Documentation Pairs Dataset
A curated dataset of 13,358 high-quality function-documentation pairs extracted from popular open-source repositories on GitHub. Designed for training models to generate documentation from code.
Dataset Description
This dataset contains functions paired with their docstrings/documentation comments from 5 programming languages, extracted from well-maintained, highly-starred GitHub repositories.
Languages Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kaanrkaraman/code2doc.kknews-dataset
KKNews.uz Dataset
Qaraqalpaqstan Xabar Agentligi (kknews.uz) maqalaları — 5 tilde.
Languages
Code
Language
ru
Russian
uz
Uzbek (Latin)
oz
Uzbek (Cyrillic)
kk
Karakalpak (Cyrillic)
qq
Karakalpak (Latin)
Columns
Column
Type
Description
id
int
WordPress post ID
lang
string
Language code
category_id
int
Category ID
category_name
string
Category name
title
string
Plain text title
content_html
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/kknews-dataset.letshipaker-dataset
Shipaker.uz Dataset
A collection of health and medicine articles in the Karakalpak language, scraped from shipaker.uz. The site is run by a public health organization in Karakalpakstan, Uzbekistan, and publishes articles on topics such as disease prevention, nutrition, psychology, and general wellness.
Columns
Column
Type
Description
id
int
Article ID
title
string
Article title
content_html
string
Full article body (original HTML)
content_text… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/shipaker-dataset.kaab_wrkfrmhme4_5eesti-kaanamiskorpus
eesti-kaanamiskorpus — Estonian inflection corpus (11,011 entries)
Word and phrase inflections with case, number, source and licence per entry.
Parallel forms kept as separate entries (crucial for fair scoring!). Built
from Riigikogu stenograms + ERR (CC-BY-SA) and Vabamorf rule-based synthesis
with round-trip validation.
On licences, stated plainly. 11,011 of 13,436 entries are published here;
entries with unresolved rights are withheld from this dataset. Withheld from… See the full description on the dataset page: https://huggingface.co/datasets/pertai/eesti-kaanamiskorpus.refusal-exp047-extraction-positionturkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.paziylet-dataset-v1
Paziylet Dataset V1
Posts from the @paziyletuz Telegram channel in the Karakalpak language.
Columns
Column
Type
Description
id
int
Original message ID from @paziyletuz channel
text
string
Post text (social media footer stripped)
date
string
Original post date (ISO 8601)
url
string
Direct link to the Telegram post
Usage
from datasets import load_dataset
ds = load_dataset("kaa-ml/paziylet-dataset-v1", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/paziylet-dataset-v1.kaa-parallel-corpus
Kaa Karakalpak-English Parallel Corpus (FineTranslations)
📌 Overview
This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan.
This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.jo-asrKA__allenai_openbookqa__openai.gpt-3.5-turbo-0125__5_5_5sugarcrm_130_documentation
Source: Sugarcrm 13.0 Dev Documentation
The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file
cl100k_base: 400 Tokens per chunk
p50k_base: 200 Tokens per chunk
TraditionalDataset4v5ultrachat_dpo_sft_deepl_kaannetty
Dataset Card for Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty
This dataset is more filtered down version of Finnish-NLP/ultrafeedback_deepl_sft_dpo_filtered
Creation process
Load data from https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized/viewer/default/train_sft
Do zero shot classification with facebook/bart-large-mnli in this kind of way (Actual implementation might be slightly different):
preds = pipe(f'{row["instruction"]} is a question about:'… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty.karakalpak-audio-datasetKA__allenai_ai2_arc__openai.gpt-3.5-turbo-0125__5_5_5aimperum_kaappiyangal-seevaga_chintamani
📕 Sivaga Chintamani Dataset (சீவக சிந்தாமணி தரவுத்தொகுப்பு)
🧾 Dataset Summary
Sivaga Chintamani (சீவக சிந்தாமணி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics) and is considered the earliest epic chronologically among them.
The epic was composed in Tamil by adapting several Sanskrit Sivagan legends. The original source is believed to be a work known as “Kshatriya Chudamani”.
This dataset presents a structured digital version of Sivaga Chintamani… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-seevaga_chintamani.dactylic-hexameter-latin-poetry-corpus
Dactylic Hexameter Latin Poetry Corpus
This repository contains a curated and processed corpus of Classical Latin poetry written in dactylic hexameter. It serves as the raw training data ("Dataset V3") for the Master's Thesis titled "A Hybrid Post Hoc Feedback Framework for Latin Dactylic Hexameter" submitted to KU Leuven (2025).
Dataset Description
This corpus was constructed to fine-tune Large Language Models (LLMs) for the generation of metrically valid Latin poetry.… See the full description on the dataset page: https://huggingface.co/datasets/KaanGoker/dactylic-hexameter-latin-poetry-corpus.turkishReviews-ds-textGeneration
Dataset Card for "turkishReviews-ds-textGeneration"
More Information needed
kochwiki-ir-datamedqa-swe-with-responsesamdQualitysalad_recipesIEEEAccessDatasetSLRVoices
