datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multidomain-kazakh-dataset
Dataset Description
Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk
Dataset Summary
MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains.
Supported Tasks
'MLM/CLM': can be used to train a model for casual and masked languange modeling
Languages
The kk code for Kazakh as generally spoken in the Kazakhstan
Data Instances
For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.Multi-Domain-Reasoning-Benchmark
Comprehensive Multi-Domain Reasoning Benchmark (CMDR-Bench)
A systematic evaluation suite comprising 100 meticulously curated test cases across 10 distinct cognitive domains, designed to assess Large Language Models' capabilities in reasoning, problem-solving, and instruction-following. Each domain features a graduated difficulty scale (Levels 1–10), enabling fine-grained analysis of capability thresholds from elementary to expert-level complexity.
Multi-Domain-Reasoning-SFT
Multi-Domain-Reasoning-SFT
Dataset Summary
The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving.
Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.multidomain-complex-text-pool
Complex Text Pool Dataset
Overview
A curated collection of complex, long-form English texts sampled from 9 diverse domains. Each document has been truncated to a maximum of 4,000 characters, preserving clean sentence boundaries. The dataset is designed to provide challenging, real-world text samples across multiple subject areas.
Categories and Sample Counts
Category
Samples
news
9,999
encyclopedic
10,000
conversational
10,000… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/multidomain-complex-text-pool.Iraqi-Arabic-multidomain-QA-text
Iraqi Arabic Multidomain QA Dataset
The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models.
This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.a2z-multidomain-glossary
A–Z Multi-Domain Glossary Dataset
This dataset is a creative collection of A-to-Z terminology across a wide range of high-level domains including Agriculture, Technology, Environment, Artificial Intelligence, Zoology, and more.Each entry includes:
domain
letter (A–Z)
word
description (short)
📊 Structure
Column
Description
domain
The high-level category (e.g. Technology, Agriculture)
letter
The alphabetical letter from A to Z
word
The concept/keyword… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/a2z-multidomain-glossary.marathi-maharashtra-multidomain-SFT-1k
Marathi Maharashtra Multidomain SFT - 1K Sample
Dataset Description
This is a carefully curated 1,000-sample subset of the comprehensive Marathi-Maharashtra multidomain supervised fine-tuning (SFT) dataset. This high-quality dataset contains question-answer pairs covering diverse aspects of Marathi language, culture, history, and Maharashtra-related topics.
Key Features
High-Quality Human Verification: All responses have been verified by Marathi language… See the full description on the dataset page: https://huggingface.co/datasets/grpathak22/marathi-maharashtra-multidomain-SFT-1k.Kurdish_Multi-Domain_Corpus_KMDC
Kurdish Multi-Domain Corpus (KMDC)
Dataset Description
The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.
