CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kz-transformers /multidomain-kazakh-dataset Dataset Description Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk Dataset Summary MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains. Supported Tasks 'MLM/CLM': can be used to train a model for casual and masked languange modeling Languages The kk code for Kazakh as generally spoken in the Kazakhstan Data Instances For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.texttext-generation10M<n<100M30 likes296 downloads1y agoHugging Face02dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes132 downloads23d agoHugging Face03bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads14d agoHugging Face04khazarai /Multi-Domain-Reasoning-Benchmark Comprehensive Multi-Domain Reasoning Benchmark (CMDR-Bench) A systematic evaluation suite comprising 100 meticulously curated test cases across 10 distinct cognitive domains, designed to assess Large Language Models' capabilities in reasoning, problem-solving, and instruction-following. Each domain features a graduated difficulty scale (Levels 1–10), enabling fine-grained analysis of capability thresholds from elementary to expert-level complexity. texttext-generationn<1K3 likes54 downloads6mo agoHugging Face05EnDevSols /Multi-Domain-Reasoning-SFT Multi-Domain-Reasoning-SFT Dataset Summary The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving. Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.texttext-generation100K<n<1M0 likes45 downloads5mo agoHugging Face06Pankaj8922 /multidomain-complex-text-pool Complex Text Pool Dataset Overview A curated collection of complex, long-form English texts sampled from 9 diverse domains. Each document has been truncated to a maximum of 4,000 characters, preserving clean sentence boundaries. The dataset is designed to provide challenging, real-world text samples across multiple subject areas. Categories and Sample Counts Category Samples news 9,999 encyclopedic 10,000 conversational 10,000… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/multidomain-complex-text-pool.texttext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face07Pangeanic /Iraqi-Arabic-multidomain-QA-text Iraqi Arabic Multidomain QA Dataset The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models. This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.textquestion-answeringn<1K1 likes27 downloads4mo agoHugging Face08sabin1234 /NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.texttext-generation100K<n<1M0 likes24 downloads1mo agoHugging Face09tejasashinde /a2z-multidomain-glossary A–Z Multi-Domain Glossary Dataset This dataset is a creative collection of A-to-Z terminology across a wide range of high-level domains including Agriculture, Technology, Environment, Artificial Intelligence, Zoology, and more.Each entry includes: domain letter (A–Z) word description (short) 📊 Structure Column Description domain The high-level category (e.g. Technology, Agriculture) letter The alphabetical letter from A to Z word The concept/keyword… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/a2z-multidomain-glossary.texttext-generationn<1K1 likes22 downloads1y agoHugging Face10grpathak22 /marathi-maharashtra-multidomain-SFT-1kgated Marathi Maharashtra Multidomain SFT - 1K Sample Dataset Description This is a carefully curated 1,000-sample subset of the comprehensive Marathi-Maharashtra multidomain supervised fine-tuning (SFT) dataset. This high-quality dataset contains question-answer pairs covering diverse aspects of Marathi language, culture, history, and Maharashtra-related topics. Key Features High-Quality Human Verification: All responses have been verified by Marathi language… See the full description on the dataset page: https://huggingface.co/datasets/grpathak22/marathi-maharashtra-multidomain-SFT-1k.textquestion-answeringn<1K0 likes7 downloads8mo agoHugging Face11shkomq /Kurdish_Multi-Domain_Corpus_KMDCgated Kurdish Multi-Domain Corpus (KMDC) Dataset Description The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.textquestion-answering100K<n<1M0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.