CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.8k downloads2mo agoHugging Face03AmelieSchreiber /toricgt-curated-splits ToricGT Curated Graph Reasoning Splits Curated working dataset repository for ToricGT. The upload contains only curated split Parquet files and metadata generated locally. Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit. Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources. Files train.parquet validation.parquet test.parquet all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.tabulartext-generation1M<n<10M0 likes1.7k downloads4mo agoHugging Face04yaakov /wikipedia-de-splits Dataset Card for yaakov/wikipedia-de-splits Dataset Description The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements. For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, .... The… See the full description on the dataset page: https://huggingface.co/datasets/yaakov/wikipedia-de-splits.text-generationn<1K0 likes605 downloads4y agoHugging Face05Xuhui /sft_processed_large_split sft_processed_large — profile-disjoint split This is the train / val / test split of Xuhui/sft_processed_large, the OdysSim midtraining corpus (21.4M interactions across 63 datasets). Split structure split rows how it's built train 21.20M what's left after val + test are carved out val 28K per-dataset random sample, in-distribution; for checkpoint selection test 128K profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.texttext-generation10M<n<100M1 likes599 downloads5mo agoHugging Face06WitchesSocialStream-Clover-B /Four-Leaf-Clover-Hyper-Split-02 Pausing & Resting Four Leaf Clover Dataset Hi! KaraKaraWitch here. For the past couple of months, I've been collecting 4chan.org posts. This used to include /r/. Fast forward to 16 Aug, I've noticed 4chan has been kind of flaky and throwing some errors at crawl time. It's was bout' time I take a pause to rework the crawller. Additionally it has came to my attention that some images in /r/ contained NCII as referenced in Openmeasures.io. For this reason, I'll be stopping 4chan… See the full description on the dataset page: https://huggingface.co/datasets/WitchesSocialStream-Clover-B/Four-Leaf-Clover-Hyper-Split-02.text-classification3 likes571 downloads3h agoHugging Face07vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split UltraData-Math L3 Textbook Exercise Synthetic Split Source dataset: openbmb/UltraData-Math Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic Each row contains: uid question answer The original content field was split using the literal markers The exercise: and The solution:. texttext-generation10M<n<100M1 likes496 downloads6mo agoHugging Face08jwkirchenbauer /fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.tabulartext-generation100K<n<1M0 likes369 downloads7mo agoHugging Face09abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes157 downloads3mo agoHugging Face10FinchResearch /pallas_splitted_18ctexttext-classification1M<n<10M0 likes149 downloads3y agoHugging Face11nikolina-p /gutenberg_clean_en_splits Dataset Card for Project Gutenberg Cleaned with splits (English Only) Dataset This dataset is a cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg, originally containing ~70,000 digitized books. The original dataset includes multiple languages, duplicate entries, and boilerplate content, all of which were removed for practicality and cleaner downstream use. This dataset containg 38.026 books. Dataset Splits The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_en_splits.texttext-generation10K<n<100K0 likes115 downloads1y agoHugging Face12Hugodonotexit /Superior-Reasoning-SFT-gpt-oss-120b-split-en Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering Summary This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields: input: the original input reasoning: the content extracted from <think> ... </think> within the original output (inner text only) output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.texttext-generation100K<n<1M2 likes113 downloads8mo agoHugging Face13tintin1027 /atomic-metrics-rm-splits Atomic Metrics RM Task Splits Preference-pair benchmark splits used by Atomic Metrics. The release contains four open-ended task families derived from public SHP, OASST1, and OASST2 preference data. Dataset structure Each configuration contains 10,000 training pairs and 2,000 test pairs. Every row has: { "sample_id": "source-specific stable ID", "source_dataset": "shp | oasst1 | oasst2", "category": "task configuration", "split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.texttext-generation10K<n<100K0 likes112 downloads21d agoHugging Face14ibm-research /Split-IFEval Split IFEval This dataset modifies the Instruction-Following Eval (IFEval) benchmark to split apart the task from the syntactic instructions in addition to fixing errors in the original dataset. It enables the use of research methods like attention steering that require access to the instruction text. To load the dataset, run: from datasets import load_dataset split_ifeval = load_dataset("ibm-research/Split-IFEval") Dataset Structure Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Split-IFEval.texttext-generationn<1K1 likes98 downloads1y agoHugging Face15Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads9d agoHugging Face16tuandunghcmut /nvidia_instruction_following_if_split_v3 Dataset Description This is the instruction_following split only (the chat split was intentionally excluded) from nvidia/Nemotron-SFT-Instruction-Following-Chat-v3, re-packaged as Parquet (sharded) instead of the original single JSONL file for faster loading and native support in the HF datasets viewer. No content was modified — this is a straight format conversion of the instruction_following subset. Source dataset: nvidia/Nemotron-SFT-Instruction-Following-Chat-v3 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3.texttext-generation100K<n<1M0 likes63 downloads3mo agoHugging Face17tuandunghcmut /nvidia_instruction_following_if_split_v3_non_thinking Dataset Description Non-thinking (no chain-of-thought) variant of tuandunghcmut/nvidia_instruction_following_if_split_v3, which is itself the instruction_following split of nvidia/Nemotron-SFT-Instruction-Following-Chat-v3. The reasoning_content field has been fully removed from every message (not just nulled) — each message now only has role and content. This is intended for training/evaluation setups that do not use chain-of-thought / reasoning traces. Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3_non_thinking.texttext-generation100K<n<1M0 likes55 downloads3mo agoHugging Face18WafaaFraih /rocov2-modality-splits ROCOv2 Modality-Specific Dataset Splits Dataset Description This dataset contains modality-specific splits of the ROCOv2 radiology dataset, organized and processed for training specialized medical image captioning models. Dataset Summary Total Samples: 1,000 Modalities: 5 Splits per Modality: train, validation, test Random Seed: 42 Processing Date: 2025-08-31 12:52:59.233482 Modality Distribution Modality Samples Percentage CT 188 18.8%… See the full description on the dataset page: https://huggingface.co/datasets/WafaaFraih/rocov2-modality-splits.image-to-text1K<n<10K0 likes54 downloads1y agoHugging Face19armand0e /qwen3.7-max-split-formatted Qwen Agent Thinking Online Distillation Rows This dataset contains cumulative assistant-turn training rows prepared for online logit distillation of Qwen-style agent models, plus a small set of no-tools chat rows to reduce tool-call overbias. Each row is a rendered-chat-ready conversation prefix ending at a target assistant turn. The trainer uses all prior messages as context and applies loss only to the final assistant span. Dataset Details Source trace repo:… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-split-formatted.texttext-generation1K<n<10K1 likes52 downloads4mo agoHugging Face20chorcat /rukh-puzzles-split chorcat/rukh-puzzles-split Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning. Part of Rukh, a chess language model built from scratch as a course on generative and agentic AI. Every derived dataset ships with the exact filters and counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.tabulartext-generation100K<n<1M0 likes50 downloads5d agoHugging Face21brandolorian /nemotron-post-training-samples-splits Nemotron Post-Training Samples with Train/Val/Test Splits This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios. Attribution This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0. Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.texttext-generation10K<n<100K0 likes49 downloads1y agoHugging Face22archit11 /verl-code-corpus-track-a-file-split archit11/verl-code-corpus-track-a-file-split Repository-specific code corpus extracted from the verl project and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_verl Total files: 214 Train files: 172 Validation files: 21 Test files: 21 File type filter: .py Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.texttext-generationn<1K0 likes47 downloads7mo agoHugging Face23nxvay /wikipedia-id-splits Dataset: Wikipedia Indonesian (Partial Splits) Motivasi Mengembangkan dan mempersiapkan dataset ini untuk fine-tuning bukanlah hal yang mudah, terutama dengan keterbatasan resource yang saya alami. Meski saya sudah berlangganan Colab Pro+ yang menjanjikan akses ke GPU berperforma tinggi (seperti A100 atau H100) dan resource lebih besar, ada beberapa tantangan signifikan yang muncul: Pembatasan Disk Space Colab VM: Saya sering menghadapi masalah "No space left on device"… See the full description on the dataset page: https://huggingface.co/datasets/nxvay/wikipedia-id-splits.texttext-generation100K<n<1M0 likes45 downloads1y agoHugging Face24disham993 /alpaca-train-validation-test-split Dataset Card for Alpaca I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows. Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.… See the full description on the dataset page: https://huggingface.co/datasets/disham993/alpaca-train-validation-test-split.texttext-generation10K<n<100K0 likes42 downloads3y agoHugging Face25AGmind /agmind-rag-splitter-ru-data RU Context-Aware Document Split Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми. Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML. Формат (Alpaca JSONL) { "instruction": "Раздели документ на смысловые части для системы поиска (RAG)...", "input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face26ecaplan /splitsThis dataset contains all four versions of the Splits! dataset, detailed in https://arxiv.org/pdf/2504.04640. See github repo https://github.com/eyloncaplan/splits. If our dataset is useful for you, please cite us: @misc{caplan2025splitsflexibledatasetevaluating, title={Splits! A Flexible Dataset for Evaluating a Model's Demographic Social Inference}, author={Eylon Caplan and Tania Chakraborty and Dan Goldwasser}, year={2025}, eprint={2504.04640}… See the full description on the dataset page: https://huggingface.co/datasets/ecaplan/splits.texttext-generation100M<n<1B0 likes37 downloads3mo agoHugging Face27Anna4242 /tool-n1-sft-unique-splits Tool-N1 SFT Unique with Train/Eval Splits This dataset contains supervised fine-tuning (SFT) data for training models on multi-hop tool usage and reasoning, with built-in train/evaluation splits. Usage from datasets import load_dataset # Load the dataset with splits dataset = load_dataset("Anna4242/tool-n1-sft-unique-splits") # Access splits train_data = dataset["train"] # 6,487 examples eval_data = dataset["eval"] # 1,622 examples # Example usage for example in… See the full description on the dataset page: https://huggingface.co/datasets/Anna4242/tool-n1-sft-unique-splits.texttext-generation1K<n<10K0 likes34 downloads1y agoHugging Face28anonymous-aardvark /submission14717_fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: omitted Paper: omitted Dataset Description This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.texttext-generation100K<n<1M0 likes29 downloads1y agoHugging Face29flavianv /deepshopper-mapper-reward-splits DeepShopper frozen splits (mapper / reward) Deterministic, leakage-safe train/test splits used across DeepShopper. Split assignment is a stable sha1(need) hash (same need never crosses train/test; reproducible). Gender-stratified. Contains: fashionrec_task1_mapper/{train,test}, amz_mapper/{female,male,other}.{train,test} (need→plan), and amz_reward_bundle/{female,male,other}.{train,test} (need→outfit, reward positives). Generated by scripts/make_mapper_reward_splits.py. Code:… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-mapper-reward-splits.text-generation0 likes29 downloads3mo agoHugging Face30Reubencf /Amharic_corpus_split Amharic Corpus — 4 x 5k Splits A randomly shuffled subset of Reubencf/Amharic_corpus, divided into four equal splits of 5,000 rows each (20,000 rows total). Splits: split_1, split_2, split_3, split_4 (5,000 rows each) Format: JSON Lines, one {"text": "..."} per line. Sampling: random without replacement (seed 42); the four splits are mutually exclusive. from datasets import load_dataset ds = load_dataset("Reubencf/Amharic_corpus_split") print(ds) # split_1..split_4, 5000 rows… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Amharic_corpus_split.texttext-generation10K<n<100K0 likes29 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.