CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.8k downloads2mo agoHugging Face03AmelieSchreiber /toricgt-curated-splits ToricGT Curated Graph Reasoning Splits Curated working dataset repository for ToricGT. The upload contains only curated split Parquet files and metadata generated locally. Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit. Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources. Files train.parquet validation.parquet test.parquet all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.tabulartext-generation1M<n<10M0 likes1.7k downloads4mo agoHugging Face04Xuhui /sft_processed_large_split sft_processed_large — profile-disjoint split This is the train / val / test split of Xuhui/sft_processed_large, the OdysSim midtraining corpus (21.4M interactions across 63 datasets). Split structure split rows how it's built train 21.20M what's left after val + test are carved out val 28K per-dataset random sample, in-distribution; for checkpoint selection test 128K profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.texttext-generation10M<n<100M1 likes599 downloads5mo agoHugging Face05vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split UltraData-Math L3 Textbook Exercise Synthetic Split Source dataset: openbmb/UltraData-Math Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic Each row contains: uid question answer The original content field was split using the literal markers The exercise: and The solution:. texttext-generation10M<n<100M1 likes496 downloads6mo agoHugging Face06jwkirchenbauer /fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.tabulartext-generation100K<n<1M0 likes369 downloads7mo agoHugging Face07nikolina-p /gutenberg_clean_en_splits Dataset Card for Project Gutenberg Cleaned with splits (English Only) Dataset This dataset is a cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg, originally containing ~70,000 digitized books. The original dataset includes multiple languages, duplicate entries, and boilerplate content, all of which were removed for practicality and cleaner downstream use. This dataset containg 38.026 books. Dataset Splits The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_en_splits.texttext-generation10K<n<100K0 likes115 downloads1y agoHugging Face08Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads9d agoHugging Face09tuandunghcmut /nvidia_instruction_following_if_split_v3 Dataset Description This is the instruction_following split only (the chat split was intentionally excluded) from nvidia/Nemotron-SFT-Instruction-Following-Chat-v3, re-packaged as Parquet (sharded) instead of the original single JSONL file for faster loading and native support in the HF datasets viewer. No content was modified — this is a straight format conversion of the instruction_following subset. Source dataset: nvidia/Nemotron-SFT-Instruction-Following-Chat-v3 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3.texttext-generation100K<n<1M0 likes63 downloads3mo agoHugging Face10tuandunghcmut /nvidia_instruction_following_if_split_v3_non_thinking Dataset Description Non-thinking (no chain-of-thought) variant of tuandunghcmut/nvidia_instruction_following_if_split_v3, which is itself the instruction_following split of nvidia/Nemotron-SFT-Instruction-Following-Chat-v3. The reasoning_content field has been fully removed from every message (not just nulled) — each message now only has role and content. This is intended for training/evaluation setups that do not use chain-of-thought / reasoning traces. Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3_non_thinking.texttext-generation100K<n<1M0 likes55 downloads3mo agoHugging Face11chorcat /rukh-puzzles-split chorcat/rukh-puzzles-split Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning. Part of Rukh, a chess language model built from scratch as a course on generative and agentic AI. Every derived dataset ships with the exact filters and counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.tabulartext-generation100K<n<1M0 likes50 downloads4d agoHugging Face12brandolorian /nemotron-post-training-samples-splits Nemotron Post-Training Samples with Train/Val/Test Splits This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios. Attribution This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0. Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.texttext-generation10K<n<100K0 likes49 downloads1y agoHugging Face13archit11 /verl-code-corpus-track-a-file-split archit11/verl-code-corpus-track-a-file-split Repository-specific code corpus extracted from the verl project and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_verl Total files: 214 Train files: 172 Validation files: 21 Test files: 21 File type filter: .py Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.texttext-generationn<1K0 likes47 downloads7mo agoHugging Face14nxvay /wikipedia-id-splits Dataset: Wikipedia Indonesian (Partial Splits) Motivasi Mengembangkan dan mempersiapkan dataset ini untuk fine-tuning bukanlah hal yang mudah, terutama dengan keterbatasan resource yang saya alami. Meski saya sudah berlangganan Colab Pro+ yang menjanjikan akses ke GPU berperforma tinggi (seperti A100 atau H100) dan resource lebih besar, ada beberapa tantangan signifikan yang muncul: Pembatasan Disk Space Colab VM: Saya sering menghadapi masalah "No space left on device"… See the full description on the dataset page: https://huggingface.co/datasets/nxvay/wikipedia-id-splits.texttext-generation100K<n<1M0 likes45 downloads1y agoHugging Face15disham993 /alpaca-train-validation-test-split Dataset Card for Alpaca I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows. Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.… See the full description on the dataset page: https://huggingface.co/datasets/disham993/alpaca-train-validation-test-split.texttext-generation10K<n<100K0 likes42 downloads3y agoHugging Face16Anna4242 /tool-n1-sft-unique-splits Tool-N1 SFT Unique with Train/Eval Splits This dataset contains supervised fine-tuning (SFT) data for training models on multi-hop tool usage and reasoning, with built-in train/evaluation splits. Usage from datasets import load_dataset # Load the dataset with splits dataset = load_dataset("Anna4242/tool-n1-sft-unique-splits") # Access splits train_data = dataset["train"] # 6,487 examples eval_data = dataset["eval"] # 1,622 examples # Example usage for example in… See the full description on the dataset page: https://huggingface.co/datasets/Anna4242/tool-n1-sft-unique-splits.texttext-generation1K<n<10K0 likes34 downloads1y agoHugging Face17anonymous-aardvark /submission14717_fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: omitted Paper: omitted Dataset Description This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.texttext-generation100K<n<1M0 likes29 downloads1y agoHugging Face18robworks-software /ccisd-teks-alignment-split [!WARNING] Deprecated - use ccisd-teks-alignment instead. This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/ccisd-teks-alignment. CCISD TEKS Alignment (pre-split) The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.texttext-classificationn<1K0 likes24 downloads2mo agoHugging Face19nikolina-p /mini_gutenberg_splits Dataset Card for Mini Project Gutenberg Dataset This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation. It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset. The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.texttext-generationn<1K0 likes18 downloads1y agoHugging Face20nelsonmaligro /alpaca-train-validation-test-split-50 Dataset Card for Alpaca I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows. Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction… See the full description on the dataset page: https://huggingface.co/datasets/nelsonmaligro/alpaca-train-validation-test-split-50.texttext-generation1K<n<10K0 likes15 downloads1mo agoHugging Face21bahaeddine09 /dz-lahja-dataset-50k-splittexttext-generation10K<n<100K0 likes11 downloads6mo agoHugging Face22br-llm-data /high_educability_training_splitgated high_educability_training_split Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis. Carregamento from datasets import load_dataset ds = load_dataset( "br-llm-data/high_educability_training_split", split="train", streaming=True, ) registro = next(iter(ds)) Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.tabulartext-generation1M<n<10M0 likes10 downloads18d agoHugging Face23nikolina-p /gutenberg_clean_tokenized_en_splits Overview This dataset is a tokenized version of the cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg. It contains full-text books in English, free of boilerplate content and duplicates, and includes a pre-tokenized version of each book's content using the GPT-2 tokenizer (tiktoken.get_encoding("gpt2")). This dataset is identical to nikolina-p/gutenberg_clean_tokenized_en except for the split configuration. Cleaning and Preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_tokenized_en_splits.texttext-generation10K<n<100K0 likes7 downloads1y agoHugging Face24guaran-ia /agustin-guarani-llm-splitsgated Guarani LLM Splits This repository contains Parquet splits used for Guarani LLM adaptation experiments. Files train.parquet: main training split synthetic.parquet: synthetic training data val_id.parquet: in-domain validation split val_ood.parquet: out-of-domain validation split test_id.parquet: in-domain test split test_ood.parquet: out-of-domain test split Loading from datasets import load_dataset repo_id = "agustin-lucas/guarani-llm-splits"… See the full description on the dataset page: https://huggingface.co/datasets/guaran-ia/agustin-guarani-llm-splits.tabulartext-generation100K<n<1M0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.