CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes379 downloads7mo agoHugging Face02swiss-ai /project_gutenberg_preprocessed Gutenberg Our version of the project gutenberg corpus, so as used to pretrain Apertus (v1 being used before 9T, v2 between 9T and 12T). More details about data provenance, preparation, and statistics can be found in our tech report. Sampling, filtering and data-preparation scripts can be found in our dedicated GitHub repository. Feel free to reach out for any questions or suggestions 😊 texttext-generation100K<n<1M8 likes288 downloads8mo agoHugging Face03lexlms /lex_files_preprocessed Dataset Card for "LexFiles" Dataset Summary Disclaimer: This is a pre-proccessed version of the LexFiles corpus (https://huggingface.co/datasets/lexlms/lexfiles), where documents are pre-split in chunks of 512 tokens. The LeXFiles is a new diverse English multinational legal corpus that we created including 11 distinct sub-corpora that cover legislation and case law from 6 primarily English-speaking legal systems (EU, CoE, Canada, US, UK, India). The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/lexlms/lex_files_preprocessed.text-generation1M<n<10M4 likes96 downloads3y agoHugging Face04jtatman /tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model: Locutusque/TinyMistral-248M-Instruct textquestion-answering1M<n<10M3 likes58 downloads3y agoHugging Face05efederici /lfqa-preprocessed-ittextquestion-answering10K<n<100K2 likes45 downloads3y agoHugging Face06DhimanBose /Bangla_Masked_Language_Model_dataset_preprocessedtext-generation1M<n<10M0 likes31 downloads3y agoHugging Face07pszemraj /LoC-PD-Books-preprocessed LoC-PD-Books: preprocessed This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps: apply clean-text package keeping casing and newlines drop OCR garbled text in first few lines of each example fix (most) 'hard' newlines w/ regex similar to gutenberg clean 'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher tabulartext-generation10K<n<100K1 likes31 downloads9mo agoHugging Face08HwanChang0106 /tulu_sft_mixture_preprocessed Tulu SFT Mixture Preprocessed This dataset was created by preprocessing the allenai/tulu-3-sft-mixture dataset for single-turn supervised fine-tuning. The preprocessing keeps English user -> assistant examples from the selected Tulu sources, applies length filtering with the official Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates. The resulting train split contains 151,292 examples with a maximum sequence length of 7,168 tokens. Each row contains the… See the full description on the dataset page: https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.tabulartext-generation100K<n<1M0 likes26 downloads2mo agoHugging Face09vishal-adithya /texthumanizer-preprocessed-datatexttext-generation10K<n<100K1 likes23 downloads1y agoHugging Face10yilmazzey /arxiv_summarization_20k_preprocessed ArXiv Summarization Dataset - 20K Preprocessed A preprocessed dataset of 20,000 ArXiv papers with their full articles and abstracts, designed for abstract generation and summarization tasks. Dataset Description This dataset contains 20,000 ArXiv papers that have been filtered and preprocessed to ensure quality for training summarization models. Each example contains the full article text and its corresponding abstract. Dataset Structure The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/yilmazzey/arxiv_summarization_20k_preprocessed.texttext-generation10K<n<100K0 likes16 downloads10mo agoHugging Face11SOULAMA /timemachine-dataset-preprocessedtexttext-generation1K<n<10K1 likes13 downloads8mo agoHugging Face12fbnhnsl /Preprocessed_Solidity_Dataset_V1This dataset consists of 4,134 unique Solidity files. The files were gathered from three sources: Etherscan, Github and DISL dataset. Six preprocessing steps were applied: Step 1 "Cleaning": Unnecessary parts such as comments or blank lines were removed from each file. Step 2 "Formatting": Each file was converted with Prettier (and the corresponding Solidity-plugin) so that the final model only generates code in a correct format. Step 3 "Slither Analysis": Each file has been checked for… See the full description on the dataset page: https://huggingface.co/datasets/fbnhnsl/Preprocessed_Solidity_Dataset_V1.texttext-generation1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.