CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes381 downloads7mo agoHugging Face02jtatman /tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model: Locutusque/TinyMistral-248M-Instruct textquestion-answering1M<n<10M3 likes59 downloads3y agoHugging Face03pszemraj /LoC-PD-Books-preprocessed LoC-PD-Books: preprocessed This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps: apply clean-text package keeping casing and newlines drop OCR garbled text in first few lines of each example fix (most) 'hard' newlines w/ regex similar to gutenberg clean 'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher tabulartext-generation10K<n<100K1 likes35 downloads9mo agoHugging Face04HwanChang0106 /tulu_sft_mixture_preprocessed Tulu SFT Mixture Preprocessed This dataset was created by preprocessing the allenai/tulu-3-sft-mixture dataset for single-turn supervised fine-tuning. The preprocessing keeps English user -> assistant examples from the selected Tulu sources, applies length filtering with the official Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates. The resulting train split contains 151,292 examples with a maximum sequence length of 7,168 tokens. Each row contains the… See the full description on the dataset page: https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.tabulartext-generation100K<n<1M0 likes30 downloads2mo agoHugging Face05DhimanBose /Bangla_Masked_Language_Model_dataset_preprocessedtext-generation1M<n<10M0 likes27 downloads3y agoHugging Face06yilmazzey /arxiv_summarization_20k_preprocessed ArXiv Summarization Dataset - 20K Preprocessed A preprocessed dataset of 20,000 ArXiv papers with their full articles and abstracts, designed for abstract generation and summarization tasks. Dataset Description This dataset contains 20,000 ArXiv papers that have been filtered and preprocessed to ensure quality for training summarization models. Each example contains the full article text and its corresponding abstract. Dataset Structure The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/yilmazzey/arxiv_summarization_20k_preprocessed.texttext-generation10K<n<100K0 likes15 downloads10mo agoHugging Face07SOULAMA /timemachine-dataset-preprocessedtexttext-generation1K<n<10K1 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.