datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zebra-cot-mistral-small-3.2-24b-preprocessed
Zebra-CoT Preprocessed — Mistral Hackathon 2026
Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct.
Format
text: formatted as [INST] question [/INST] <think> reasoning </think> answer
image: PIL JPEG image for the corresponding visual task
Usage
Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning.
Hackathon
Created for Mistral Hackaton 2026 — Fine-tuning track with W&B.
project_gutenberg_preprocessed
Gutenberg
Our version of the project gutenberg corpus, so as used to pretrain Apertus (v1 being used before 9T, v2 between 9T and 12T).
More details about data provenance, preparation, and statistics can be found in our tech report.
Sampling, filtering and data-preparation scripts can be found in our dedicated GitHub repository.
Feel free to reach out for any questions or suggestions 😊
lex_files_preprocessed
Dataset Card for "LexFiles"
Dataset Summary
Disclaimer: This is a pre-proccessed version of the LexFiles corpus (https://huggingface.co/datasets/lexlms/lexfiles), where documents are pre-split in chunks of 512 tokens.
The LeXFiles is a new diverse English multinational legal corpus that we created including 11 distinct sub-corpora that cover legislation and case law from 6 primarily English-speaking legal systems (EU, CoE, Canada, US, UK, India).
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/lexlms/lex_files_preprocessed.tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model:
Locutusque/TinyMistral-248M-Instruct
lfqa-preprocessed-itBangla_Masked_Language_Model_dataset_preprocessedLoC-PD-Books-preprocessed
LoC-PD-Books: preprocessed
This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps:
apply clean-text package keeping casing and newlines
drop OCR garbled text in first few lines of each example
fix (most) 'hard' newlines w/ regex similar to gutenberg clean
'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher
tulu_sft_mixture_preprocessed
Tulu SFT Mixture Preprocessed
This dataset was created by preprocessing the
allenai/tulu-3-sft-mixture
dataset for single-turn supervised fine-tuning.
The preprocessing keeps English user -> assistant examples from the selected
Tulu sources, applies length filtering with the official
Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates.
The resulting train split contains 151,292 examples with a maximum sequence
length of 7,168 tokens.
Each row contains the… See the full description on the dataset page: https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.texthumanizer-preprocessed-dataarxiv_summarization_20k_preprocessed
ArXiv Summarization Dataset - 20K Preprocessed
A preprocessed dataset of 20,000 ArXiv papers with their full articles and abstracts, designed for abstract generation and summarization tasks.
Dataset Description
This dataset contains 20,000 ArXiv papers that have been filtered and preprocessed to ensure quality for training summarization models. Each example contains the full article text and its corresponding abstract.
Dataset Structure
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/yilmazzey/arxiv_summarization_20k_preprocessed.timemachine-dataset-preprocessedPreprocessed_Solidity_Dataset_V1This dataset consists of 4,134 unique Solidity files. The files were gathered from three sources: Etherscan, Github and DISL dataset. Six preprocessing steps were applied:
Step 1 "Cleaning": Unnecessary parts such as comments or blank lines were removed from each file.
Step 2 "Formatting": Each file was converted with Prettier (and the corresponding Solidity-plugin) so that the final model only generates code in a correct format.
Step 3 "Slither Analysis": Each file has been checked for… See the full description on the dataset page: https://huggingface.co/datasets/fbnhnsl/Preprocessed_Solidity_Dataset_V1.
