datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NOAH-mini
MOAH mini
The dataset prest here is a very samll sample of NOAH dataset.
In the original dataset each satellite image is ~650MB with 234,089 images present in 11 bands.
It is not feasible to upload the complete dataset.
A sample of the dataset across diffrent modalities can be seen in the figure below:
The diffrence between NOAH and NOAH mini is hilighted in the figure below.
Each subplot is a band of Landsat 8 in NOAH.
The region hilighted in red is the region available in NOAH… See the full description on the dataset page: https://huggingface.co/datasets/mutakabbirCarleton/NOAH-mini.Amelia42-Mini
Dataset Overview
The Amelia42-Mini dataset provides air traffic position reports for 42 major U.S. airports, including the following airports:
KATL (Hartsfield-Jackson Atlanta International Airport)
KBDL (Bradley International Airport)
KBOS (Boston Logan International Airport)
KBWI (Baltimore/Washington International Thurgood Marshall Airport)
KCLE (Cleveland Hopkins International Airport)
KCLT (Charlotte Douglas International Airport)
KDCA (Washington National Airport)
KDEN… See the full description on the dataset page: https://huggingface.co/datasets/AmeliaCMU/Amelia42-Mini.OxHyperMinerals_MINISeismicX-Cont-mini
SeismicX-Cont Mini Two-Hour Subset
This folder is the compact, Zenodo-archived two-hour mini release for
SeismicX-Cont. It is designed for quick download, tutorial use, software smoke
tests, and checking that the HDF5, annotation, SQLite, dataloader, picker, and
validation workflow all fit together before using the full 14-day data product.
Zenodo record: https://zenodo.org/records/21331024
DOI: https://doi.org/10.5281/zenodo.21331024
Hugging Face record:… See the full description on the dataset page: https://huggingface.co/datasets/cangyeone/SeismicX-Cont-mini.starcop_allbands_mini
MINI version of the STARCOP dataset
For full details please refer to https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1
MINI_21poc-mini-trade-game-dataset
Dataset Card for Mini Trade Game NPC Dataset
Dataset Summary
This dataset contains synthetic training examples for simulating NPC (Non-Player Character) merchant behavior in a trading game scenario. The dataset is designed to train language models to generate contextually appropriate trading responses based on item properties, relationship status, and player interactions.
All examples are in Traditional Chinese (zh-TW), with player inputs and NPC responses using… See the full description on the dataset page: https://huggingface.co/datasets/aotoki/poc-mini-trade-game-dataset.vn-par-ministry
Vietnam PAR Index ministry panel (2012-2025)
Public Administration Reform Index (PAR INDEX / Chỉ số CCHC) for Vietnam's ministries and ministerial-level agencies (bộ, cơ quan ngang bộ), 2012-2025. Core fields: MoHA-verified points, survey points, composite PAR Index (0-100), and rank when officially published. Universe size changes by year (19→12). 2024 has scores but no official ranking. 2025 uses the post-restructure 12-agency panel including Bộ Nông nghiệp và Môi trường… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-par-ministry.S2-TOMG-Bench-mini
S^2-Bench Dataset (TOMG) (mini version, 4.5k entries)
Official Huggingface Datasets for S^2-Bench: "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation"
Please refer to our Github Repo for more usage and useful information.
This mini version offers a choice for researchers with limited resources.
Configurations
Each configuration represents a different task:
MolCustom_AtomNum: Molecular customized generation by atom number… See the full description on the dataset page: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench-mini.fineweb-ultra-mini
Dataset Card for Fineweb Ultra Mini
Fineweb Ultra Mini is a dataset derived from the original Fineweb dataset made by huggingface (see here: https://huggingface.co/datasets/HuggingFaceFW/fineweb).
The dataset focuses on extracting high quality data from the Fineweb dataset, from the 2-3% range. If you would like even more high-quality data, keep out for our next release, fineweb ultra mini pro, which focuses on the 0-1% of high quality data originally found in fineweb.… See the full description on the dataset page: https://huggingface.co/datasets/motionlabs/fineweb-ultra-mini.job-descriptions-dataset-miniProtGlycanDock-Minimal
ProtGlycanDock Dataset – Minimal Version
NOTE: This is the minimal version of ProtGlycanDock, uploaded for the NeurIPS 2026 Datasets and Benchmarks Track due to the < 4 GB size requirement. It contains only the mmCIF structural files and the metadata tables needed to define splits and case information.The full dataset (including PDB structures, JSON inputs, and reconstruction scripts) is available in the accompanying complete repository… See the full description on the dataset page: https://huggingface.co/datasets/ProtGlycanDock/ProtGlycanDock-Minimal.malicious-prompts-minilm-embeddingsmini_MMPDEnglish-Mini
LLM-English-100MB — Compact & Dense English Teaching Corpus
A 100MB, extremely clean CSV designed to teach an LLM English from scratch via instruction-tuning. No noise, no HTML, no duplicates — just pure grammar, vocabulary, and syntax transformations.
Generated with a single paste-and-run Python script in Google Colab.
Why this teaches English
Instead of raw text, the dataset is instruction -> input -> output pairs that force the model to learn rules:
Grammar… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/English-Mini.MINI_SAVIpatents-classified-2106-gpt5-minilmsys-chat-lewd-minimalThis dataset is extracted from lmsys/lmsys-chat-1m.
Multiple filters were used to extract 800+ pieces of sex-related data.
Removed:
prompt words generated by role-playing programs.
Jailbreak prompts.
Answers that are too "appropriate"
Meditation-miniset-v0.2
Synthetic Meditation Dataset v0.2 🧘♂️🧘♀️✨
Welcome to the Synthetic Meditation Dataset v0.2, a comprehensive collection of meditation guidance prompts designed to assist users in different stages of emotional wellbeing and mindfulness. This dataset aims to help developers train AI models to provide empathetic and personalized meditation guidance. The dataset focuses on inclusivity, personalization, and diverse meditation contexts to cater to a wide audience.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/BuildaByte/Meditation-miniset-v0.2.qwen9b-coop-mini-swe-agent
qwen9b-coop-mini-swe-agent
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote.
The matched solo version is at
CooperBench/qwen9b-solo-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.English-Mini
LLM-English-100MB — Compact & Dense English Teaching Corpus
A 100MB, extremely clean CSV designed to teach an LLM English from scratch via instruction-tuning. No noise, no HTML, no duplicates — just pure grammar, vocabulary, and syntax transformations.
Generated with a single paste-and-run Python script in Google Colab.
Why this teaches English
Instead of raw text, the dataset is instruction -> input -> output pairs that force the model to learn rules:
Grammar… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/English-Mini.CAD-MiniBooNe
CAD-MiniBooNe
Dataset Summary
CAD-MiniBooNe is a single-source continual anomaly detection benchmark scenario derived from the MiniBooNE Particle Identification dataset. It recasts the dataset's original binary classification task (distinguishing electron-neutrino "signal" events from muon-neutrino "background" events) as an anomaly detection problem and converts the tabular data into a sequence of concept-grouped tasks.
The dataset contains 66,116 samples, 5… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/CAD-MiniBooNe.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.ai-tool-prompts-mini
AI Tool Prompts Mini
A tiny synthetic dataset for experimenting with tool routing and text classification.
Columns
id — row identifier
prompt — user request
tool — expected tool category
Labels
search
calculator
weather
translation
summarize
code
email
calendar
Intended use
This dataset is designed for:
Hugging Face demos
text-classification experiments
tool-routing prototypes
educational projects
All examples are synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/Beratung/ai-tool-prompts-mini.mini_etymologymining-silicosis
Mining Communities & Silicosis in Sub-Saharan Africa
Abstract
Synthetic dataset modelling silicosis, TB co-morbidity, and occupational lung disease among miners and former miners across three settings in SSA. SA gold mines have a silicosis epidemic; silica dust causes progressive fibrotic lung disease. Silicosis increases TB risk 3-4x. Mining drives HIV and TB epidemics in southern Africa.
Scenarios
Deep Gold Mining: SA-type underground gold mines with ~22%… See the full description on the dataset page: https://huggingface.co/datasets/Laksh51/mining-silicosis.Math-Mini
Clean Math Reasoning Dataset
A clean dataset for training and evaluating language models on mathematical problem solving.
The dataset contains concise mathematical question-and-answer pairs designed to improve model performance on structured numerical reasoning tasks.
Dataset Structure
The dataset contains two fields:
Column
Description
prompt
A mathematical problem or question
response
The corresponding solution
Example:
prompt:
48392+92831=?… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Math-Mini.Math-IIO-68K-Mini
Mathematics Dataset for AI Model Training
This dataset contains 68,000 rows of mathematical questions and their corresponding solutions. It is designed for training AI models capable of solving mathematical problems or providing step-by-step explanations for a variety of mathematical concepts. The dataset is structured into three columns: input, instruction, and output.
Dataset Overview
Input: A mathematical question or problem statement (e.g., arithmetic, algebra… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-IIO-68K-Mini.Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills!
This dataset contains:
a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale.
provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.Math-Mini
Clean Math Reasoning Dataset
A clean dataset for training and evaluating language models on mathematical problem solving.
The dataset contains concise mathematical question-and-answer pairs designed to improve model performance on structured numerical reasoning tasks.
Dataset Structure
The dataset contains two fields:
Column
Description
prompt
A mathematical problem or question
response
The corresponding solution
Example:
prompt:
48392+92831=?… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Math-Mini.
