datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
witadni-mini
ADNI mini v1.3 — SynthSeg-masked
This is a local derivative of medarc/adni-mini-v1-3.
It preserves the v1.3-r2 row order, metadata columns, labels, image geometry,
and float32 values inside the brain. The only image change is:
image[synthseg_dseg == 0] = 0.0
The brain mask is therefore defined strictly as nonzero labels in the matching
SynthSeg discrete segmentation.
See comparison.json and per_scan_stats.csv for measured storage and mask
statistics. This derivative is not the… See the full description on the dataset page: https://huggingface.co/datasets/medarc/adni-mini.open-imagesbitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2
Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2"
More Information needed
MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.MiniCPM5-1B-atlas
juiceb0xc0de/MiniCPM5-1B-atlas
A brain atlas for openbmb/MiniCPM5-1B, a 1B on-device model with a 130k bilingual vocabulary. This is not a chat dataset or a benchmark. It is an internal-mechanics map, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know which parts of this model are safe to edit, where its output-vocabulary directions live, or which layers are carrying the most… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/MiniCPM5-1B-atlas.msmarco-msmarco-MiniLM-L6-v3
MS MARCO with hard negatives from msmarco-MiniLM-L6-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-MiniLM-L6-v3.gelsight-mini-pretrain
GelSight Mini Pretrain
~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.
Frames
Sources
Real
536K
FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad
Sim
317K
sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated)
NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.the-stack-mini-nonshuffledThis repo contains (up to) 30k samples of 21 languages (top 20 languages by StackOverflow survey, html/css was splited).
'javascript', 'html', 'css', 'python', 'sql', 'typescript', 'shell', 'java', 'c-sharp', 'cpp', 'c', 'php', 'powershell', 'go', 'rust', 'kotlin', 'lua', 'dart', 'assembly', 'ruby', 'swift'
full-math-private-n256-Phi-4-mini-instruct-bonrole-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.asia-energy-world-bank-energy-and-mining-indicators
India - Energy and Mining
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
The world economy needs ever-increasing amounts of energy to sustain economic growth, raise living standards, and reduce poverty. But today's trends in energy use are not sustainable. As the world's population grows and economies become more industrialized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-energy-world-bank-energy-and-mining-indicators.aime_1983_2023_grok-3-mini-high_traces_32768full-aime_2026-n256-Phi-4-mini-instruct-bonREASONING_evalchemy_64_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_64_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.You are given a 0-indexed array nums of n integers and an integer target.\nYou are initially positioned at index 0. In one step, you can… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_64_sharded_gpt-4o-mini.miniVLA-Nav
MiniVLA-Nav v1
A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation
Demo
All-scenes montage
Nova Carter navigating to named objects across all four Isaac Sim environments.
Dataset Summary
MiniVLA-Nav v1 is a simulation dataset for the Language-Conditioned Object Approach (LCOA) task: given a short natural-language instruction, an NVIDIA Nova Carter differential-drive robot must navigate to the named object and stop within 1 m.… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/miniVLA-Nav.mini-sinhala-flanMiniCPM-RobotManip-LIBERO
MiniCPM-RobotManip LIBERO
This dataset contains the four LIBERO suites converted to LeRobot v3 format
for the MiniCPM-RobotManip LIBERO full-parameter fine-tuning example in
starVLA.
Dataset summary
Suite
Episodes
Frames
Videos
LIBERO-10
358
95,740
716
LIBERO-Goal
405
48,131
810
LIBERO-Object
450
66,294
900
LIBERO-Spatial
423
51,707
846
Total
1,636
261,872
3,272
Format: LeRobot v3
Frequency: 20 Hz
Cameras: agent view and wrist view
Video… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/MiniCPM-RobotManip-LIBERO.rebot_mini_towel_folding2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/rebot_mini_towel_folding2.preprocessed-full-MATH-500-n256-Phi-4-mini-instruct-bonpreprocessed-full-gsm8k-private-n256-Phi-4-mini-instruct-bonmini_kittifull-MATH-500-n256-Phi-4-mini-instruct-bonmixtraltoken_fineweb_edu_mini_combinedthe-stack-minicomparia-conversations
comparia-conversations : one of the largest prompt + text completion datasets in French
Origin of the data: what is compar:IA?
Compar:IA is a conversational AI comparison tool (a "chatbot arena") developed within the French Ministry of Culture with a dual mission:
To educate and raise awareness about the diversity of models, cultural and linguistic biases, and the environmental impact of conversational AIs.
To improve French-language conversational AIs by… See the full description on the dataset page: https://huggingface.co/datasets/ministere-culture/comparia-conversations.randomized_clean_miniwob_episodes__image0_5000_v2
Dataset Card for "randomized_clean_miniwob_episodes__image0_5000_v2"
More Information needed
rebot_mini_towel_folding1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/rebot_mini_towel_folding1.aime_1983_2023_grok-3-mini-high_traces
