datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually.
maze
This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS.
It has two columns:
'Maze': representation of maze in a list of string.shape is 30*30
visual example
'Path': solution from start point to end point in a list of string, each item represent a position in the maze.
0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
Mazesynthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj.
Original README
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.smoltalk2-thinkdigit_mask_false_positive_cv12_rawpolka-pretrain-en-pl-v1Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format
This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
into the ShareGPT format while preserving the original splits and columns.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"messages": [
{"role": "user", "content": "User message"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.maze-30x30-hard-1kMAZEL16Sitesdigit_mask_augmented_raw
Dataset Card for "digit_mask_augmented_raw"
More Information needed
details_MaziyarPanahi__calme-2.7-qwen2-7b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.digit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
MazePlanning-Testsift-audio
SIFT Audio Dataset
Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models.
Dataset Description
This dataset contains audio samples paired with LLM-generated responses following the
AZeroS multi-mode approach. Each audio sample is processed in three different modes
to train models that can both respond conversationally AND describe/analyze audio.
SIFT Modes
Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.acoustic_scattering_mazeThis Dataset is part of The Well Collection.
How To Load from HuggingFace Hub
Be sure to have the_well installed (pip install the_well)
Use the WellDataModule to retrieve data as follows:
from the_well.benchmark.data import WellDataModule
# The following line may take a couple of minutes to instantiate the datamodule
datamodule = WellDataModule(
"hf://datasets/polymathic-ai/",
"acoustic_scattering_maze",
)
train_dataloader = datamodule.train_dataloader()
for batch in… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/acoustic_scattering_maze.SYNTHETIC-1-800Kcloning-deterministic-worlds-maze9
Maze9 Dataset
This dataset is released with the CVPR 2026 paper Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models.
Contents
9x9_n0.8/: Maze9 training set
9x9_val/: Maze9 validation set
Each split contains:
info.hdf5: trajectory metadata, states, and actions
obs/*.npy: observation videos for each trajectory
Usage
Set the paths in the training code:
export TRAIN_DIR=/path/to/9x9_n0.8
export… See the full description on the dataset page: https://huggingface.co/datasets/xiazaishuo/cloning-deterministic-worlds-maze9.piper_dataset20251222抓取物体数据
0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
libritts-r-mimi-latentsLlama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
converted to ShareGPT format and merged into a single dataset.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"original_split": "code|math|science|chat|safety",
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.OpenMathReasoning_ShareGPTOriginal README:
OpenMathReasoning
OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs).
This dataset contains
540K unique mathematical problems sourced from AoPS forums,
3.2M long chain-of-thought (CoT) solutions
1.7M long tool-integrated reasoning (TIR) solutions
566K samples that select the most promising solution out of many candidates (GenSelect)
We used Qwen2.5-32B-Instruct to preprocess problems, and
DeepSeek-R1 and QwQ-32B… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenMathReasoning_ShareGPT.cranfield-synthetic-drone-detectionIf you use this, pleace cite Drone Detection using Deep Neural Networks Trained on Pure Synthetic Data (https://arxiv.org/abs/2411.09077)
hermes-function-calling-v1-allubr-maze-nav
ubr-maze-nav — vision + instrument waypoint planning for a small tracked robot
Synthetic navigation corpus for fine-tuning small vision-language models to
plan local waypoint paths for a 0.3 m tracked ground robot in corridor/maze
environments, plus the frozen evaluation suite used in our internal reports.
Each sample is one first-person RGB frame (640×480) from the robot's camera
in a procedurally generated MuJoCo scene, an instruction carrying the goal
(bearing/range) and a… See the full description on the dataset page: https://huggingface.co/datasets/ubr-physical-ai/ubr-maze-nav.mazerunner
MazeRunner
This repository contains the MazeRunner datasets as used in "Retrieval-augmented Decision Transformer: External Memory for In-context RL":
Datasets for grid-size 15x15.
The 15x15 folder contains 300 .npz files. Not all of them were used for our experiments.
Download the dataset using:
huggingface-cli download ml-jku/mazerunner --local-dir=./mazerunner --repo-type dataset
For dataloading we refer to our Github repository: https://github.com/ml-jku/RA-DT
Citation:… See the full description on the dataset page: https://huggingface.co/datasets/ml-jku/mazerunner.toy_maze_2d_v2kimodo-rigs
Kimodo Avatar Rigs
Two Unreal Engine 5 Mannequin skeletons and one rigid-parts robot, in glTF
binary form, hosted so the
Kimodo Avatar Motion Studio
can offer them as built-in rigs. The tool fetches them on click; nothing else
reads this repo.
They live here rather than beside the app because a Space repo rejects any
commit containing a binary file, and a dataset is the repo type built for them.
File
Size
Model
ue5-manny.glb
5.2 MB
UE5 male Mannequin, "Manny"… See the full description on the dataset page: https://huggingface.co/datasets/MazenVR/kimodo-rigs.riddles_evolved
Dataset Card for "riddles_evolved"
More Information needed
