datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually.
maze
This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS.
It has two columns:
'Maze': representation of maze in a list of string.shape is 30*30
visual example
'Path': solution from start point to end point in a list of string, each item represent a position in the maze.
Mazesynthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj.
Original README
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.smoltalk2-thinkpolka-pretrain-en-pl-v1Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format
This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
into the ShareGPT format while preserving the original splits and columns.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"messages": [
{"role": "user", "content": "User message"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.maze-30x30-hard-1kMAZEL16Sitesdetails_MaziyarPanahi__calme-2.7-qwen2-7b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.sift-audio
SIFT Audio Dataset
Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models.
Dataset Description
This dataset contains audio samples paired with LLM-generated responses following the
AZeroS multi-mode approach. Each audio sample is processed in three different modes
to train models that can both respond conversationally AND describe/analyze audio.
SIFT Modes
Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.SYNTHETIC-1-800Klibritts-r-mimi-latentsLlama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
converted to ShareGPT format and merged into a single dataset.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"original_split": "code|math|science|chat|safety",
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.OpenMathReasoning_ShareGPTOriginal README:
OpenMathReasoning
OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs).
This dataset contains
540K unique mathematical problems sourced from AoPS forums,
3.2M long chain-of-thought (CoT) solutions
1.7M long tool-integrated reasoning (TIR) solutions
566K samples that select the most promising solution out of many candidates (GenSelect)
We used Qwen2.5-32B-Instruct to preprocess problems, and
DeepSeek-R1 and QwQ-32B… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenMathReasoning_ShareGPT.hermes-function-calling-v1-allriddles_evolved
Dataset Card for "riddles_evolved"
More Information needed
Maze-ReasoningMaze-Reasoning-v0.1arxiv.org/abs/2502.14669
OpenCodeReasoning_ShareGPT
Added cnversations column in ShareGPT format
Original README from nvidia/OpenCodeReasoning
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report - Discover the methodology and… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenCodeReasoning_ShareGPT.MAZEL8Siteslibritts-mimi
Dataset with Mimi Codes
This dataset adds Mimi codec codes to parler-tts/libritts_r_filtered.
Dataset Description
Each sample contains:
audio: Audio resampled to 24kHz (Mimi's native rate)
codes: 8-layer Mimi codec codes (list of 8 lists of integers)
text: Text transcription (from text_normalized column)
Additional columns preserved from source dataset
Stats
Source: parler-tts/libritts_r_filtered
Splits: train.clean.360
Samples: 112,326
Audio Sample Rate:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/libritts-mimi.Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT
This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
converted to ShareGPT format and merged into a single dataset.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"original_split": "code|math|science|chat|safety",
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT.AM-DeepSeek-R1-0528-Distilled-with-Systemdetails_MaziyarPanahi__calme-2.3-llama3-70b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3-70b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-llama3-70b.Maze-Reasoning-Reset-v0.1arxiv.org/abs/2502.14669
maze-solving-for-gemma-4test
mazhdrak/test — Mixed Instruction Dataset
A personal mixed-domain instruction dataset compiled from JSON files, CSV tables, Word documents, hardware reports, chatbot histories, and production manuals.
Languages: English + Bulgarian. Built for fine-tuning, RAG, and LLM evaluation.
Dataset Stats
Subset
File
Records
Description
Master (all)
train.jsonl
811
Complete unified dataset
Chat Exports
chat_exports.jsonl
235
Tabular Data
tabular_data.jsonl
233… See the full description on the dataset page: https://huggingface.co/datasets/mazhdrak/test.mazinger-dubber-profiles
Mazinger Dubber — Voice Profiles
Voice profiles for mazinger-dubber. Hosted on HuggingFace:
https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles
Adding a New Profile
1. Prepare your files
Create a folder named after the profile:
profiles/
└── my-name/
├── script.txt # Plain-text transcript matching the audio exactly
└── voice.m4a # Voice sample (supported: .m4a, .wav, .mp3)
Tips: 10–30 seconds of clear speech, minimal… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles.Maze-Reasoninglm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private
Dataset Card for Evaluation run of MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26
Dataset automatically created during the evaluation run of model MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private.
