datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pytorch-Code-10K
Hot Coco Training Dataset
A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!)
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.Monster-Piano
Monster Piano
580204 solo Piano MIDI scores representations from Monster MIDI dataset
Installation and use
Load dataset
#===================================================================
from datasets import load_dataset
#===================================================================
monster_piano = load_dataset('asigalov61/Monster-Piano')
dataset_split = 'train'
dataset_entry_index = 0
dataset_entry =… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Monster-Piano.colored-monsters
Colored Monsters
A toy dataset for unconditional image generation. It consists of 3 million renders of 3D monsters at a resolution of 256x256 pixels.
Method
Randomly select 3 out of 27 monsters.
Monsters:
alien
alpaking
armabee
birb
blue_demon
bunny
cactoro
demon
dino
dragon_evolved
dragon
fish
frog
ghost
ghost_skull
glub_evolved
glub
goleling_evolved
goleling
monkroose
mushnub
mushroom_king
orc_skull
pigeon
squidle
tribale
yeti
Randomly assign 1 of 9 colors to each… See the full description on the dataset page: https://huggingface.co/datasets/hayden-donnelly/colored-monsters.galahad-deconf-verb
Deconfounded verb set (RoboCasa)
Part of the Galahad release · Project page · Code + battery + generator
Lift-versus-slide with the same object under both verbs, role-balanced, so the verb is the only cue. Behind the verb-selectivity result (+44.3pp, told-slide→lift 2.0%).
LeRobot v3 format. Generated by the deconfounding generators in the release repo; the generator produces the exam (the battery) and the medicine (the training set) from the same code.
code_search_netcode-search-net/code_search_net , already loaded and converted to paraquet so you dont have to enable remote execution to use it .
OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/monster75/OpenThoughts3-1.2M.Monster-Chords
Monster Chords
High quality tones chords and pitches chords from select MIDIs in the Monster MIDI Dataset
Load dataset
#===================================================================
from datasets import load_dataset
#===================================================================
monster_chords = load_dataset('asigalov61/Monster-Chords')
Decode tones chords
#===================================================================
#… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Monster-Chords.toxic-monstergalahad-deconf-object
Deconfounded object-identity set (LIBERO)
Part of the Galahad release · Project page · Code + battery + generator
900 episodes, 90 per object across all 10 LIBERO objects, perfect role balance: every object appears as the target and as a distractor across randomized positions, so the instruction is the only predictor of the target. The training set behind the object-substitution cure (base 13.5% → 94.5%).
LeRobot v3 format. Generated by the deconfounding generators in the… See the full description on the dataset page: https://huggingface.co/datasets/phi-monster/galahad-deconf-object.monster-girl-encyclopedia-stories
All user-submitted Monster Girl Encyclopedia (魔物娘図鑑) fanfictions/short stories from kurobinega as of October 2025, or roughly 19000 documents. Japanese-only content for now. It can be safely assumed that all stories on the source website have been approved by KC as being in line with the basic principles behind the MGE universe.
The stories are roughly sorted by date, with consecutive story chapters naturally ordered one after the other. User comments and author commentaries have been… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/monster-girl-encyclopedia-stories.JFLD_punipuni_monster
Dataset Card for "JFLD_punipuni_monster"
See here for the details of this corpus.
For the whole of the project, see our project page.
More Information needed
Monster-Caps
Monster Caps
Detailed captions for all MIDIs from Monster MIDI Dataset
Project Los Angeles
Tegridy Code 2025
MonsterInstructA Curated superset of 12 of the best LLM instruct opensource datasets available today
Below is a list of the datasets and the examples picked from them
MonsterInstruct-gemma2-formattedmonster-tray-pickplace
Monster Tray Pick-and-Place (Unitree G1)
A teleoperated manipulation dataset collected on a Unitree G1 humanoid
(dual-arm, Dex1 grippers, dual camera) for a pick-and-place task: pick up a
monster energy drink can and place it on a serving tray. Two object
variants are covered — a green can and an orange can — each with its
own language-conditioned task instruction.
This dataset was collected to fine-tune
GR00T N1.7 for the same task; the
resulting checkpoint is published at… See the full description on the dataset page: https://huggingface.co/datasets/tysyuvraj/monster-tray-pickplace.MonsterInstruct-llama3.2-formattedmonster-sanctuary-conversationsygo_monstersDND-Monster-Diffusionlmfdb-monster-71
lmfdb-monster-71
L-functions and Modular Forms Database mapped to Monster Group 71 shards
Dataset Description
Part of the Monster OSM Quest project, which maps the Monster Group's 71 shards to OpenStreetMap data.
Data Fields
See individual JSON files for schema.
Source Data
OSM: OpenStreetMap planet data (ODbL license)
LMFDB: L-functions and Modular Forms Database (CC BY-SA 4.0)
Mathematical: Ramanujan biographical data and FRACTRAN encodings… See the full description on the dataset page: https://huggingface.co/datasets/h4/lmfdb-monster-71.no_robots_chatformatted_version1
Dataset Card for "no_robots_FalconChatFormated"
More Information needed
no_robots_chatformatted_version2
Dataset Card for "no_robots_MistralChatFormated"
More Information needed
imnet1k_Gila_monster_Heloderma_suspectumgalahad-deconf-goal
Deconfounded goal set (LIBERO)
Part of the Galahad release · Project page · Code + battery + generator
600 episodes, objects displaced 6–13 cm with the instruction held fixed, so completion requires reading the referent's current position rather than a remembered one. Cures position-memorization (base 31.2% → 83.3% on displaced referents).
LeRobot v3 format. Generated by the deconfounding generators in the release repo; the generator produces the exam (the battery) and the… See the full description on the dataset page: https://huggingface.co/datasets/phi-monster/galahad-deconf-goal.monster_toyMonsterInstruct-llama3.1-formattedfasthtml_monsterui_docsdnd_monster_samplemonsterapi__Llama-3_1-8B-Instruct-orca-ORPOlingolee1
