datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.HR-Bench
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
🌐Homepage | 📖 Paper
📊 HR-Bench
We find that the highest resolution in existing multimodal benchmarks is only 2K. To address the current lack of high-resolution multimodal benchmarks, we construct HR-Bench. HR-Bench consists two sub-tasks: Fine-grained Single-instance Perception (FSP) and Fine-grained Cross-instance Perception (FCP).… See the full description on the dataset page: https://huggingface.co/datasets/DreamMr/HR-Bench.finqadataset_info:
features:
name: id
dtype: string
name: post_text
sequence: string
name: pre_text
sequence: string
name: question
dtype: string
name: answers
dtype: string
name: table
sequence:
sequence: string
splits:
name: train
num_bytes: 26984130
num_examples: 6251
name: validation
num_bytes: 3757103
num_examples: 883
name: test
num_bytes: 4838430
num_examples: 1147
download_size: 21240722
dataset_size: 35579663
dreamzero-egoverse-360-pretrainuscode
United States Code, versioned by release point
Every section of the United States Code, as published by the Office of the Law
Revision Counsel (OLRC) at uscode.house.gov, across
every release point from 113-21 (July 18, 2013) through the present. A release
point is OLRC's republication of the Code after a batch of Public Laws is
classified; this dataset covers 381 of them over 58 titles.
Each row carries the section's plain text, its verbatim USLM XML, its
citation, its place in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/uscode.dresscode_agnostic_and_densepose
DressCode Agnostic & DensePose Dataset
Agnostic images, corresponding masks, and DensePose images for the DressCode dataset.
Information about the usage can be found at:
https://github.com/jiwoohong93/ita-mdt_code
License
The Dress Code Dataset is proprietary to and © Yoox Net-a-Porter Group S.p.A. and its licensors.It is distributed by the University of Modena and Reggio Emilia and is available for non-commercial academic use under the licence terms provided… See the full description on the dataset page: https://huggingface.co/datasets/jiwoohong93/dresscode_agnostic_and_densepose.omni-dreams-samples
AlpaDreams Samples
Curated single-view driving sequences for evaluating the
nvidia/alpadreams-dit world model.
Layout
data/
└── single_view/
├── <clip-id>/
| ├── <clip-id_...>.mp4 # ground truth video
│ ├── <clip-id_..._hdmap>.mp4 # HD-map rasterized conditioning video
│ ├── first_frame.png # RGB first frame, extracted from ground truth video
│ └── prompt.txt # text prompt
└──… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/omni-dreams-samples.dreambooth
Dataset Card for "dreambooth"
Dataset of the Google paper DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
The dataset includes 30 subjects of 15 different classes. 9 out of these subjects are live subjects (dogs and cats) and 21 are objects. The dataset contains a variable number of images per subject (4-6). Images of the subjects are usually captured in different conditions, environments and under different angles.
We include a file… See the full description on the dataset page: https://huggingface.co/datasets/google/dreambooth.DreamOmni2Bench
DreamOmni2: Multimodal Instruction-based Editing and Generation Benchmark
This repository contains the DreamOmni2Bench benchmark dataset, introduced in the paper DreamOmni2: Multimodal Instruction-based Editing and Generation.
The DreamOmni2 project proposes two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical… See the full description on the dataset page: https://huggingface.co/datasets/xiabs/DreamOmni2Bench.Dreamer-V1-DataAfter heavier cleaning, the remaining data size is 3.12M.
WebDreamer: Model-Based Planning for Web Agents
WebDreamer is a planning framework that enables efficient and effective planning for real-world web agent tasks. Check our paper for more details.
This work is a collaboration between OSUNLP and Orby AI.
Repository: https://github.com/OSU-NLP-Group/WebDreamer
Paper: https://arxiv.org/abs/2411.06559
Point of Contact: Kai Zhang
Models
Dreamer-7B:
General… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Dreamer-V1-Data.trajectory_data_dream_32
d3LLM Trajectory Dataset
Project Page | Paper | GitHub | Blog
This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation".
Introduction
d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.DREAMS-AVATAR
DREAMS-AVATAR
The DREAMS-Avatar dataset from the DEGAS paper
(3DV 2025), re-registered to pure SMPL-X.
These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS
(Fig. 1b); what is new here is the registration.
32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted
to all views at once by our multiview tracker: 300 shape coefficients, 100 expression
coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.dream-coder
Program Synthesis Data
Generated program synthesis datasets used to train dreamcoder.
Currently just supports text & list data.
D-Rep
Summary
This is the dataset proposed in our paper Image Copy Detection for Diffusion Models (NeurIPS 2024).
D-Rep consists of 40, 000 image-replica pairs, in which each replica is generated by a diffusion model. The 40, 000 image-replica pairs are manually labeled with 6 replication levels ranging from 0 (no replication) to 5 (total replication). We divide D-Rep into a training set with 90% (36, 000) pairs and a test set with the remaining 10% (4, 000) pairs.… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/D-Rep.DreamOmni3BenchDREAM-1K
DREAM-1K
DREAM-1K (Description
with Rich Events, Actions, and Motions) is a challenging video description benchmark. It contains a collection of 1,000 short (around 10 seconds) video clips with diverse complexities from five different origins: live-action movies, animated movies, stock videos, long YouTube videos, and TikTok-style short videos. We provide a fine-grained manual annotation for each video.
Bellow is the dataset statistics:
us-statutes-at-large
United States Statutes at Large
Bound volumes of the United States Statutes at Large as published by the U.S. Government Publishing
Office on GovInfo, collection STATUTE. Each
volume is one PDF and one USLM XML file.
Files
metadata.jsonl one row per volume
pdfs/STATUTE-{n}.pdf volume PDF as served by GovInfo
xmls/STATUTE-{n}.xml volume USLM XML as served by GovInfo
granules/ per-law PDFs used by the conversion benchmark (a sample… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/us-statutes-at-large.dreamlip-gpt4v-500kdreamt
Dataset Description
DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices.
Version: 2.1.0
Repository: PhysioNet: DREAMT v2.1.0
Access Policy & Licensing
Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.Dress-ED
Dress-ED
Instruction-Guided Editing for Virtual Try-On and Try-Off
Overview
Dress-ED is a large-scale garment-editing dataset for instruction-guided virtual try-on and clothing manipulation research. Given a person image and a garment image, each example asks a model to apply a specific textual edit to the garment while keeping it worn on the person.
The dataset provides three complementary instructions per example — one relative to the garment… See the full description on the dataset page: https://huggingface.co/datasets/davidelobba/Dress-ED.DirectContacts2
DirectContacts2: A network of direct physical protein interactions derived from high throughput mass spectrometry experiments
Proteins carry out cellular functions by self-assembling into functional complexes, a process that depends on direct physical interactions
between components. While tools like AlphaFold and RoseTTAFold have advanced structure prediction, they remain limited in scaling to the full
human proteome. DirectContacts2 addresses this challenge by integrating… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/DirectContacts2.PRISM-100K
PRISM: Multi-View Multi-Capability Video SFT Dataset for Retail Embodied AI
Dataset Details
Dataset Description
PRISM is a video Supervised Fine-Tuning (SFT) dataset designed for training Vision-Language Models (VLMs) on retail-domain physical AI tasks. It features synchronized egocentric and exocentric video from real retail environments, annotated across 21 task types spanning embodied reasoning, common-sense reasoning, spatial perception, and… See the full description on the dataset page: https://huggingface.co/datasets/DreamVu/PRISM-100K.dreaddit
Dreaddit: A Reddit Dataset for Stress Analysis in Social Media
Consists of 3.5k labeled texts from five different categories of Reddit communities.
Citation
@inproceedings{turcan-mckeown-2019-dreaddit,
title = "{D}readdit: A {R}eddit Dataset for Stress Analysis in Social Media",
author = "Turcan, Elsbeth and
McKeown, Kathy",
editor = "Holderness, Eben and
Jimeno Yepes, Antonio and
Lavelli, Alberto and
Minard, Anne-Lyse and… See the full description on the dataset page: https://huggingface.co/datasets/andreagasparini/dreaddit.sangyo_no_yume_industrial_dreams
From the Frontier Research Team at Takara.ai we present the "Sangyo no Yume Industrial Dreams" dataset, a collection of AI-generated industrial dreamscapes.
Sangyo no Yume Industrial Dreams
Dataset Details
"Sangyo no Yume Industrial Dreams" is a collection of images generated using SDXL Lightning with specialized prompt engineering techniques. These images balance industrial themes with dreamlike qualities, creating a unique aesthetic that sits at the intersection of… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/sangyo_no_yume_industrial_dreams.Arabic-Dialects
Arabic Dialects Dataset (Bivalency & Code-Switching)
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties:
EGY – Egyptian Arabic
GLF – Gulf Arabic
LAV – Levantine Arabic
NOR – North African / Tunisian Arabic
MSA – Modern Standard Arabic
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.EASC
EASC: The Essex Arabic Summaries Corpus
Mo El-Haj, Udo Kruschwitz, Chris FoxUniversity of Essex, UK
This repository hosts EASC — the Essex Arabic Summaries Corpus — a collection of 153 Arabic source documents and 765 human-generated extractive summaries, created using Amazon Mechanical Turk.
EASC is one of the earliest publicly available datasets for Arabic single-document summarisation and remains widely used in research on Arabic NLP, extractive summarisation, sentence ranking… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/EASC.dreamdojo-ego-view
DreamDojo ego-view — video + instruction for Cosmos-Predict2.5 post-training
3,168 ego-view robot manipulation clips in the flat videos/ + metas/ layout that
cosmos-predict2.5's VideoDataset reads directly, with no conversion step.
Only the observation.images.ego_view camera is included. Episodes shorter than 94
frames are excluded: VideoDataset samples a random 93-frame window, and on a
93-frame video its np.random.randint(0, 0) raises rather than returning 0.… See the full description on the dataset page: https://huggingface.co/datasets/Eurong2/dreamdojo-ego-view.data_libero_dreamgen_stage1_s10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 200,
"total_frames": 33587,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/data_libero_dreamgen_stage1_s10.DreamNLPDream_Train
