CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01drengskapur /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.text1K<n<10K19 likes9.9k downloads2y agoHugging Face02DreamMr /HR-Bench Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models 🌐Homepage | 📖 Paper 📊 HR-Bench We find that the highest resolution in existing multimodal benchmarks is only 2K. To address the current lack of high-resolution multimodal benchmarks, we construct HR-Bench. HR-Bench consists two sub-tasks: Fine-grained Single-instance Perception (FSP) and Fine-grained Cross-instance Perception (FCP).… See the full description on the dataset page: https://huggingface.co/datasets/DreamMr/HR-Bench.textvisual-question-answering1K<n<10K17 likes7.8k downloads10mo agoHugging Face03dreamerdeo /finqadataset_info: features: name: id dtype: string name: post_text sequence: string name: pre_text sequence: string name: question dtype: string name: answers dtype: string name: table sequence: sequence: string splits: name: train num_bytes: 26984130 num_examples: 6251 name: validation num_bytes: 3757103 num_examples: 883 name: test num_bytes: 4838430 num_examples: 1147 download_size: 21240722 dataset_size: 35579663 text1K<n<10K28 likes4.8k downloads4y agoHugging Face04liang12121 /dreamzero-egoverse-360-pretraintabular10M<n<100M1 likes3.2k downloads5mo agoHugging Face05dreamproit /uscode United States Code, versioned by release point Every section of the United States Code, as published by the Office of the Law Revision Counsel (OLRC) at uscode.house.gov, across every release point from 113-21 (July 18, 2013) through the present. A release point is OLRC's republication of the Code after a batch of Public Laws is classified; this dataset covers 381 of them over 58 titles. Each row carries the section's plain text, its verbatim USLM XML, its citation, its place in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/uscode.tabulartext-retrieval100K<n<1M0 likes2k downloads1mo agoHugging Face06jiwoohong93 /dresscode_agnostic_and_densepose DressCode Agnostic & DensePose Dataset Agnostic images, corresponding masks, and DensePose images for the DressCode dataset. Information about the usage can be found at: https://github.com/jiwoohong93/ita-mdt_code License The Dress Code Dataset is proprietary to and © Yoox Net-a-Porter Group S.p.A. and its licensors.It is distributed by the University of Modena and Reggio Emilia and is available for non-commercial academic use under the licence terms provided… See the full description on the dataset page: https://huggingface.co/datasets/jiwoohong93/dresscode_agnostic_and_densepose.image100K<n<1M0 likes1.5k downloads1y agoHugging Face07nvidia /omni-dreams-samplesgated AlpaDreams Samples Curated single-view driving sequences for evaluating the nvidia/alpadreams-dit world model. Layout data/ └── single_view/ ├── <clip-id>/ | ├── <clip-id_...>.mp4 # ground truth video │ ├── <clip-id_..._hdmap>.mp4 # HD-map rasterized conditioning video │ ├── first_frame.png # RGB first frame, extracted from ground truth video │ └── prompt.txt # text prompt └──… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/omni-dreams-samples.imageimage-to-videon<1K4 likes1.3k downloads4mo agoHugging Face08google /dreambooth Dataset Card for "dreambooth" Dataset of the Google paper DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation The dataset includes 30 subjects of 15 different classes. 9 out of these subjects are live subjects (dogs and cats) and 21 are objects. The dataset contains a variable number of images per subject (4-6). Images of the subjects are usually captured in different conditions, environments and under different angles. We include a file… See the full description on the dataset page: https://huggingface.co/datasets/google/dreambooth.imagen<1K56 likes1.2k downloads3y agoHugging Face09xiabs /DreamOmni2Bench DreamOmni2: Multimodal Instruction-based Editing and Generation Benchmark This repository contains the DreamOmni2Bench benchmark dataset, introduced in the paper DreamOmni2: Multimodal Instruction-based Editing and Generation. The DreamOmni2 project proposes two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical… See the full description on the dataset page: https://huggingface.co/datasets/xiabs/DreamOmni2Bench.imageimage-to-imagen<1K0 likes1.1k downloads11mo agoHugging Face10osunlp /Dreamer-V1-DataAfter heavier cleaning, the remaining data size is 3.12M. WebDreamer: Model-Based Planning for Web Agents WebDreamer is a planning framework that enables efficient and effective planning for real-world web agent tasks. Check our paper for more details. This work is a collaboration between OSUNLP and Orby AI. Repository: https://github.com/OSU-NLP-Group/WebDreamer Paper: https://arxiv.org/abs/2411.06559 Point of Contact: Kai Zhang Models Dreamer-7B: General… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Dreamer-V1-Data.text1M<n<10M5 likes975 downloads1y agoHugging Face11d3LLM /trajectory_data_dream_32 d3LLM Trajectory Dataset Project Page | Paper | GitHub | Blog This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation". Introduction d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.tabulartext-generation100K<n<1M0 likes819 downloads4mo agoHugging Face12initialneil /DREAMS-AVATAR DREAMS-AVATAR The DREAMS-Avatar dataset from the DEGAS paper (3DV 2025), re-registered to pure SMPL-X. These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS (Fig. 1b); what is new here is the registration. 32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted to all views at once by our multiview tracker: 300 shape coefficients, 100 expression coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.imageimage-to-3dn<1K0 likes793 downloads2mo agoHugging Face13Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes742 downloads4y agoHugging Face14WenhaoWang /D-Rep Summary This is the dataset proposed in our paper Image Copy Detection for Diffusion Models (NeurIPS 2024). D-Rep consists of 40, 000 image-replica pairs, in which each replica is generated by a diffusion model. The 40, 000 image-replica pairs are manually labeled with 6 replication levels ranging from 0 (no replication) to 5 (total replication). We divide D-Rep into a training set with 90% (36, 000) pairs and a test set with the remaining 10% (4, 000) pairs.… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/D-Rep.imagetext-to-image10K<n<100K3 likes663 downloads2y agoHugging Face15xiabs /DreamOmni3Benchimagen<1K1 likes648 downloads2mo agoHugging Face16omni-research /DREAM-1Kgated DREAM-1K DREAM-1K (Description with Rich Events, Actions, and Motions) is a challenging video description benchmark. It contains a collection of 1,000 short (around 10 seconds) video clips with diverse complexities from five different origins: live-action movies, animated movies, stock videos, long YouTube videos, and TikTok-style short videos. We provide a fine-grained manual annotation for each video. Bellow is the dataset statistics: tabular1K<n<10K34 likes577 downloads2y agoHugging Face17dreamproit /us-statutes-at-large United States Statutes at Large Bound volumes of the United States Statutes at Large as published by the U.S. Government Publishing Office on GovInfo, collection STATUTE. Each volume is one PDF and one USLM XML file. Files metadata.jsonl one row per volume pdfs/STATUTE-{n}.pdf volume PDF as served by GovInfo xmls/STATUTE-{n}.xml volume USLM XML as served by GovInfo granules/ per-law PDFs used by the conversion benchmark (a sample… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/us-statutes-at-large.documentn<1K0 likes526 downloads17d agoHugging Face18ljnlonoljpiljm /dreamlip-gpt4v-500kimage100K<n<1M0 likes394 downloads2y agoHugging Face19bsaenz /dreamt Dataset Description DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices. Version: 2.1.0 Repository: PhysioNet: DREAMT v2.1.0 Access Policy & Licensing Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.tabularother100M<n<1B4 likes392 downloads6mo agoHugging Face20davidelobba /Dress-EDgated Dress-ED Instruction-Guided Editing for Virtual Try-On and Try-Off Overview Dress-ED is a large-scale garment-editing dataset for instruction-guided virtual try-on and clothing manipulation research. Given a person image and a garment image, each example asks a model to apply a specific textual edit to the garment while keeping it worn on the person. The dataset provides three complementary instructions per example — one relative to the garment… See the full description on the dataset page: https://huggingface.co/datasets/davidelobba/Dress-ED.imageimage-to-image100K<n<1M1 likes334 downloads1mo agoHugging Face21DrewLab /DirectContacts2 DirectContacts2: A network of direct physical protein interactions derived from high throughput mass spectrometry experiments Proteins carry out cellular functions by self-assembling into functional complexes, a process that depends on direct physical interactions between components. While tools like AlphaFold and RoseTTAFold have advanced structure prediction, they remain limited in scaling to the full human proteome. DirectContacts2 addresses this challenge by integrating… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/DirectContacts2.tabular10K<n<100K0 likes319 downloads3mo agoHugging Face22DreamVu /PRISM-100Kgated PRISM: Multi-View Multi-Capability Video SFT Dataset for Retail Embodied AI Dataset Details Dataset Description PRISM is a video Supervised Fine-Tuning (SFT) dataset designed for training Vision-Language Models (VLMs) on retail-domain physical AI tasks. It features synchronized egocentric and exocentric video from real retail environments, annotated across 21 task types spanning embodied reasoning, common-sense reasoning, spatial perception, and… See the full description on the dataset page: https://huggingface.co/datasets/DreamVu/PRISM-100K.textvideo-text-to-text100K<n<1M7 likes316 downloads5mo agoHugging Face23andreagasparini /dreaddit Dreaddit: A Reddit Dataset for Stress Analysis in Social Media Consists of 3.5k labeled texts from five different categories of Reddit communities. Citation @inproceedings{turcan-mckeown-2019-dreaddit, title = "{D}readdit: A {R}eddit Dataset for Stress Analysis in Social Media", author = "Turcan, Elsbeth and McKeown, Kathy", editor = "Holderness, Eben and Jimeno Yepes, Antonio and Lavelli, Alberto and Minard, Anne-Lyse and… See the full description on the dataset page: https://huggingface.co/datasets/andreagasparini/dreaddit.tabular1K<n<10K4 likes314 downloads1y agoHugging Face24takara-ai /sangyo_no_yume_industrial_dreams From the Frontier Research Team at Takara.ai we present the "Sangyo no Yume Industrial Dreams" dataset, a collection of AI-generated industrial dreamscapes. Sangyo no Yume Industrial Dreams Dataset Details "Sangyo no Yume Industrial Dreams" is a collection of images generated using SDXL Lightning with specialized prompt engineering techniques. These images balance industrial themes with dreamlike qualities, creating a unique aesthetic that sits at the intersection of… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/sangyo_no_yume_industrial_dreams.image1K<n<10K2 likes301 downloads2y agoHugging Face25drelhaj /Arabic-Dialects Arabic Dialects Dataset (Bivalency & Code-Switching) The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.texttext-classification10K<n<100K4 likes301 downloads10mo agoHugging Face26drelhaj /EASC EASC: The Essex Arabic Summaries Corpus Mo El-Haj, Udo Kruschwitz, Chris FoxUniversity of Essex, UK This repository hosts EASC — the Essex Arabic Summaries Corpus — a collection of 153 Arabic source documents and 765 human-generated extractive summaries, created using Amazon Mechanical Turk. EASC is one of the earliest publicly available datasets for Arabic single-document summarisation and remains widely used in research on Arabic NLP, extractive summarisation, sentence ranking… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/EASC.textsummarization1K<n<10K0 likes295 downloads10mo agoHugging Face27Eurong2 /dreamdojo-ego-view DreamDojo ego-view — video + instruction for Cosmos-Predict2.5 post-training 3,168 ego-view robot manipulation clips in the flat videos/ + metas/ layout that cosmos-predict2.5's VideoDataset reads directly, with no conversion step. Only the observation.images.ego_view camera is included. Episodes shorter than 94 frames are excluded: VideoDataset samples a random 93-frame window, and on a 93-frame video its np.random.randint(0, 0) raises rather than returning 0.… See the full description on the dataset page: https://huggingface.co/datasets/Eurong2/dreamdojo-ego-view.textrobotics1K<n<10K0 likes258 downloads15d agoHugging Face28vrfai /data_libero_dreamgen_stage1_s10This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 200, "total_frames": 33587, "total_tasks": 40, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:200" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/data_libero_dreamgen_stage1_s10.imagerobotics10K<n<100K0 likes250 downloads3mo agoHugging Face29zluvolyote /DreamNLPtext1M<n<10M0 likes231 downloads4y agoHugging Face30zluvolyote /Dream_Traintext1M<n<10M0 likes226 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.