datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dream-coder
Program Synthesis Data
Generated program synthesis datasets used to train dreamcoder.
Currently just supports text & list data.
dreamt
Dataset Description
DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices.
Version: 2.1.0
Repository: PhysioNet: DREAMT v2.1.0
Access Policy & Licensing
Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.Dream_TrainDREAM_SAMPLE_600KDream_NLP_FineTunetargetdreamlip_long_captions
Dataset Card for DreamLIP-30M
Dataset Summary
DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.DreamBank-annotated
Presentation
DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English.
Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper:
Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.DREAM-Red-Teaming-Prompts
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems
This dataset contains red-teaming prompts generated by the DREAM framework, as presented in the IEEE S&P 2026 paper: "DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling". These prompts are designed to evaluate and stress-test the safety mechanisms of Text-to-Image (T2I) generative systems.
⚠️ Disclaimer / WarningContent Warning: This dataset contains prompts that may be… See the full description on the dataset page: https://huggingface.co/datasets/DREAM-17k/DREAM-Red-Teaming-Prompts.A-share-market_ETF-daily-datacptac-3hcmci-cmdcDream_NLP_ValidationBanG_Dream_150k
BanG Dream Dialogue 150K
A large-scale multilingual dialogue corpus featuring 150,000+ interactions across 40 characters from the BanG Dream!本数据集收录 BanG Dream! 系列 40 角色 的 150,000+ 条对话,适用于自然语言处理任务,深度还原角色性格与互动模式。
Dataset Statistics|统计
· ako: 4854条
· anon: 2251条
· arisa: 6713条
· aya: 5479条
· chisato: 5074条
· chuchu: 2432条
· eve: 4639条
· hagumi: 4144条
· himari: 5266条
· hina: 5102条
· kanon: 4022条
· kaoru: 4083条
· kasumi: 7233条
· kokoro: 4493条
· layer: 2225条
· lisa: 6057条
· lock:… See the full description on the dataset page: https://huggingface.co/datasets/Takaharadesu/BanG_Dream_150k.cgci-htmcp-ccMedQAdream-databasedream-observatory
Dream Observatory
Dream Observatory is a structured dataset of 1,646 publicly available dream narratives transformed into 16 cognitive and phenomenological dimensions using an automated LLM-based scoring pipeline.
Unlike a static labeled corpus, Dream Observatory is designed as a living dataset: the underlying pipeline continuously collects new public dream reports, performs quality validation, extracts structured cognitive features, and can be re-run to produce updated versions… See the full description on the dataset page: https://huggingface.co/datasets/kumpank/dream-observatory.edueval-benchmarksencodeInterpretation-of-dreamsDREAM-1k-VLMEvalKitibnsirin-dreamtextdreambank-netDreamcoder_naming_flipNector_featuresmed_50MultiModal-Planning-Benchmarkhighway_fast_v0_dreamworldhighway_fast_v0_dreamworld_large
