datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.dream-coder
Program Synthesis Data
Generated program synthesis datasets used to train dreamcoder.
Currently just supports text & list data.
dreamt
Dataset Description
DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices.
Version: 2.1.0
Repository: PhysioNet: DREAMT v2.1.0
Access Policy & Licensing
Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.DirectContacts2
DirectContacts2: A network of direct physical protein interactions derived from high throughput mass spectrometry experiments
Proteins carry out cellular functions by self-assembling into functional complexes, a process that depends on direct physical interactions
between components. While tools like AlphaFold and RoseTTAFold have advanced structure prediction, they remain limited in scaling to the full
human proteome. DirectContacts2 addresses this challenge by integrating… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/DirectContacts2.Arabic-Dialects
Arabic Dialects Dataset (Bivalency & Code-Switching)
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties:
EGY – Egyptian Arabic
GLF – Gulf Arabic
LAV – Levantine Arabic
NOR – North African / Tunisian Arabic
MSA – Modern Standard Arabic
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.Dream_TrainDREAM_SAMPLE_600Krussian_gec
📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs)
A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing.
✨ Dataset Summary
Metric
Value
Sentence pairs
25 362
Avg. tokens / sentence
≈ 12
File size
~5 MB (CSV, UTF‑8)
Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.Arabic-news-and-management-corpus
Arabic Management, Economics & Financial News Corpus (1,200 Articles)
This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available.
The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.Dream_NLP_FineTuneAraFinNews
AraFinNews: The Arabic Financial News Dataset (212K)
For the JSON file format please check our AraFinNews GitHub repo
AraFinNews is the largest openly available dataset of Arabic financial news, comprising 212,500 full-length articles collected from Argaam.com — a leading financial news portal in the Arab world.The dataset provides structured, machine-readable text suitable for research in financial NLP, abstractive summarisation, event extraction, and domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/AraFinNews.dreamlip_long_captions
Dataset Card for DreamLIP-30M
Dataset Summary
DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.DreamBank-annotated
Presentation
DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English.
Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper:
Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.DREAM-Red-Teaming-Prompts
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems
This dataset contains red-teaming prompts generated by the DREAM framework, as presented in the IEEE S&P 2026 paper: "DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling". These prompts are designed to evaluate and stress-test the safety mechanisms of Text-to-Image (T2I) generative systems.
⚠️ Disclaimer / WarningContent Warning: This dataset contains prompts that may be… See the full description on the dataset page: https://huggingface.co/datasets/DREAM-17k/DREAM-Red-Teaming-Prompts.hu.MAP_3.0
hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments.
Proteins interact with each other and organize themselves into macromolecular machines (ie. complexes)
to carry out essential functions of the cell. We have a good understanding of a few complexes such as
the proteasome and the ribosome but currently we have an incomplete view of all protein complexes as
well as their functions. The hu.MAP attempts to address this lack of understanding… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/hu.MAP_3.0.dreadditArabJobs
ArabJobs: A Multinational Corpus of Arabic Job Advertisements
📖 Overview
ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE.
It contains:
8,546 job postings
550,000+ words
Coverage across numerous sectors and dialects
Rich metadata including salary, profession, gender indicators, and job categories
This dataset supports research on:
Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.A-share-market_ETF-daily-dataD_Regrading_Dataset_J2C
Regrading_Dataset_2JC
This dataset contains short-answer responses with rubric-based grades, shared directly by the authors for research use. It is prepared for use with the S-GRADES benchmark. This is the train, test, and validation split. Ground truth labels of test split have been removed to prevent leakage during evaluation.
Citation
If you use this dataset, please cite the original authors:
@article{gao2024towards,
title={Towards scalable automated grading:… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Regrading_Dataset_J2C.household_power_cleanBanG_Dream_150k
BanG Dream Dialogue 150K
A large-scale multilingual dialogue corpus featuring 150,000+ interactions across 40 characters from the BanG Dream!本数据集收录 BanG Dream! 系列 40 角色 的 150,000+ 条对话,适用于自然语言处理任务,深度还原角色性格与互动模式。
Dataset Statistics|统计
· ako: 4854条
· anon: 2251条
· arisa: 6713条
· aya: 5479条
· chisato: 5074条
· chuchu: 2432条
· eve: 4639条
· hagumi: 4144条
· himari: 5266条
· hina: 5102条
· kanon: 4022条
· kaoru: 4083条
· kasumi: 7233条
· kokoro: 4493条
· layer: 2225条
· lisa: 6057条
· lock:… See the full description on the dataset page: https://huggingface.co/datasets/Takaharadesu/BanG_Dream_150k.Dream_NLP_ValidationtargetMedQAEdgar-Cayce_Readingsdreaddit_traindream-observatory
Dream Observatory
Dream Observatory is a structured dataset of 1,646 publicly available dream narratives transformed into 16 cognitive and phenomenological dimensions using an automated LLM-based scoring pipeline.
Unlike a static labeled corpus, Dream Observatory is designed as a living dataset: the underlying pipeline continuously collects new public dream reports, performs quality validation, extracts structured cognitive features, and can be re-run to produce updated versions… See the full description on the dataset page: https://huggingface.co/datasets/kumpank/dream-observatory.edueval-benchmarksdreadit-validationcptac-3
