CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01drengskapur /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.text1K<n<10K19 likes9.9k downloads2y agoHugging Face02Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes742 downloads4y agoHugging Face03bsaenz /dreamt Dataset Description DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices. Version: 2.1.0 Repository: PhysioNet: DREAMT v2.1.0 Access Policy & Licensing Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.tabularother100M<n<1B4 likes392 downloads6mo agoHugging Face04DrewLab /DirectContacts2 DirectContacts2: A network of direct physical protein interactions derived from high throughput mass spectrometry experiments Proteins carry out cellular functions by self-assembling into functional complexes, a process that depends on direct physical interactions between components. While tools like AlphaFold and RoseTTAFold have advanced structure prediction, they remain limited in scaling to the full human proteome. DirectContacts2 addresses this challenge by integrating… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/DirectContacts2.tabular10K<n<100K0 likes319 downloads3mo agoHugging Face05drelhaj /Arabic-Dialects Arabic Dialects Dataset (Bivalency & Code-Switching) The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.texttext-classification10K<n<100K4 likes301 downloads10mo agoHugging Face06zluvolyote /Dream_Traintext1M<n<10M0 likes226 downloads4y agoHugging Face07zluvolyote /DREAM_SAMPLE_600Ktext1M<n<10M0 likes193 downloads4y agoHugging Face08dreuxx26 /russian_gec 📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs) A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing. ✨ Dataset Summary Metric Value Sentence pairs 25 362 Avg. tokens / sentence ≈ 12 File size ~5 MB (CSV, UTF‑8) Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.text10K<n<100K3 likes123 downloads1y agoHugging Face09drelhaj /Arabic-news-and-management-corpus Arabic Management, Economics & Financial News Corpus (1,200 Articles) This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available. The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.texttext-classification1K<n<10K0 likes120 downloads10mo agoHugging Face10zluvolyote /Dream_NLP_FineTunetabular100K<n<1M0 likes104 downloads4y agoHugging Face11drelhaj /AraFinNews AraFinNews: The Arabic Financial News Dataset (212K) For the JSON file format please check our AraFinNews GitHub repo AraFinNews is the largest openly available dataset of Arabic financial news, comprising 212,500 full-length articles collected from Argaam.com — a leading financial news portal in the Arab world.The dataset provides structured, machine-readable text suitable for research in financial NLP, abstractive summarisation, event extraction, and domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/AraFinNews.textsummarization100K<n<1M1 likes81 downloads10mo agoHugging Face12qidouxiong619 /dreamlip_long_captions Dataset Card for DreamLIP-30M Dataset Summary DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.imagetext-to-image10M<n<100M19 likes77 downloads2y agoHugging Face13gustavecortal /DreamBank-annotated Presentation DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English. Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper: Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.texttext-generation10K<n<100K12 likes75 downloads8mo agoHugging Face14DREAM-17k /DREAM-Red-Teaming-Promptsgated DREAM: Scalable Red Teaming for Text-to-Image Generative Systems This dataset contains red-teaming prompts generated by the DREAM framework, as presented in the IEEE S&P 2026 paper: "DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling". These prompts are designed to evaluate and stress-test the safety mechanisms of Text-to-Image (T2I) generative systems. ⚠️ Disclaimer / WarningContent Warning: This dataset contains prompts that may be… See the full description on the dataset page: https://huggingface.co/datasets/DREAM-17k/DREAM-Red-Teaming-Prompts.text10K<n<100K6 likes70 downloads10mo agoHugging Face15DrewLab /hu.MAP_3.0 hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments. Proteins interact with each other and organize themselves into macromolecular machines (ie. complexes) to carry out essential functions of the cell. We have a good understanding of a few complexes such as the proteasome and the ribosome but currently we have an incomplete view of all protein complexes as well as their functions. The hu.MAP attempts to address this lack of understanding… See the full description on the dataset page: https://huggingface.co/datasets/DrewLab/hu.MAP_3.0.tabular100K<n<1M0 likes56 downloads1y agoHugging Face16asmaab /dreaddittabular1K<n<10K3 likes55 downloads2y agoHugging Face17drelhaj /ArabJobs ArabJobs: A Multinational Corpus of Arabic Job Advertisements 📖 Overview ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE. It contains: 8,546 job postings 550,000+ words Coverage across numerous sectors and dialects Rich metadata including salary, profession, gender indicators, and job categories This dataset supports research on: Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.tabulartext-classification1K<n<10K0 likes49 downloads10mo agoHugging Face18Dreaming668 /A-share-market_ETF-daily-datagatedtabular1M<n<10M1 likes43 downloads1y agoHugging Face19nlpatunt /D_Regrading_Dataset_J2C Regrading_Dataset_2JC This dataset contains short-answer responses with rubric-based grades, shared directly by the authors for research use. It is prepared for use with the S-GRADES benchmark. This is the train, test, and validation split. Ground truth labels of test split have been removed to prevent leakage during evaluation. Citation If you use this dataset, please cite the original authors: @article{gao2024towards, title={Towards scalable automated grading:… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Regrading_Dataset_J2C.tabular1K<n<10K0 likes35 downloads7mo agoHugging Face20drengskapur /household_power_cleantabular1M<n<10M0 likes31 downloads3y agoHugging Face21Takaharadesu /BanG_Dream_150k BanG Dream Dialogue 150K A large-scale multilingual dialogue corpus featuring 150,000+ interactions across 40 characters from the BanG Dream!本数据集收录 BanG Dream! 系列 40 角色 的 150,000+ 条对话,适用于自然语言处理任务,深度还原角色性格与互动模式。 Dataset Statistics|统计 · ako: 4854条 · anon: 2251条 · arisa: 6713条 · aya: 5479条 · chisato: 5074条 · chuchu: 2432条 · eve: 4639条 · hagumi: 4144条 · himari: 5266条 · hina: 5102条 · kanon: 4022条 · kaoru: 4083条 · kasumi: 7233条 · kokoro: 4493条 · layer: 2225条 · lisa: 6057条 · lock:… See the full description on the dataset page: https://huggingface.co/datasets/Takaharadesu/BanG_Dream_150k.texttranslation100K<n<1M1 likes31 downloads1y agoHugging Face22zluvolyote /Dream_NLP_Validationtabular100K<n<1M0 likes30 downloads4y agoHugging Face23dreamland4dnam /targettabular1M<n<10M0 likes30 downloads1mo agoHugging Face24DreamyP /MedQAtext1K<n<10K0 likes27 downloads3y agoHugging Face25Dregandor /Edgar-Cayce_Readingstabularquestion-answering10K<n<100K0 likes20 downloads3y agoHugging Face26asmaab /dreaddit_traintext1K<n<10K0 likes18 downloads2y agoHugging Face27kumpank /dream-observatory Dream Observatory Dream Observatory is a structured dataset of 1,646 publicly available dream narratives transformed into 16 cognitive and phenomenological dimensions using an automated LLM-based scoring pipeline. Unlike a static labeled corpus, Dream Observatory is designed as a living dataset: the underlying pipeline continuously collects new public dream reports, performs quality validation, extracts structured cognitive features, and can be re-run to produce updated versions… See the full description on the dataset page: https://huggingface.co/datasets/kumpank/dream-observatory.document1K<n<10K1 likes17 downloads2mo agoHugging Face28DreamsHunter /edueval-benchmarkstabularn<1K0 likes15 downloads2mo agoHugging Face29MihaiIonascu /dreadit-validationtextn<1K0 likes14 downloads3y agoHugging Face30dreamland4dnam /cptac-3tabular100K<n<1M0 likes14 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.