CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01d3LLM /trajectory_data_dream_32 d3LLM Trajectory Dataset Project Page | Paper | GitHub | Blog This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation". Introduction d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.tabulartext-generation100K<n<1M0 likes819 downloads4mo agoHugging Face02drelhaj /Tarab Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry Tarab is a large-scale Arabic creative-text corpus that unifies song lyrics and poetry in a single verse-level representation.It contains 2,557,311 verses and 13,509,336 tokens, spanning Classical Arabic, MSA, and six major regional dialect groups, and covering both modern countries and historical eras. Dataset Overview Each row corresponds to a single verse with structured metadata linking it to its parent… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Tarab.text-classification10M<n<100M4 likes241 downloads7mo agoHugging Face03drewparo /bigquery-swift-unfiltered GitHub Swift Repositories Dataset Description Dataset Summary This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license. Source Data Initial Data Collection and Normalization The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.tabulartext-generation100K<n<1M1 likes223 downloads3y agoHugging Face04Lots-of-LoRAs /task246_dream_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task246_dream_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task246_dream_question_generation.texttext-generation1K<n<10K0 likes158 downloads2y agoHugging Face05huang3eng /Dressage-Claw Dressage-Claw Dressage-Claw is a synthetic collection of 441 tool-use tasks for black-box agent reinforcement learning and evaluation. It is designed for Accio-Lab/Dressage, with OpenClaw as the agent harness and deterministic local mock HTTP services as the execution environment. Each task bundles its prompt, tool definitions, service fixtures, lifecycle scripts, workspace, and grader. Tasks require agents to retrieve, reconcile, or update state while respecting explicit safety… See the full description on the dataset page: https://huggingface.co/datasets/huang3eng/Dressage-Claw.tabulartext-generationn<1K0 likes97 downloads2mo agoHugging Face06Lots-of-LoRAs /task247_dream_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task247_dream_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task247_dream_answer_generation.texttext-generation1K<n<10K0 likes95 downloads2y agoHugging Face07dreamproit /bill_text_us Dataset Card for "bill_text_us" Dataset Summary Dataset for US Congressional bills (bill_text_us). Supported Tasks and Leaderboards More Information Needed Languages English Dataset Structure Data Instances default Data Fields id: id of the bill in format(congress number + bill type + bill number + bill version). congress: number of the congress. bill_type: type of the bill. bill_number: number of the… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_text_us.tabulartext-generation100K<n<1M2 likes87 downloads3y agoHugging Face08PursuitOfDataScience /dream-of-the-red-chamber-continuations 红楼梦续写 · Dream of the Red Chamber: 100 AI Continuations 项目简介 本数据集包含 92 个独立的AI续写版本,续写中国古典文学巅峰之作《红楼梦》的第八十一回至第一百零八回(共28回)。所有续写严格遵循曹雪芹前八十回中埋下的伏笔、谶语和人物命运,完全拒绝高鹗续书。 为什么做这个数据集 《红楼梦》的结局是世界文学史上最大的悬案之一。曹雪芹约于1763年去世前未能完成全书,仅留下前八十回。1791年左右,高鹗发表了一百二十回本,补写了后四十回,但红学研究日益表明高鹗续书严重违背了曹雪芹在前八十回中精心布置的伏笔。 曹雪芹原意 vs 高鹗续书 情节 曹雪芹原意 高鹗续书 黛玉之死 泪尽而亡,呼应"绛珠还泪"神话 焚稿断痴情 宝玉宝钗婚姻 "纵然是齐眉举案,到底意难平" 掉包计骗婚 贾府败落 政治牵连,锦衣军抄家,"忽喇喇似大厦倾" 败而复兴,"兰桂齐芳" 结局… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/dream-of-the-red-chamber-continuations.tabulartext-generationn<1K2 likes80 downloads6mo agoHugging Face09S-Dreamer /ai-governance-synthetic-glm5 AI Governance Synthetic Dataset (GLM-5.3-Flash) A ~1,000-example synthetic dataset on AI governance and frontier AI safety, generated with zai-org/GLM-5.3-Flash via the Hugging Face Inference Providers API. Configs Config Rows Schema Use policy_qa 400 messages (user/assistant chat), topic SFT of governance assistants risk_classification 300 scenario, risk_category (10-way enum), severity (low/medium/high/critical), rationale, topic Training risk… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/ai-governance-synthetic-glm5.textquestion-answering1K<n<10K0 likes78 downloads9d agoHugging Face10gustavecortal /DreamBank-annotated Presentation DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English. Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper: Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.texttext-generation10K<n<100K12 likes75 downloads8mo agoHugging Face11trentmkelly /dread-crime-forum Dread Forum Archive A near-complete capture of public content from Dread, a Reddit-style discussion forum hosted as a Tor hidden service. Dread is one of the longest-running darknet community forums and a primary site for discussion of darknet markets, operational security, cryptocurrency, and related topics. The archive covers content posted between April 2018 and September 2025 and is structured as three Parquet-backed splits: posts, comments, and users (with parsed PGP key… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/dread-crime-forum.tabulartext-classification1M<n<10M1 likes70 downloads4mo agoHugging Face12dreamproit /bill_labels_us Dataset Card for "bill_labels_us" Dataset Summary Dataset for US Congressional bills with policy area and legislative subjects information (bill_labels_us). Contains data for bills from the 108th to the 118th Congress, approximately 119,000 documents. Supported Tasks and Leaderboards More Information Needed Languages English Dataset Structure Data Instances default Data Fields id: id of the bill in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_labels_us.tabulartext-generation100K<n<1M6 likes57 downloads2y agoHugging Face13drewoodward /spanglish-sentences Spanglish Sentences A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models. Data format Each line of spanglish_sentences.jsonl is a JSON object with two fields: field description sentence A Spanglish utterance (mixed Spanish / English, or monolingual in either language). english_translation The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.texttranslation10K<n<100K0 likes51 downloads5mo agoHugging Face14dreeseaw /cleo-process-analytics-v1 Cleo Process Analytics v1 cleo-process-analytics-v1 is a 260-example SQL analytics dataset built for process-heavy analyst workflows. The questions are designed to require multi-step SQL behavior such as joins, aggregations, rankings, CTEs, windows, and occasional semantic-view use, while keeping answers deterministic and execution-verified. This dataset was created for the Cleo SQL analyst project: github.com/Dreeseaw/cleo. Contents path description… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-process-analytics-v1.text-generationn<1K0 likes46 downloads4mo agoHugging Face15Zigeng /dParallel_Dream_Distill_Data dParallel-Dream-Distill Dataset: This dataset is used for the certainty-forcing distillation process in dParallel. We use prompts from publicly available training datasets and let the pretrained model generate its own responses as training data. For LLaDA-8B-Instruct, we sample prompts from the GSM8K, PRM12K training set, and part of the Numina-Math dataset. We generate target trajectories using a semi-autoregressive strategy with a sequence length of 256 and block length of 32. We… See the full description on the dataset page: https://huggingface.co/datasets/Zigeng/dParallel_Dream_Distill_Data.texttext-generation100K<n<1M2 likes42 downloads1y agoHugging Face16S-Dreamer /my-distiset-3be4288b S-Dreamer/my-distiset-3be4288b Overview This synthetic dataset is designed for multiple natural language processing tasks, including Text Generation, Text2Text Generation, and Question Answering. With a lightweight size (fewer than 1K rows) and an auto-converted Parquet format, it is ideal for rapid prototyping, model development, and educational experiments. Key Details Modalities: Text Format: Parquet Size: < 1K rows Tags: Synthetic, distilabel, rlaif, datacraft… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/my-distiset-3be4288b.text-generationn<1K1 likes40 downloads1y agoHugging Face17DReAMy-lib /DreamBank-dreams DreamBank - Dreams The dataset is a collection of ~30k textual reports of dreams, originally scraped from the DreamBank databased by mattbierner. The DreamBank reports are divided into series, which are collections of individuals or research projects/groups that have gathered the dreams. The vast majority of the series are in the English language, but a small part of the are in German. These series are indicated by the presence of .de in their name. Content The… See the full description on the dataset page: https://huggingface.co/datasets/DReAMy-lib/DreamBank-dreams.texttext-generation10K<n<100K1 likes36 downloads4y agoHugging Face18mujo-labs /sandman-dream_multitask_train Sandman dream multitask v1 — train split The first version of the train split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation100K<n<1M0 likes36 downloads7d agoHugging Face19mujo-labs /sandman-dream_multitask_test Sandman dream multitask v1 — test split The first version of the test split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation10K<n<100K0 likes36 downloads7d agoHugging Face20drelhaj /KALIMAT Kalimat - a multipurpose Arabic Corpus This repository provides a cleaned and consolidated version of the Kalimat - a multipurpose Arabic Corpus, containing 18,256 Arabic news articles collected from a diverse range of domains. The original material consisted of thousands of individual .txt files organised across multiple category folders. These have been reconstructed, normalised, and compiled into modern machine-learning-friendly formats. 📚 Corpus Overview The… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/KALIMAT.texttext-generation10K<n<100K0 likes35 downloads10mo agoHugging Face21mujo-labs /sandman-dream_multitask_val Sandman dream multitask v1 — val split The first version of the val split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation10K<n<100K0 likes32 downloads7d agoHugging Face22DreamingBumblebee /ultrachat-100-ko Dataset Card for ultrachat-mini-ko Dataset Description This is a mini translated version of the UltraChat 200k. @misc{ding2023enhancing, title={Enhancing Chat Language Models by Scaling High-quality Instructional Conversations}, author={Ning Ding and Yulin Chen and Bokai Xu and Yujia Qin and Zhi Zheng and Shengding Hu and Zhiyuan Liu and Maosong Sun and Bowen Zhou}, year={2023}, eprint={2305.14233}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/DreamingBumblebee/ultrachat-100-ko.texttext-generationn<1K0 likes29 downloads3y agoHugging Face23mujo-labs /sandman-dream_multitask_v2_train Sandman dream multitask v2 — train split 17,300 instruction-following examples for fine-tuning Sandman's on-device dream-analysis model, built from sandman-dreambank-v2. Every row is a single-turn conversation (messages) covering one of three tasks: Summarize — read a dream, return a one- or two-sentence summary as JSON. Extract symbols — return only the concrete nouns literally present in the dream text, as a JSON array, with an explicit instruction not to infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.texttext-generation10K<n<100K0 likes27 downloads7d agoHugging Face24dreeseaw /cleo-value-discovery Cleo Value-Discovery Benchmark A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss: questions whose correct SQL depends on a literal that lives in the data, not the schema. The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}. The schema shows to_date; only the data reveals that "current" is encoded as the sentinel '9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.texttable-question-answeringn<1K0 likes23 downloads4mo agoHugging Face25mujo-labs /sandman-dream_multitask_v2_test Sandman dream multitask v2 — test split The test split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes19 downloads7d agoHugging Face26mujo-labs /sandman-dreambank-v2 Sandman dreambank v2 5,220 dream reports drawn from DreamBank, the dream-report archive maintained by Dr. G. William Domhoff and Adam Schneider for dream research, each enriched with structured annotations: a title, one-line summary, mood, 2-4 themes, the concrete symbols mentioned (people, places, things), a meaning for each symbol, and two styles of written interpretation (oracle_interpretation, more evocative; brief_interpretation, more grounded). _meta on every row carries… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dreambank-v2.texttext-generation1K<n<10K0 likes18 downloads7d agoHugging Face27mujo-labs /sandman-dream_multitask_v2_val Sandman dream multitask v2 — val split The val split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes17 downloads7d agoHugging Face28Lots-of-LoRAs /task283_dream_incorrect_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task283_dream_incorrect_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task283_dream_incorrect_answer_generation.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face29dreamproit /bill_committees_us Dataset Card for "bill_committees_us" Dataset Summary Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents. Supported Tasks and Leaderboards More Information Needed Languages English Dataset Structure Data Instances default Data Fields id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.tabulartext-generation100K<n<1M5 likes14 downloads2y agoHugging Face30JuanMallo /dream-75pct DREAM: Dialogue to REAlistic Multicultural Image Sequences DREAM is a multicultural multimodal dataset linking persona-grounded dialogues with photorealistic portrait images and storyboard-like dialogue scene sequences. The dataset was introduced in: DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image SequencesLREC 2026. Dataset Overview DREAM is a fully synthetic multimodal resource designed to support research on: visual grounding of… See the full description on the dataset page: https://huggingface.co/datasets/JuanMallo/dream-75pct.text-to-image1K<n<10K0 likes14 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.