datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trajectory_data_dream_32
d3LLM Trajectory Dataset
Project Page | Paper | GitHub | Blog
This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation".
Introduction
d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.Tarab
Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry
Tarab is a large-scale Arabic creative-text corpus that unifies song lyrics and poetry in a single verse-level representation.It contains 2,557,311 verses and 13,509,336 tokens, spanning Classical Arabic, MSA, and six major regional dialect groups, and covering both modern countries and historical eras.
Dataset Overview
Each row corresponds to a single verse with structured metadata linking it to its parent… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Tarab.bigquery-swift-unfiltered
GitHub Swift Repositories
Dataset Description
Dataset Summary
This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license.
Source Data
Initial Data Collection and Normalization
The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.task246_dream_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task246_dream_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task246_dream_question_generation.Dressage-Claw
Dressage-Claw
Dressage-Claw is a synthetic collection of 441 tool-use tasks for black-box
agent reinforcement learning and evaluation. It is designed for
Accio-Lab/Dressage, with OpenClaw as
the agent harness and deterministic local mock HTTP services as the execution
environment.
Each task bundles its prompt, tool definitions, service fixtures, lifecycle
scripts, workspace, and grader. Tasks require agents to retrieve, reconcile, or
update state while respecting explicit safety… See the full description on the dataset page: https://huggingface.co/datasets/huang3eng/Dressage-Claw.task247_dream_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task247_dream_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task247_dream_answer_generation.bill_text_us
Dataset Card for "bill_text_us"
Dataset Summary
Dataset for US Congressional bills (bill_text_us).
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number + bill type + bill number + bill version).
congress: number of the congress.
bill_type: type of the bill.
bill_number: number of the… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_text_us.dream-of-the-red-chamber-continuations
红楼梦续写 · Dream of the Red Chamber: 100 AI Continuations
项目简介
本数据集包含 92 个独立的AI续写版本,续写中国古典文学巅峰之作《红楼梦》的第八十一回至第一百零八回(共28回)。所有续写严格遵循曹雪芹前八十回中埋下的伏笔、谶语和人物命运,完全拒绝高鹗续书。
为什么做这个数据集
《红楼梦》的结局是世界文学史上最大的悬案之一。曹雪芹约于1763年去世前未能完成全书,仅留下前八十回。1791年左右,高鹗发表了一百二十回本,补写了后四十回,但红学研究日益表明高鹗续书严重违背了曹雪芹在前八十回中精心布置的伏笔。
曹雪芹原意 vs 高鹗续书
情节
曹雪芹原意
高鹗续书
黛玉之死
泪尽而亡,呼应"绛珠还泪"神话
焚稿断痴情
宝玉宝钗婚姻
"纵然是齐眉举案,到底意难平"
掉包计骗婚
贾府败落
政治牵连,锦衣军抄家,"忽喇喇似大厦倾"
败而复兴,"兰桂齐芳"
结局… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/dream-of-the-red-chamber-continuations.ai-governance-synthetic-glm5
AI Governance Synthetic Dataset (GLM-5.3-Flash)
A ~1,000-example synthetic dataset on AI governance and frontier AI safety, generated with zai-org/GLM-5.3-Flash via the Hugging Face Inference Providers API.
Configs
Config
Rows
Schema
Use
policy_qa
400
messages (user/assistant chat), topic
SFT of governance assistants
risk_classification
300
scenario, risk_category (10-way enum), severity (low/medium/high/critical), rationale, topic
Training risk… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/ai-governance-synthetic-glm5.DreamBank-annotated
Presentation
DreamBank, an open corpus of more than 27,000 dream narratives, mostly written in English.
Annotations were produced using dream-t5, a LaMini-Flan-T5 model finetuned on Hall and Van de Castle annotations to predict character and emotion. I've introduced this task in this paper:
Gustave Cortal. 2024. Sequence-to-Sequence Language Models for Character and Emotion Detection in Dream Narratives. In Proceedings of the 2024 Joint International Conference on Computational… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/DreamBank-annotated.dread-crime-forum
Dread Forum Archive
A near-complete capture of public content from Dread, a Reddit-style discussion forum hosted as a Tor hidden service. Dread is one of the longest-running darknet community forums and a primary site for discussion of darknet markets, operational security, cryptocurrency, and related topics.
The archive covers content posted between April 2018 and September 2025 and is structured as three Parquet-backed splits: posts, comments, and users (with parsed PGP key… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/dread-crime-forum.bill_labels_us
Dataset Card for "bill_labels_us"
Dataset Summary
Dataset for US Congressional bills with policy area and legislative subjects information (bill_labels_us). Contains data for bills from the 108th to the 118th Congress, approximately 119,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_labels_us.spanglish-sentences
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.cleo-process-analytics-v1
Cleo Process Analytics v1
cleo-process-analytics-v1 is a 260-example SQL analytics dataset built for process-heavy analyst
workflows. The questions are designed to require multi-step SQL behavior such as joins, aggregations,
rankings, CTEs, windows, and occasional semantic-view use, while keeping answers deterministic and
execution-verified.
This dataset was created for the Cleo SQL analyst project:
github.com/Dreeseaw/cleo.
Contents
path
description… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-process-analytics-v1.dParallel_Dream_Distill_Data
dParallel-Dream-Distill Dataset:
This dataset is used for the certainty-forcing distillation process in dParallel. We use prompts from publicly available training datasets and let the pretrained model generate its own responses as training data. For LLaDA-8B-Instruct, we sample prompts from the GSM8K, PRM12K training set, and part of the Numina-Math dataset. We generate target trajectories using a semi-autoregressive strategy with a sequence length of 256 and block length of 32. We… See the full description on the dataset page: https://huggingface.co/datasets/Zigeng/dParallel_Dream_Distill_Data.my-distiset-3be4288b
S-Dreamer/my-distiset-3be4288b
Overview
This synthetic dataset is designed for multiple natural language processing tasks, including Text Generation, Text2Text Generation, and Question Answering. With a lightweight size (fewer than 1K rows) and an auto-converted Parquet format, it is ideal for rapid prototyping, model development, and educational experiments.
Key Details
Modalities: Text
Format: Parquet
Size: < 1K rows
Tags: Synthetic, distilabel, rlaif, datacraft… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/my-distiset-3be4288b.DreamBank-dreams
DreamBank - Dreams
The dataset is a collection of ~30k textual reports of dreams, originally scraped from the DreamBank databased by
mattbierner. The DreamBank reports are divided into series,
which are collections of individuals or research projects/groups that have gathered the dreams. The vast majority of the series are in the
English language, but a small part of the are in German. These series are indicated by the presence of .de in their name.
Content
The… See the full description on the dataset page: https://huggingface.co/datasets/DReAMy-lib/DreamBank-dreams.sandman-dream_multitask_train
Sandman dream multitask v1 — train split
The first version of the train split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
sandman-dream_multitask_test
Sandman dream multitask v1 — test split
The first version of the test split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
KALIMAT
Kalimat - a multipurpose Arabic Corpus
This repository provides a cleaned and consolidated version of the Kalimat - a multipurpose Arabic Corpus, containing 18,256 Arabic news articles collected from a diverse range of domains. The original material consisted of thousands of individual .txt files organised across multiple category folders. These have been reconstructed, normalised, and compiled into modern machine-learning-friendly formats.
📚 Corpus Overview
The… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/KALIMAT.sandman-dream_multitask_val
Sandman dream multitask v1 — val split
The first version of the val split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
ultrachat-100-ko
Dataset Card for ultrachat-mini-ko
Dataset Description
This is a mini translated version of the UltraChat 200k.
@misc{ding2023enhancing,
title={Enhancing Chat Language Models by Scaling High-quality Instructional Conversations},
author={Ning Ding and Yulin Chen and Bokai Xu and Yujia Qin and Zhi Zheng and Shengding Hu and Zhiyuan Liu and Maosong Sun and Bowen Zhou},
year={2023},
eprint={2305.14233},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/DreamingBumblebee/ultrachat-100-ko.sandman-dream_multitask_v2_train
Sandman dream multitask v2 — train split
17,300 instruction-following examples for fine-tuning Sandman's on-device
dream-analysis model, built from
sandman-dreambank-v2.
Every row is a single-turn conversation (messages) covering one of three
tasks:
Summarize — read a dream, return a one- or two-sentence summary as JSON.
Extract symbols — return only the concrete nouns literally present in
the dream text, as a JSON array, with an explicit instruction not to
infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.cleo-value-discovery
Cleo Value-Discovery Benchmark
A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss:
questions whose correct SQL depends on a literal that lives in the data, not the schema.
The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}.
The schema shows to_date; only the data reveals that "current" is encoded as the sentinel
'9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.sandman-dream_multitask_v2_test
Sandman dream multitask v2 — test split
The test split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
sandman-dreambank-v2
Sandman dreambank v2
5,220 dream reports drawn from DreamBank, the
dream-report archive maintained by Dr. G. William Domhoff and Adam Schneider
for dream research, each enriched with structured annotations: a title,
one-line summary, mood, 2-4 themes, the concrete symbols mentioned (people,
places, things), a meaning for each symbol, and two styles of written
interpretation (oracle_interpretation, more evocative; brief_interpretation,
more grounded).
_meta on every row carries… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dreambank-v2.sandman-dream_multitask_v2_val
Sandman dream multitask v2 — val split
The val split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
task283_dream_incorrect_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task283_dream_incorrect_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task283_dream_incorrect_answer_generation.bill_committees_us
Dataset Card for "bill_committees_us"
Dataset Summary
Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.dream-75pct
DREAM: Dialogue to REAlistic Multicultural Image Sequences
DREAM is a multicultural multimodal dataset linking persona-grounded dialogues with photorealistic portrait images and storyboard-like dialogue scene sequences.
The dataset was introduced in:
DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image SequencesLREC 2026.
Dataset Overview
DREAM is a fully synthetic multimodal resource designed to support research on:
visual grounding of… See the full description on the dataset page: https://huggingface.co/datasets/JuanMallo/dream-75pct.
