datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dreamDREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset. In contrast to existing reading comprehension datasets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding.uscode
United States Code, versioned by release point
Every section of the United States Code, as published by the Office of the Law
Revision Counsel (OLRC) at uscode.house.gov, across
every release point from 113-21 (July 18, 2013) through the present. A release
point is OLRC's republication of the Code after a batch of Public Laws is
classified; this dataset covers 381 of them over 58 titles.
Each row carries the section's plain text, its verbatim USLM XML, its
citation, its place in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/uscode.mdlens-realdocs-v1
mdlens-realdocs-v1
A held-out Markdown QA / retrieval eval built entirely from real open-source
project documentation. It measures whether an agent can answer documentation
questions from the right evidence with fewer irrelevant reads and fewer tokens.
Questions are deliberately low lexical overlap (paraphrased), so they stress
retrieval rather than string matching. Every non-abstention question has its
answer keywords verified to appear in the cited source file.… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/mdlens-realdocs-v1.DreamFlow-AI-Dataai-governance-synthetic-glm5
AI Governance Synthetic Dataset (GLM-5.3-Flash)
A ~1,000-example synthetic dataset on AI governance and frontier AI safety, generated with zai-org/GLM-5.3-Flash via the Hugging Face Inference Providers API.
Configs
Config
Rows
Schema
Use
policy_qa
400
messages (user/assistant chat), topic
SFT of governance assistants
risk_classification
300
scenario, risk_category (10-way enum), severity (low/medium/high/critical), rationale, topic
Training risk… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/ai-governance-synthetic-glm5.ArabJobs
ArabJobs: A Multinational Corpus of Arabic Job Advertisements
📖 Overview
ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE.
It contains:
8,546 job postings
550,000+ words
Coverage across numerous sectors and dialects
Rich metadata including salary, profession, gender indicators, and job categories
This dataset supports research on:
Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.MMDA_BenchThe dataset proposed in MoDora
Some documents involving sensitive data are hidden. If you believe any content in this open source dataset infringes upon your copyright, please contact us, and we will remove it.
my-distiset-3be4288b
S-Dreamer/my-distiset-3be4288b
Overview
This synthetic dataset is designed for multiple natural language processing tasks, including Text Generation, Text2Text Generation, and Question Answering. With a lightweight size (fewer than 1K rows) and an auto-converted Parquet format, it is ideal for rapid prototyping, model development, and educational experiments.
Key Details
Modalities: Text
Format: Parquet
Size: < 1K rows
Tags: Synthetic, distilabel, rlaif, datacraft… See the full description on the dataset page: https://huggingface.co/datasets/S-Dreamer/my-distiset-3be4288b.dream
DREAM‑CFB · Dialogue-based Reading Comprehension Examination through Machine Reading (Conversation Fact Benchmark Format)
DREAM‑CFB is a 6,444 example dataset derived from the original DREAM dataset, transformed and adapted for the Conversation Fact Benchmark framework. Each item consists of multi-turn dialogues with associated multiple-choice questions that test reading comprehension and conversational understanding.
The dataset focuses on dialogue-based reading comprehension:… See the full description on the dataset page: https://huggingface.co/datasets/onionmonster/dream.Edgar-Cayce_ReadingsFinSynth_data
FinSynth_data
本数据集有三个,分别解决三个领域的问题:
客户服务聊天机器人:生成可以有效理解和回应广泛客户询问的训练数据。
欺诈检测:从交易数据中提取模式和异常,以训练可以识别和预防欺诈行为的模型。
合规监控:总结法规和合规文件,以帮助模型确保遵守金融法规。
微调大模型参考
Fintech-Dreamer/FinSynth_model_chatbot · Hugging Face
Fintech-Dreamer/FinSynth_model_fraud · Hugging Face
Fintech-Dreamer/FinSynth_model_compliance · Hugging Face
前端框架参考
Fintech-Dreamer/FinSynth
数据处理方式参考
Fintech-Dreamer/FinSynth-Data-Processing
HMO
HMO Personal Memory Evolution Benchmark (PME-L-2K)
PME-L-2K is the long-horizon retrieval benchmark accompanying
Hierarchical Memory Orchestration (HMO). It evaluates whether a persistent
agent can retain and retrieve useful evidence across 2,000 interactions per
user while avoiding an expensive full-archive search.
The release contains synthetic English personalized trajectories and a
phase-aware retrieval set. Every query is asked at the final point of a user's
history and… See the full description on the dataset page: https://huggingface.co/datasets/DreamingLow/HMO.
