datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dao_taoReviewRebuttal
Introduction
This dataset is the largest real-world consistency-ensured dataset for peer review, which features the widest range of conferences and the most complete review stages, including initial submissions, reviews, ratings and confidence, aspect ratings, rebuttals, discussions, score changes, meta-reviews, and final decisions.
Paper: https://arxiv.org/abs/2505.07920
If our dataset can help you, please consider include the following citation in your publications:… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/ReviewRebuttal.Multi-SWE-bench
SWE-bench-Java: A GitHub Issue Resolving Benchmark for Java
📰 News
[Aug. 27, 2024]:We’ve released the JAVA version of SWE-bench! Check it out on Hugging Face. For more details, see our paper!
📄 Abstract
GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/Daoguang/Multi-SWE-bench.wb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.Qwen3.8-27B-Drafter-SFT
Qwen3.8-27B Drafter SFT Corpus
Supervised fine-tuning data released for training speculative drafters for Qwen/Qwen3.8-27B. All completions were generated with Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
The dataset contains 367,535 source conversations and 450,401 train rows, totaling 1,953,218,671 tokens after filtering and evaluation decontamination. Rows contain Qwen3.8-27B-tokenized prompts and target-generated completions, together with loss… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Qwen3.8-27B-Drafter-SFT.DaoZang
DaoZang — 道藏经文语料数据集
《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应,
供检索评测、微调与 RAG 使用。
数据分片
分片
行数
粒度
字段
train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片)
285,117
文本块 (chunk)
source / title / chunk_index / chars / text / embedding
标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动;
每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer
在线浏览(单次扫描上限 ~300MB);
source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.g1_desktop_organize_table_correction
g1_desktop_organize_table_correction
Bimanual tabletop teleoperation on a Unitree G1 with Inspire dexterous hands,
in LeRobot v2.1 format.
Task — Follow the human's corrective gesture and hand the indicated object to the human.
Episodes
192
Frames
67405 (37.4 min @ 30 fps)
Episode length
268–519 frames (median 346)
State / action
26-D / 26-D
Cameras
2 × 640×480
State and action
Both vectors are 26-D and share the same layout:
0– 6… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/g1_desktop_organize_table_correction.dao_tam_audioMAEtoanmath.com-alphawb_add_push_cart
wb_add_push_cart
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_add_push_cart
Task — push the cart, stop when the human raises a hand, and continue when the human waves again… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_add_push_cart.wb_push_cart_stop_go
wb_push_cart_stop_go
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_push_cart_stop_go
Task — push the cart, stop when the human raises a hand, and continue when the human waves… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_push_cart_stop_go.CodeM-Multilinugal-Data
CodeM: Can Programming Languages Boost Each Other via Instruction Tuning?
Paper GitHub
Abstract
When human programmers have mastered a programming language, it would be easier when they learn a new programming language. In this report, we focus on exploring whether programming languages can boost each other during the instruction fine-tuning phase of code large language models. We conduct extensive experiments of 8 popular programming languages (Python, JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/Daoguang/CodeM-Multilinugal-Data.mangadex-images-30kdao_tam_truyenarchive_stuff_no1Robot-Feeding
Robot-Feeding
Wheeled humanoid robot model for feeding assistance research.Contains URDF, STL meshes, and Isaac Lab USD assets.
Robot Overview
Type: Wheeled humanoid (differential drive base + dual 7-DOF arms)
Total DOF: 20 (waist ×3, chest ×1, neck ×1, head ×1, left arm ×7, right arm ×7)
Base: Two-wheeled chassis with integrated wheel geometry (chassis.STL)
Joint List
#
Joint
Group
1
waist_link1
Waist
2
waist_link2
Waist
3… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/Robot-Feeding.synthetic-seq-labelling-vi-exam-v2
Vietnamese Exam Sequence Labelling Dataset (v2)
A comprehensive, production-grade dataset for token-level and span-level sequence labelling of Vietnamese educational examination documents (covering grades 8–12 across Mathematics, Physics, Chemistry, Biology, History, Geography, Literature, and English).
Generated and curated by the Vietnamese Sequence Labelling v2 pipeline, combining real OCR-annotated examination documents (both scanned and digital PDFs) with synthetic… See the full description on the dataset page: https://huggingface.co/datasets/daominhwysi/synthetic-seq-labelling-vi-exam-v2.openfwiMuse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.MM-Bench-E-CommerceThis is the HuggingFace repository of the paper named MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding in WSDM 2026 (oral).
In this paper, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning.
We propose the first generative MLLM-based model named MOON for product representation learning.
Furthermore, we contruct and publish a large-scale real-world… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/MM-Bench-E-Commerce.bsbidaibisaismol-smoltalk-sftlang-adversarial-inference-01
Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse
An exploration on how to secure Model-as-a-Service infrastructure at the inference layer against runtime exploits, ranging from automated bot farms to adversarial extraction and agentic misuse
More details about the project: https://www.daoist.dev/posts/adversarial-inference-security-1
CondAmbigQA-2K
CondAmbigQA-2K Dataset
Dataset Description
This is an expanded version of the CondAmbigQA dataset, growing from the original 200 entries to 2000 entries.
Dataset Summary
CondAmbigQA-2K contains 2000 question-answering pairs with conditional contexts and ground truth answers. Each entry includes:
Question: The ambiguous question
Properties: Contains condition, groundtruth, and citations
Context (ctxs): Retrieved relevant passages with scores
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Apocalypse-AGI-DAO/CondAmbigQA-2K.tieu_dao_cacx23-DAO-governance-forums
X23 Governance Forum Dataset
The X23 Governance Forum Dataset is a normalized corpus of public DAO governance forum discussions collected from publicly accessible governance forums, covering discussions from September 2015 through September 2025. It includes topics, posts, public author profiles, timestamps, engagement/reputation fields, tags, summaries, extracted links, and source forum/domain metadata.
The dataset covers 44 public forum domains across 42 protocol/community labels… See the full description on the dataset page: https://huggingface.co/datasets/daveytea/x23-DAO-governance-forums.grad-news-media
News Media Image Text Data Notes
Dataset summary
This data card accompanies a lightweight News Media loader for Image Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/daoliveirana/grad-news-media.
