dao
Datasets
All datasets matching “dao”Dao_taoReviewRebuttal
Introduction
This dataset is the largest real-world consistency-ensured dataset for peer review, which features the widest range of conferences and the most complete review stages, including initial submissions, reviews, ratings and confidence, aspect ratings, rebuttals, discussions, score changes, meta-reviews, and final decisions.
Paper: https://arxiv.org/abs/2505.07920
If our dataset can help you, please consider include the following citation in your publications:… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/ReviewRebuttal.Multi-SWE-bench
SWE-bench-Java: A GitHub Issue Resolving Benchmark for Java
📰 News
[Aug. 27, 2024]:We’ve released the JAVA version of SWE-bench! Check it out on Hugging Face. For more details, see our paper!
📄 Abstract
GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/Daoguang/Multi-SWE-bench.wb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.Qwen3.8-27B-Drafter-SFT
Qwen3.8-27B Drafter SFT Corpus
Supervised fine-tuning data released for training speculative drafters for Qwen/Qwen3.8-27B. All completions were generated with Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
The dataset contains 367,535 source conversations and 450,401 train rows, totaling 1,953,218,671 tokens after filtering and evaluation decontamination. Rows contain Qwen3.8-27B-tokenized prompts and target-generated completions, together with loss… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Qwen3.8-27B-Drafter-SFT.DaoZang
DaoZang — 道藏经文语料数据集
《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应,
供检索评测、微调与 RAG 使用。
数据分片
分片
行数
粒度
字段
train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片)
285,117
文本块 (chunk)
source / title / chunk_index / chars / text / embedding
标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动;
每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer
在线浏览(单次扫描上限 ~300MB);
source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.
