datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenCodeReasoning2DeepMath-12kdanbooru_2025_recaption
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.public-domain-poetry
Overview
This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/.
Language
The language of this dataset is English.
License
All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so.
GeneratingQuestions
HVU_QA
HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.openpi_sim_pick_place
SO-101 Pick and Place Dataset (OpenPi Format)
This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models.
Source
Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep
Format
Each episode is stored as an NPZ file containing:
Key
Shape
Type
Description
observation/state
(N, 6)
float32
Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.CREBench
CREBench
CREBench is a benchmark for evaluating large language models (LLMs) on cryptographic binary reverse engineering.
Paper: arXiv:2604.03750
Code: wangyu-ovo/CREBench
Project Page: CREBench Homepage
Dataset Description
CREBench measures reverse-engineering performance on cryptographic binaries across four evaluation levels:
Level
Task
L1
Algorithm identification
L2
Key (and IV) extraction
L3
Wrapper-level code reimplementation
L4
Flag… See the full description on the dataset page: https://huggingface.co/datasets/Danny-1223/CREBench.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.AviationLLMDans-Logicmaxx-SAT-APWhiteboards
Whiteboards
A small, single-author whiteboard corpus for evaluating vision-language
models on handwritten OCR accuracy and for studying pseudotext
hallucination — the failure mode where a VLM invents plausible-but-wrong
words for ambiguous handwriting.
Every image is the same wall-mounted whiteboard, same marker, same author,
photographed with a phone. This is deliberate: the dataset exists to
measure whether a small number of human-authored ground-truth pairs can
improve… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Whiteboards.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.phase2bench-marques
bench-marques — local LLM inference measurements
Every number here was measured on one machine, with the method written down and
the refutations kept. This dataset is the source of truth for the results;
the harness that produces them lives at
danielrmarques/bench-marques.
The machine
AMD Ryzen AI MAX+ 395 · Radeon 8060S · 128 GB unified LPDDR5X (~205 GB/s
effective, 80% of the 256 GB/s theoretical ceiling) · llama.cpp, Vulkan backend.
Every run is on AC power… See the full description on the dataset page: https://huggingface.co/datasets/danielrmarques/bench-marques.Dans-Prosemaxx-Opus-Writingnatural_reasoning_rubricstravel_hang_llama3_ttArxivBulkDataset
Arxiv Bulk Dataset
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset
ArXiv Bulk Dataset
This dataset is a bulk fetch of ArXiv articles, based on the official metadata dataset maintained and updated by Cornell https://www.kaggle.com/datasets/Cornell-University/arxiv/data.
This dataset was created to provide cross-domain academic training data, with existing datasets being domain-specific, and… See the full description on the dataset page: https://huggingface.co/datasets/dankeg/ArxivBulkDataset.anime-caption-danbooru-2021-sfw-5m-hq
Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq
Dataset Summary
This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated.
Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.Dans-Benchmaxx-COTTech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.Dans-TaskmaxxOpenMathReasoning2scientific_papers_DANCERattention-uq-800q-colab
Attention/UQ 800-question Colab bundle
A deterministic 200-question subset for each of MultiModalQA, WebQA, HotpotQA, and TAT-QA. See manifest.json for exact upstream sources, hashes, counts, and the explicitly constructed WebQA distractor setting.
seedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.MontageLie
MontageLie: Information Alignment Evaluation Benchmark
To investigate this vulnerability, we introduce MontageLie, a novel benchmark designed to test the limitations of current information alignment evaluators.
Drawing inspiration from the cinematic concept of montage, which creates new meaning by rearranging
real scenes in novel sequences, MontageLie constructs "montage-style lies": deceptive texts composed entirely of truthful statements,
deliberately reordered to imply… See the full description on the dataset page: https://huggingface.co/datasets/Dannalily/MontageLie.zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora.
Download
You can download the latest Chinese Wikipedia dump from the following link:
Chinese Wikipedia Dump
English Wikipedia Dump (For reference)
Extraction
After you download the dump, you can extract the data using the following commands:
# install wikiextractor
pip install wikiextractor
# extract the data
wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2
Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.
