datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.dragonballdaima
Bangumi Image Base of Dragon Ball Daima
This is the image base of bangumi Dragon Ball Daima, we detected 46 characters, 8351 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/dragonballdaima.DragonData-Finance-Corpus
DragonData Finance Corpus
The world's largest open, permissively-licensed finance corpus for LLM pretraining.
Built by Dragon Limited - 100% free, publicly available sources.
Overview
Attribute
Value
Repository
dragonlimited/DragonData-Finance-Corpus
Tokenizer
bigcode/starcoder2-3b (vocab 49,152)
Format
Binary shards (data/train-XXXXX-of-00001.bin, uint16)
Target
46T tokens across 25 domains
Current
~33B tokens, 246 shards
License
Permissive… See the full description on the dataset page: https://huggingface.co/datasets/dragonlimited/DragonData-Finance-Corpus.ksponspeechCIFAKE-image-datasetControlArt-Bench
ControlArt-Bench
ControlArt-Bench is a benchmark for evaluating intent-driven image
retouching and visual focus enhancement, introduced with EyeControl.
Dataset Contents
The dataset contains 200 evaluation samples.
Each sample includes:
a input image.
a user intent masks.
a retouching instructions (we support a long instruction and a short instruction).
a reference image.
Directory Structure
Replace this example with the actual dataset layout:
.… See the full description on the dataset page: https://huggingface.co/datasets/Dragoniss/ControlArt-Bench.ksponspeech_03dragonflydragonwell_filesindian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian jewelry… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/indian-traditional-artificial-jewellery.FiFA-pickapic-v2
Pick-a-Pic v2 · FiFA Filtered Subsets
These subsets were produced by filtering the original Pick-a-Pic v2 dataset using FiFA, a data filtering algorithm proposed in the paper Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models.
Overview
The filtering process is based on three key metrics:
Preference Margin: Estimated using PickScore
Text Quality: Estimated through LLM scoring
Text Diversity: Estimated using K-NN distance… See the full description on the dataset page: https://huggingface.co/datasets/Dragonjinny/FiFA-pickapic-v2.DragOn
DragOn: A Drag-Grounding Benchmark and Training Dataset for GUI Agents
Drag-grounding dataset for GUI agents. Each example = one screenshot + one instruction +
start/end bounding boxes for the drag. Four domains: text_highlight, sheet (cell
selection), slide_resize (element resize/rotate/crop), slider.
Dataset statistics
Domain
Train images
Train tasks
Eval (public)
Dominant resolution
Text Highlighting
100,000
1,000,000
250
1275×1650 (portrait)
Cell… See the full description on the dataset page: https://huggingface.co/datasets/Hcompany/DragOn.ksponspeech_05ksponspeech_04LandCover.ai
The Dataset
The LandCover.ai (Land Cover from Aerial Imagery) dataset is a dataset for automatic mapping of buildings, woodlands, water and roads from aerial images.
Dataset features
land cover from Poland, Central Europe (1)
three spectral bands - RGB
33 orthophotos with 25 cm per pixel resolution (~9000x9500 px)
8 orthophotos with 50 cm per pixel resolution (~4200x4700 px)
total area of 216.27 km2
Dataset format
rasters are three-channel GeoTiffs with… See the full description on the dataset page: https://huggingface.co/datasets/dragon7/LandCover.ai.ksponspeech_04_preprocessOurData22026-09-03_shake4it_bench_5sensors_v3_dragonfly_10kHz_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-03_shake4it_bench_5sensors_v3_dragonfly_10kHz_nfft_512.2026-09-02_shake4it_bench_5sensors_v2_dragonfly_10kHz_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-02_shake4it_bench_5sensors_v2_dragonfly_10kHz_nfft_512.details_ibivibiv__alpaca-dragon-72b-v1
Dataset Card for Evaluation run of ibivibiv/alpaca-dragon-72b-v1
Dataset automatically created during the evaluation run of model ibivibiv/alpaca-dragon-72b-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ibivibiv__alpaca-dragon-72b-v1.kniv-corpus-en
kniv-corpus-en
A multi-domain English NLP corpus with four annotation layers: Named Entity Recognition (18 types), POS tagging (17 UPOS tags), dependency parsing, and dialog act classification (9 types). All data is commercially licensed (CC BY-SA 4.0 compatible).
Built for training kniv multi-task NLP models that power the uniko cognitive memory system.
Quick Start
from datasets import load_dataset
# Load the full corpus via HuggingFace
ds =… See the full description on the dataset page: https://huggingface.co/datasets/dragonscale-ai/kniv-corpus-en.dragonfly2026-08-31_shake4it_bench_5sensors_dragonfly_10kHz_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-08-31_shake4it_bench_5sensors_dragonfly_10kHz_nfft_512.ksponspeech_preprocessDragonTCM
DragonTCM Dataset
Dataset Description
Overview
DragonTCM is a comprehensive Traditional Chinese Medicine (TCM) knowledge base derived from the American Dragon website (www.americandragon.com) and the foundational work of Dr. Joel Penner. The dataset captures the complex relationships between TCM conditions, herbal formulas, and individual herbs, making it valuable for both research and educational purposes in traditional medicine.
This dataset is an incomplete… See the full description on the dataset page: https://huggingface.co/datasets/f-galkin/DragonTCM.2026-08-04_shake4it_bench_dragonfly_10kHz_nfft_512_big_croppedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-08-04_shake4it_bench_dragonfly_10kHz_nfft_512_big_cropped.so100_sorting_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 56929,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dragon-95/so100_sorting_1.dragon-ai-vector-embeddingshotpot_qa
Dataset Card for "hotpot_qa"
Dataset Summary
HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason… See the full description on the dataset page: https://huggingface.co/datasets/dragonprogramer/hotpot_qa.
