datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dancedangmai2005DangDevVHcheboksary-public-transport-gps-daily
ChebTransport Daily GPS Dataset
This dataset contains daily GPS and metadata records of public transport vehicles in Cheboksary, Russia, for the period from 20.04.2025 to 21.09.2026.
Each file corresponds to a "transport day" (which may start and end at different times depending on the actual end of public transport service, not at midnight).
Data Source
The data was parsed from the website buscheb.ru, which aggregates public transport data for the city of… See the full description on the dataset page: https://huggingface.co/datasets/daniilakk/cheboksary-public-transport-gps-daily.danghuy2004dangphong1998dangquyen1986dangkhoa2006dangquynh2003danbooru-1024-eq-captioned
Danbooru 1024 e/q Captioned Dataset
59,495 high-resolution (1024px) anime-style images from Danbooru's explicit and questionable rated pools. Each image includes comprehensive JSON captions generated via MiniMax-M3 with structured per-character state-of-dress inventories, camera notes, mood palettes, and post-processing detections.
Directory Structure
danbooru-1024-eq-captioned.parquet <- consolidated metadata manifest
originals/ <-… See the full description on the dataset page: https://huggingface.co/datasets/quarterturn/danbooru-1024-eq-captioned.LSDIRDance2Hesitateviet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.danghieu2003China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.danbooru
Danbooru 2024 Dataset
Danbooru 2024 数据集
A collection of images from Danbooru website, organized and packaged by ID sequence. This dataset is for research and learning purposes only.
本数据集收集了来自 Danbooru 网站的图像,按 ID 顺序组织打包。该数据集仅用于研究和学习目的。
Dataset Description
数据集描述
This dataset contains image resources from Danbooru website, updated to ID 8380648 (Update time: 2024-11-03).
本数据集包含来自 Danbooru 网站的图像资源,更新至 ID 8380648(更新时间:2024-11-03)。
Data… See the full description on the dataset page: https://huggingface.co/datasets/picollect/danbooru.danbooru2026
Danbooru2026: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset [WIP]
Dataset Description
Danbooru2026 is a large-scale anime illustration dataset containing over 10 million community-annotated images. It is intended for research and development in anime-style image generation, image classification, multimodal learning, and related tasks.
Danbooru is a long-running image board known for its extensive tagging system and community-maintained… See the full description on the dataset page: https://huggingface.co/datasets/nyanko-devs/danbooru2026.danish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.HumanAndRobot
H&R
Dataset of paper Human2Robot: Learning Robot Actions from Paired Human-Robot Videos (AAAI 2026 Oral)
Data Details
/cam_data
/human_camera: obs of human
/robot_camera: obs of robot
/end_position: eef 6-dof position, [x, y, z, roll, pitch, yaw] represent position and orientation, where the rotation is expressed in Euler angles (degrees) with the order XYZ.
/gripper_state: 0/1, 1 for gripper open,0 for close
/action: This attribute is not available in version v0… See the full description on the dataset page: https://huggingface.co/datasets/dannyXSC/HumanAndRobot.dangle2006dangquynh2000dangthu2001dangbao2004dangthihang1995dangquanghuy1985dangthao2002dangquanghuy1992dangthuy1987dangthu2006dangminhchau01
