datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru_2025_recaption
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.AuroraCap-recaption
AuroraCap-recaption
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
Video recaption data by AuroraCap. Continue updating...
For some video source, we could upload the raw videos but for the others we could only provide the url since the well-known reason.
Citation
@article{chai2024auroracap,
title={AuroraCap: Efficient, Performant Video Detailed… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-recaption.human-recaption
Dataset Card for Human Recaption
This dataset contains 240,146 recaptioned images focusing on human subjects, derived from the HumanCaption-HQ-311K dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of OpenFace-CQUPT/HumanCaption-HQ-311K. The original dataset contained 313,482 samples. This version contains 240,146 samples; the reduction is… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/human-recaption.
