hdhacker/connan_30k
Connan 30K Video Caption Dataset Private Hugging Face backup of 30,617 short anime video clips and English visual captions. Layout metadata/manifest.parquet # sample index, captions, original paths, shard/member mapping wds/*.tar # WebDataset shards, about 10GB each dataset_info.json # summary metadata Each WebDataset sample uses a stable key such as 001_shot_00003 and contains: 001_shot_00003.mp4 001_shot_00003.json The JSON… See the full description on the dataset page: https://huggingface.co/datasets/hdhacker/connan_30k.
Connan 30K Video Caption Dataset
Private Hugging Face backup of 30,617 short anime video clips and English visual captions.
Layout
metadata/manifest.parquet # sample index, captions, original paths, shard/member mapping
wds/*.tar # WebDataset shards, about 10GB each
dataset_info.json # summary metadataEach WebDataset sample uses a stable key such as 001_shot_00003 and contains:
001_shot_00003.mp4
001_shot_00003.jsonThe JSON sidecar contains the caption and original source paths. The Parquet manifest is the canonical table for captions and shard lookup.
Summary
- Samples: 30617
- Shards: 10
- Caption language: English
- Recommended LoRA trigger:
DCANIME
Captions are generic visual descriptions and intentionally avoid character names, franchise names, and trigger words.
