CoolFace
Datasetpublic

hdhacker/connan_30k

Connan 30K Video Caption Dataset Private Hugging Face backup of 30,617 short anime video clips and English visual captions. Layout metadata/manifest.parquet # sample index, captions, original paths, shard/member mapping wds/*.tar # WebDataset shards, about 10GB each dataset_info.json # summary metadata Each WebDataset sample uses a stable key such as 001_shot_00003 and contains: 001_shot_00003.mp4 001_shot_00003.json The JSON… See the full description on the dataset page: https://huggingface.co/datasets/hdhacker/connan_30k.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes13downloads
Dataset Card

Connan 30K Video Caption Dataset

Private Hugging Face backup of 30,617 short anime video clips and English visual captions.

Layout

text
metadata/manifest.parquet   # sample index, captions, original paths, shard/member mapping
wds/*.tar                   # WebDataset shards, about 10GB each
dataset_info.json           # summary metadata

Each WebDataset sample uses a stable key such as 001_shot_00003 and contains:

text
001_shot_00003.mp4
001_shot_00003.json

The JSON sidecar contains the caption and original source paths. The Parquet manifest is the canonical table for captions and shard lookup.

Summary

  • Samples: 30617
  • Shards: 10
  • Caption language: English
  • Recommended LoRA trigger: DCANIME

Captions are generic visual descriptions and intentionally avoid character names, franchise names, and trigger words.