Khoale11hcmut/gpt4o_captions_1k5samples
PACO WebDataset export for PACO-style localized caption data. Summary Samples: 1500 Shards: 2 Payload format inside each shard: pickle Split: train Config: PACO Layout Media files are stored in WebDataset tar shards. Each sample key is stable and becomes __key__ in the dataset viewer. Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents. Manifest file: PACO/annotations.json
15
