datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CounterStrike-1K-360-wds
CounterStrike-1K — 360p WebDataset shards
This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets.
360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code.
Quickstart
Start a fresh uv project and add the loader:
mkdir cs1k-demo && cd cs1k-demo
uv init
uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.wds_country211wds_vtab-clevr_count_allwds_country211_test
Country-211 (Test set only)
Original paper: Learning Transferable Visual Models From Natural Language Supervision
Homepage: https://github.com/openai/CLIP/blob/main/data/country211.md
Derived from YFCC100M: https://multimediacommons.wordpress.com/yfcc100m-core-dataset/
Bibtex:
@article{DBLP:journals/corr/abs-2103-00020,
author = {Alec Radford and
Jong Wook Kim and
Chris Hallacy and
Aditya Ramesh and
Gabriel Goh and… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_country211_test.wds_vtab-clevr_count_all_test
CLEVR Count All Webdataset (Test set only)
Original paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Homepage: https://cs.stanford.edu/people/jcjohns/clevr/
Bibtex:
@article{DBLP:journals/corr/JohnsonHMFZG16,
author = {Justin Johnson and
Bharath Hariharan and
Laurens van der Maaten and
Li Fei{-}Fei and
C. Lawrence Zitnick and
Ross B. Girshick},
title =… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_vtab-clevr_count_all_test.wds_clevr_count_allwds_countbenchAS-CountingQAwds_country211counterfactual-activations
