datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_objectnetwds_imagenet_sketchwds_imagenet-rshiur-clips-flacwds_imagenet-ainstructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.wds_imagenetv2hdr-demo-clips
HDR Demo Clips (Lightricks SDR→HDR)
Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2).
Each clip contains:
hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred)
sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred)
thumbnail.jpg — 280px preview from the middle frame
Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.beir-nl-cqadupstack
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.wds_imagenet1ksea-clip-eval-predictionswds_fer2013wan22-processed-clipswds_flickr30kwds_carswds_vtab-eurosatwds_vtab-caltech101wds_fgvc_aircraftwds_vtab-cifar10wds_vtab-dtdwds_vtab-petswds_vtab-cifar100xtd_11
Dataset Summary
The expanded XTD-11 dataset, now including Arabic, enhances the original XTD collection. This dataset introduces a 1,000-image multi-lingual MSCOCO2014 caption to test multimodel in zeroshot image or text retrieval in 11 Languages.
Dataset Details
Citation
@misc{aggarwal2020zeroshot,
title={Towards Zero-shot Cross-lingual Image Retrieval},
author={Pranav Aggarwal and Ajinkya Kale},
year={2020},
eprint={2012.05107}… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/xtd_11.ucf-crime-clip-featurescode-clippy-tfrecordsclipdemo_openai_clip_index
CLIP index — demo_openai_clip
Precomputed image embeddings (openai/clip-vit-base-patch32) for the static Space
diegoolguinw/demo_openai_clip.
Current contents: 9000 images from the train split of
detection-datasets/coco.
(HF datasets cap a directory at 10 000 files, so thumbs/ stays below that.)
File
Description
embeddings.f16.bin
[N, 512] row-major float16, L2-normalized
metadata.json
index-aligned list: {id, file, thumb, width, height}
manifest.json
model, count… See the full description on the dataset page: https://huggingface.co/datasets/diegoolguinw/demo_openai_clip_index.mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.
