datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BVD-V-55M-URLs
LAION-BVD - 55M Video Clips (URL Release)
This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW.
The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect.
No video or audio files are included; only URLs, timestamps, and text annotations are provided.
Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.img_urlMS_COCO_2017_URL_TEXTYFCC15M_page_and_download_urls
YFCC15M subset used for VLMs
This dataset contains the ~15M subset of YFCC100M used for training the models in the paper Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP. The metadata provided in this repo contains both the page-urls and image-download-urls for downloading the dataset.
This dataset can be easily downloaded with img2dataset:
img2dataset --url_list yfcc15m_final_split_pageandimageurls.csv --input_format "csv" --output_format… See the full description on the dataset page: https://huggingface.co/datasets/vishaal27/YFCC15M_page_and_download_urls.preference_data_openrlhf_with_urls_len_8kcc-filtered-urlspreference_data_llama_factory_len_15k_text_with_urlsmetrics-danbooru2025-id-url-pairsMinimum version of danbooru meta for downloading images;
Saves ram and is suitable for receiving from scraper machines.
Code to reproduce:
import unibox as ub
dset = ub.loads("hf://trojblue/danbooru2025-metadata")
dset = dset.remove_columns(list(set(dset.column_names) - set(["id", "file_url"])))
ub.saves(dset, "hf://dataproc5/metircs-danbooru2025-id-url-pairs", private=False)
upper-gi-tractmetircs-danbooru2025-id-url-pairs
dataproc5/metircs-danbooru2025-id-url-pairs
(Auto-generated summary)
Basic Info:
Shape: 9113285 rows × 2 columns
Total Memory Usage: 1.14 GB
Duplicates: 1 (0.00%)
Column Stats:
Error generating stats table.
Column Summaries:
→ file_url (object)
Unique values: 8836519
→ id (int64)
Min: 1.000, Max: 9158800.000, Mean: 4590528.448, Std: 2635818.323
Sample Rows (first 3):… See the full description on the dataset page: https://huggingface.co/datasets/dataproc5/metircs-danbooru2025-id-url-pairs.preference_data_llama_factory_len_8k_text_with_urlsocr_image_urlsLOGO_URLpreference_data_llama_factory_with_urlsforniture_train_dataset_cannycc-full-image-urls
Common Crawl Image URL Batches
This gated dataset contains deduplicated image URL text files extracted from Common Crawl.
Source S3 prefix: s3://core-cn-oss.canva.com/usr/xinyang/datasets/cc-full/raw/url_batches_filtered_10m/cyrusli-cc-url-filter-10m-08121819/
URL count: 144,911,156,406
File count: 14,492
URLs per full file: 10,000,000
Final partial file URLs: 1,156,406
Compressed byte count: 8,095,053,791,842
File format: gzip-compressed text, one URL per line
Created from… See the full description on the dataset page: https://huggingface.co/datasets/xinyangli/cc-full-image-urls.bnf_journaux_urlSource gallica.bnf.fr / Bibliothèque nationale de France
preference_data_openrlhf_with_urls_len_8k_wo_checklistforniture_datasetyfcc-urlsScienceQA_Image_urldisplay-urlsdisplay_urls_truedisplay_urls_falsePokemon_Card_Image_URL_And_Captionswaon-cc-pair-url-deduplicatedds573-img-urlcast probe
laion_url_downloadpixmo_points_dataset_url_downloaded_but_maybe_htmlforniture_train_dataset_hed
