datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BVD-I-300M-URLs
LAION-BVD - 300M Video Frame URLs
This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame.
Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content.
Dataset structure
Column
Type
Description
webpage_url
string
URL of the source video
frame_pts_time
float
Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.BVD-V-55M-URLs
LAION-BVD - 55M Video Clips (URL Release)
This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW.
The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect.
No video or audio files are included; only URLs, timestamps, and text annotations are provided.
Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.gneissweb-annotation-url-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.BVD-A-10M-URLs
LAION-BVD — 10M Audio Clip URLs
This repository contains the metadata and captions for ~10 million audio clips randomly sampled
from BVD-V-55M for large-scale audio pre-training.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated
laion/BVD-A-10M repository.
Dataset structure
One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.CC_eng_urlBVD-URLs
LAION-BVD — 1.3B Video URLs
This repository contains 1.3 billion platform-specific video URLs collected from CommonCrawl. No video content is included — only URLs and associated crawl metadata.
These URLs form the source corpus for LAION-BVD (LAION — Big Video Dataset). From this collection, 80M videos were successfully downloaded, totalling approximately 10 million hours of video.
Loading the data
import datasets
ds = datasets.load_dataset("laion/BVD-URLs"… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-URLs.all-portable-apps-and-ai-in-one-urlSaving you time and space on drive!
Экономлю ваше время и место на диске!
"-cl" = clear (no models / other languages) / очишенное (без моделей / доп языков)
Моя личная подборка портативных приложений и ИИ!Перепаковывал и уменьшал размер архивов лично я!Поддержите меня: Boosty или Donationalerts
My personal selection of portable apps and AI's!I personally repacked and reduced the size of the archives!Support me: Boosty or Donationalerts
… See the full description on the dataset page: https://huggingface.co/datasets/Derur/all-portable-apps-and-ai-in-one-url.urls
URLs
74,918,894,107 deduplicated, validated URLs, sorted by
SURT key
and split into 2,334 range shards.
As plain text the URLs are 5.8 TiB, averaging 84 characters each. Sorted by
SURT key and delta-encoded they fit in 665.6 GiB — 9.54 bytes per URL, a
8.85× reduction. That is the whole point of the ordering: SURT puts URLs
from the same site next to each other, DELTA_LENGTH_BYTE_ARRAY then stores only
where each row differs from the one above it, and zstd compresses what is… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls.url-atlas
URL Atlas
257,548,097,528 URLs from 105 web corpora, each kept as
its own separately-loadable config, plus the raw source dumps two of them were
extracted from. 4.35 TiB across 52,244 files.
This is the input side of a URL-compression corpus: every source reduced to
its URL column and nothing else. It is deliberately not deduplicated or
merged — sources are kept intact and overlapping so you can measure what each
one contributes, pick the subset you want, and dedup on your own… See the full description on the dataset page: https://huggingface.co/datasets/ks48/url-atlas.BVD-A-1.7M-URLs
LAION-BVD — 1.7M Audio Clip URLs
This repository contains the metadata and captions for ~1.7 million audio clips taken from
BVD-V-55M and sampled for uniqueness of the
source video, so that the subset maximises source diversity rather than clip count.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.fineweb_urls
Dataset Card for fineweb_urls
This dataset provides the URLs and top-level domains associated with training records in HuggingFaceFW/fineweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/fineweb_urls.zyda_urls
Dataset Card for zyda_urls
This dataset provides the URLs and top-level domains associated with training records in Zyphra/Zyda. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/zyda_urls.redpajama-data-v2_urls
Dataset Card for redpajama-data-v2_urls
This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-V2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-v2_urls.hplt2.0_cleaned_urls
Dataset Card for hplt2.0_cleaned_urls
This dataset provides the URLs and top-level domains associated with training records in HPLT/HPLT2.0_cleaned. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt2.0_cleaned_urls.c4_urls_en.noblocklist
Dataset Card for c4_urls_en.noblocklist
This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (English no blocklist variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_en.noblocklist.img_urlculturax_urls
Dataset Card for culturax_urls
This dataset provides the URLs and top-level domains associated with training records in uonlp/CulturaX. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/culturax_urls.txt360_urls
Dataset Card for txt360_urls
This dataset provides the URLs and top-level domains associated with training records in LLM360/TxT360. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/txt360_urls.ccrawl-recrawl-urls
Common Crawl URL Recrawl
Live refetches of pages from Common Crawl's URL index, with the body inline and the text already extracted
What is it?
Common Crawl publishes which URLs it saw and when, but the page bodies live in WARC archives that are awkward to query and are as old as the crawl that made them. This dataset takes the URL index for a single monthly crawl and fetches the pages again now, storing each response as one Parquet row with the body, the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-urls.ccrawl-urls
Common Crawl URL Index
Every URL Common Crawl has seen, as a slim columnar table, ready to seed a crawler frontier
What is it?
This dataset is the URL-level index of Common Crawl, republished as clean Parquet. Common Crawl is a non-profit that crawls the web every month and freely publishes its archives. Each crawl ships a columnar URL index that lists every captured page with its host, fetch status, content type, detected language, and a pointer into the WARC… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-urls.malicious_urlcolossal-oscar-1.0_urls
Dataset Card for colossal-oscar-1.0_urls
This dataset provides the URLs and top-level domains associated with training records in oscar-corpus/colossal-oscar-1.0. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/colossal-oscar-1.0_urls.urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.falcon_urlsdolma_urls_v1.6
Dataset Card for dolma_urls_v1.6
This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.6.phishing_urlsc4_urls_en
Dataset Card for c4_urls_en
This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (English variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_en.c4_urls_multilingual
Dataset Card for c4_urls_multilingual
This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (multilingual variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_multilingual.Phishing_urls
Dataset Card for "Phishing_urls"
More Information needed
ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.
