CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /BVD-I-300M-URLs LAION-BVD - 300M Video Frame URLs This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame. Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content. Dataset structure Column Type Description webpage_url string URL of the source video frame_pts_time float Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.textimage-text-to-text100M<n<1B2 likes22k downloads26d agoHugging Face02laion /BVD-V-55M-URLs LAION-BVD - 55M Video Clips (URL Release) This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW. The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect. No video or audio files are included; only URLs, timestamps, and text annotations are provided. Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.imagevideo-text-to-text10M<n<100M5 likes16k downloads26d agoHugging Face03commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face04laion /BVD-A-10M-URLs LAION-BVD — 10M Audio Clip URLs This repository contains the metadata and captions for ~10 million audio clips randomly sampled from BVD-V-55M for large-scale audio pre-training. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated laion/BVD-A-10M repository. Dataset structure One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.tabulartext-to-audio10M<n<100M1 likes7.5k downloads26d agoHugging Face05leiwx52 /CC_eng_urltext100M<n<1B0 likes7.1k downloads2y agoHugging Face06laion /BVD-URLs LAION-BVD — 1.3B Video URLs This repository contains 1.3 billion platform-specific video URLs collected from CommonCrawl. No video content is included — only URLs and associated crawl metadata. These URLs form the source corpus for LAION-BVD (LAION — Big Video Dataset). From this collection, 80M videos were successfully downloaded, totalling approximately 10 million hours of video. Loading the data import datasets ds = datasets.load_dataset("laion/BVD-URLs"… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-URLs.text1B<n<10B12 likes5.8k downloads26d agoHugging Face07Derur /all-portable-apps-and-ai-in-one-urlSaving you time and space on drive! Экономлю ваше время и место на диске! "-cl" = clear (no models / other languages) / очишенное (без моделей / доп языков) Моя личная подборка портативных приложений и ИИ!Перепаковывал и уменьшал размер архивов лично я!Поддержите меня: Boosty или Donationalerts My personal selection of portable apps and AI's!I personally repacked and reduced the size of the archives!Support me: Boosty or Donationalerts &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/Derur/all-portable-apps-and-ai-in-one-url.13 likes5.8k downloads23d agoHugging Face08ks46 /urls URLs 74,918,894,107 deduplicated, validated URLs, sorted by SURT key and split into 2,334 range shards. As plain text the URLs are 5.8 TiB, averaging 84 characters each. Sorted by SURT key and delta-encoded they fit in 665.6 GiB — 9.54 bytes per URL, a 8.85× reduction. That is the whole point of the ordering: SURT puts URLs from the same site next to each other, DELTA_LENGTH_BYTE_ARRAY then stores only where each row differs from the one above it, and zstd compresses what is… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls.texttext-generation10B<n<100B1 likes3k downloads1mo agoHugging Face09ks48 /url-atlas URL Atlas 257,548,097,528 URLs from 105 web corpora, each kept as its own separately-loadable config, plus the raw source dumps two of them were extracted from. 4.35 TiB across 52,244 files. This is the input side of a URL-compression corpus: every source reduced to its URL column and nothing else. It is deliberately not deduplicated or merged — sources are kept intact and overlapping so you can measure what each one contributes, pick the subset you want, and dedup on your own… See the full description on the dataset page: https://huggingface.co/datasets/ks48/url-atlas.texttext-retrieval100B<n<1T0 likes2.9k downloads1mo agoHugging Face10laion /BVD-A-1.7M-URLs LAION-BVD — 1.7M Audio Clip URLs This repository contains the metadata and captions for ~1.7 million audio clips taken from BVD-V-55M and sampled for uniqueness of the source video, so that the subset maximises source diversity rather than clip count. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.tabulartext-to-audio1M<n<10M1 likes2.6k downloads26d agoHugging Face11nhagar /fineweb_urls Dataset Card for fineweb_urls This dataset provides the URLs and top-level domains associated with training records in HuggingFaceFW/fineweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/fineweb_urls.texttext-generation10B<n<100B2 likes2.4k downloads1y agoHugging Face12nhagar /zyda_urls Dataset Card for zyda_urls This dataset provides the URLs and top-level domains associated with training records in Zyphra/Zyda. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/zyda_urls.texttext-generation1B<n<10B0 likes2k downloads1y agoHugging Face13nhagar /redpajama-data-v2_urls Dataset Card for redpajama-data-v2_urls This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-V2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-v2_urls.text1B<n<10B0 likes1.7k downloads1y agoHugging Face14nhagar /hplt2.0_cleaned_urls Dataset Card for hplt2.0_cleaned_urls This dataset provides the URLs and top-level domains associated with training records in HPLT/HPLT2.0_cleaned. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt2.0_cleaned_urls.text10B<n<100B0 likes1.4k downloads1y agoHugging Face15nhagar /c4_urls_en.noblocklist Dataset Card for c4_urls_en.noblocklist This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (English no blocklist variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_en.noblocklist.texttext-generation100M<n<1B1 likes1.4k downloads1y agoHugging Face16Liskk /img_urlimagen<1K0 likes1.2k downloads2mo agoHugging Face17nhagar /culturax_urls Dataset Card for culturax_urls This dataset provides the URLs and top-level domains associated with training records in uonlp/CulturaX. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/culturax_urls.texttext-generation1B<n<10B0 likes848 downloads1y agoHugging Face18nhagar /txt360_urls Dataset Card for txt360_urls This dataset provides the URLs and top-level domains associated with training records in LLM360/TxT360. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/txt360_urls.text1M<n<10M0 likes822 downloads1y agoHugging Face19open-index /ccrawl-recrawl-urls Common Crawl URL Recrawl Live refetches of pages from Common Crawl's URL index, with the body inline and the text already extracted What is it? Common Crawl publishes which URLs it saw and when, but the page bodies live in WARC archives that are awkward to query and are as old as the crawl that made them. This dataset takes the URL index for a single monthly crawl and fetches the pages again now, storing each response as one Parquet row with the body, the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-urls.tabulartext-generation1M<n<10M0 likes708 downloads1mo agoHugging Face20open-index /ccrawl-urls Common Crawl URL Index Every URL Common Crawl has seen, as a slim columnar table, ready to seed a crawler frontier What is it? This dataset is the URL-level index of Common Crawl, republished as clean Parquet. Common Crawl is a non-profit that crawls the web every month and freely publishes its archives. Each crawl ships a columnar URL index that lists every captured page with its host, fetch status, content type, detected language, and a pointer into the WARC… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-urls.tabulartext-retrieval1B<n<10B0 likes689 downloads2mo agoHugging Face21surajshelke /malicious_urltext100K<n<1M0 likes644 downloads2y agoHugging Face22nhagar /colossal-oscar-1.0_urls Dataset Card for colossal-oscar-1.0_urls This dataset provides the URLs and top-level domains associated with training records in oscar-corpus/colossal-oscar-1.0. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/colossal-oscar-1.0_urls.text10B<n<100B0 likes621 downloads1y agoHugging Face23ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes620 downloads16d agoHugging Face24nhagar /falcon_urlstext100M<n<1B1 likes604 downloads2y agoHugging Face25nhagar /dolma_urls_v1.6 Dataset Card for dolma_urls_v1.6 This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.6.text1B<n<10B0 likes581 downloads1y agoHugging Face26alexkstern /phishing_urlstext100K<n<1M4 likes563 downloads3y agoHugging Face27nhagar /c4_urls_en Dataset Card for c4_urls_en This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (English variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_en.texttext-generation100M<n<1B0 likes497 downloads1y agoHugging Face28nhagar /c4_urls_multilingual Dataset Card for c4_urls_multilingual This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (multilingual variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_multilingual.texttext-generation1B<n<10B1 likes467 downloads1y agoHugging Face29kmack /Phishing_urls Dataset Card for "Phishing_urls" More Information needed text100K<n<1M5 likes446 downloads2y agoHugging Face30latmay /ats-career-page-urls ATS Career Page URLs 69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR. Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines. Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.text10K<n<100K1 likes429 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.