CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /audiofolder_two_configs_in_metadataaudion<1K1 likes101k downloads3y agoHugging Face02hf-internal-testing /audiofolder_single_config_in_metadataaudion<1K0 likes93k downloads3y agoHugging Face03hf-internal-testing /imagefolder_with_metadataimagen<1K0 likes51k downloads2y agoHugging Face04hf-internal-testing /audiofolder_no_configs_in_metadataaudion<1K0 likes49k downloads3y agoHugging Face05polinaeterna /audiofolder_two_configs_in_metadataaudion<1K0 likes20k downloads3y agoHugging Face06adollahamini1998 /voxblink2-metadata0 likes20k downloads16d agoHugging Face07model-metadata /code_execution_filestextn<1K0 likes13k downloads7mo agoHugging Face08model-metadata /code_python_files0 likes13k downloads7mo agoHugging Face09hf-internal-testing /audiofolder_two_configs_in_metadata_with_defaultaudion<1K0 likes12k downloads3y agoHugging Face10hf-internal-testing /imagefolder_with_metadata_no_splitsimagen<1K0 likes9.2k downloads3y agoHugging Face11saeedzouashkiani /yodas_fa_nosub_metadatatabular100K<n<1M0 likes9k downloads28d agoHugging Face12model-metadata /custom_code_py_files1 likes6.4k downloads11mo agoHugging Face13webshart /conceptual-captions-12m-webdataset-metadata Conceptual Captions 12M — Webshart metadata indices Per-shard webshart metadata indices for laion/conceptual-captions-12m-webdataset: 1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout. Each index records every tar member's byte offset and length (enabling ranged reads without downloading whole shards), image geometry (width/height for aspect bucketing), and — as of August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.1 likes4.8k downloads1mo agoHugging Face14librarian-bots /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.texttext-generation1M<n<10M22 likes4.7k downloads18h agoHugging Face15bigcode /the-stack-metadata Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.tabulartext-generation10B<n<100B10 likes4.7k downloads4y agoHugging Face16model-metadata /custom_code_execution_filestext1K<n<10K1 likes3.2k downloads11mo agoHugging Face17huggingface /transformers-metadata Transformers metadata text1K<n<10K43 likes3k downloads14h agoHugging Face18bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.8k downloads4y agoHugging Face19wallstoneai /civitai-top-nsfw-images-with-metadata CivitAI Top NSFW Images Dataset This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X. Original forum post: https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images Dataset collection date June 2025 Dataset structure: ├── 📂 images/ │ ├── 1.jpg │ ├── 2.jpg │ ├── 3.jpg │ ├── .... ├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-nsfw-images-with-metadata.imageimage-classification1K<n<10K69 likes2k downloads1y agoHugging Face20labofsahil /pypi-packages-metadata-datasettext10M<n<100M0 likes2k downloads8mo agoHugging Face21huggingface /diffusers-metadatatextn<1K34 likes2k downloads8h agoHugging Face22trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.8k downloads1y agoHugging Face23dmmagdal /enwiki-2024-04-20-metadata0 likes1.7k downloads2y agoHugging Face24silatus /1k_Website_Screenshots_and_Metadata Dataset Card for 1000 Website Screenshots with Metadata Dataset Summary Silatus is sharing, for free, a segment of a dataset that we are using to train a generative AI model for text-to-mockup conversions. This dataset was collected in December 2022 and early January 2023, so it contains very recent data from 1,000 of the world's most popular websites. You can get our larger 10,000 website dataset for free at: https://silatus.com/datasets This dataset includes: High-res… See the full description on the dataset page: https://huggingface.co/datasets/silatus/1k_Website_Screenshots_and_Metadata.imagetext-to-image1K<n<10K20 likes1.4k downloads4y agoHugging Face25smartcat /Amazon_Sample_Metadata_2023 Dataset Card for Dataset Name Original datasets can be found on: https://amazon-reviews-2023.github.io/ Dataset Details This dataset was made as sample of several datasets from the link above. Dataset Description This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors, Health and Personal Care, Amazon Clothing Shoes and Jewlery, Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.tabular1M<n<10M1 likes1.3k downloads2y agoHugging Face26mmarone /fineweb-edu-full-metadata[WIP] FineWeb-Edu with Metadata This repo contains 3 versions of the FineWeb-Edu v1 dataset: fwedu1-metaonly/ fwedu1-text-content-zstd/ fineweb-edu-1.0.0-meta-and-text/ These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.tabular100M<n<1B0 likes1.3k downloads1y agoHugging Face27stma /danbooru-metadata danbooru-metadata Dump of various portions danbooru's metadata, as of Feburary 2023. Everything was taken directly from their JSON API. The directory structure follows the Danbooru20XX format of each subfolder for a record type being the record's ID modulo 1000. The .zip files can sometimes hold hundreds of thousands of small JSON files when put together, so use caution when extracting. 2 likes1.3k downloads4y agoHugging Face28laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes1.3k downloads2y agoHugging Face29amazon-sagemaker /repository-metadata1 likes1.2k downloads17h agoHugging Face30ReadyAi /5000-podcast-conversations-with-metadata-and-embedding-dataset 🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications. This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network. AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.text10K<n<100K8 likes1.2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.